HappyHorse 1.0 pricing, plans and limits

Alibaba's joint audio-video generation model, built for short-form ad, e-commerce and social clips that need built-in lip-synced dialogue.

  • ga
  • proprietary
  • multimodal
  • HappyHorse family
checked

Alibaba's ATH Business Unit shipped HappyHorse 1.0 in limited beta on April 28, 2026, and its own benchmark clocks a full clip at about 38 seconds on a single H100 GPU. Teams building ad creative, e-commerce product videos or short-form social clips get lip-synced audio and motion baked in without pairing a separate voice model.

HappyHorse 1.0 is Alibaba's video generation model, built by its ATH Business Unit, that renders text, image and reference prompts into synchronized video and audio in one forward pass. It ranked #1 in Artificial Analysis's April 2026 Video Generation Arena for text-to-video and image-to-video without audio.

Provider: Alibaba Cloud · Family: HappyHorse

More about Alibaba Cloud on HokAI

Input modalities: text, image, video · Output: video, audio

About HappyHorse 1.0

HappyHorse 1.0 is a large-scale, 40-layer self-attention Transformer built by Alibaba's ATH Business Unit (Alibaba Token Hub, formed March 2026 under CEO Eddie Wu, combining Tongyi Lab, the Qwen unit, the Wukong unit, the MaaS business line and an AI innovation unit). Alibaba Cloud rolled it out in limited beta on April 28, 2026, positioning it as a joint audio-video generator: rather than rendering a silent clip and dubbing it afterward, the model produces video frames and their matching audio track (dialogue, ambient sound, Foley) in a single forward pass across a unified text/image/video/audio token sequence.

The model appeared anonymously on Artificial Analysis's Video Generation Arena on April 7, 2026 and took the top spot within days in both the text-to-video and image-to-video no-audio categories, ahead of ByteDance's Seedance 2.0 and Kuaishou's Kling 3.0 at the time; it placed second in the audio-enabled variant of both categories, with Elo scores of 1205 and 1161. Alibaba did not confirm authorship until April 10, 2026, after a since-verified account referenced the ATH unit; Alibaba's own press did not name an individual project lead.

Four generation modes are exposed as separate endpoints on Alibaba Cloud Model Studio and through fal.ai, the model's official third-party API partner: text-to-video, image-to-video (first frame), reference-to-video (multiple reference images plus a prompt), and natural-language video editing. The model natively lip-syncs generated dialogue across seven languages, spanning three East Asian languages and English, German and French.

Alibaba's June 23, 2026 HappyHorse 1.1 update made native audio fully available across every endpoint (1.0's production API had not exposed it everywhere at launch), added multi-image reference input, and rebuilt the motion system for smoother, more physically grounded movement. Alibaba Cloud Model Studio now recommends 1.1 by default for text-to-video and image-to-video, while its own documentation still points to the 1.0 endpoint for video editing specifically. HappyHorse 1.0 stays live and billed on both fal.ai and Alibaba Cloud Model Studio as of this writing.

Alibaba has stated an intention to release HappyHorse under an Apache 2.0 open-weight license, but no independently downloadable weights exist: the public HuggingFace repository has returned an authentication error since at least April 2026, unchanged as of this writing. In practice the model is API-only, billed per second of rendered output rather than per token.

Teams choosing between Alibaba's own video lineup and outside rivals should weigh the tradeoffs directly: Wan 3.0, Alibaba Tongyi Lab's other video model, targets longer single continuous shots at the cost of HappyHorse 1.0's native lip-synced dialogue; Wan Animate 2 is Alibaba's separate character-animation line built for a narrower motion-transfer job; Kling 3.0 counters with 4K capture and a wider set of launch-day audio languages; Veo 3.1 answers with tighter Google Cloud integration and richer conversational audio; and MiniMax H3 is the pick for teams that need downloadable, self-hostable weights, which HappyHorse 1.0 does not offer.

HappyHorse sits alongside Alibaba's Qwen3.8-Max text model under the same ATH Business Unit; both ship through Alibaba Cloud's Model Studio console rather than as separate vendor accounts. It is listed on hokai's Alibaba Cloud model directory and among proprietary video models.

No published system card, red-team partner list, or public training-data disclosure accompanies HappyHorse 1.0; Alibaba Cloud Model Studio's general acceptable-use terms govern the endpoint, and no HappyHorse-specific content-moderation document was found at the time of writing.

Screenshots

fal.ai's official HappyHorse-1.0 product page hero section reading The Top Ranked AI Video Model, with the number one Artificial Analysis Video Arena ranking claim
fal.ai's official HappyHorse-1.0 product page (fal.ai/happyhorse-1.0)

Pricing

Billed per second of rendered video, not per token. fal.ai (Alibaba's official API partner) and Alibaba Cloud Model Studio both bill pay-as-you-go with no subscription or minimum spend; exact per-second rates by resolution are in the pricing table above.

TierRate
fal.ai API - 720p$0.14 per second of generated video
fal.ai API - 1080p$0.28 per second of generated video
Alibaba Cloud Model StudioPay-as-you-go, billed per second; console rate for the 1.0 endpoint specifically was not separately published

Key Features

  • Joint Audio-Video Generation: Renders video frames and their audio track, including dialogue, ambience and Foley, in a single forward pass through a 15-billion-parameter, 40-layer Transformer rather than dubbing a silent clip afterward.
  • Seven-Language Native Lip-Sync: Matches mouth movement to generated dialogue in Mandarin, Cantonese, English, Japanese, Korean, German and French without a separate voice-sync pass.
  • Four Generation Modes: Covers text-to-video, image-to-video, reference-to-video and natural-language video editing, each exposed as its own Alibaba Cloud Model Studio endpoint.
  • 1080p in About 38 Seconds: Produces a full 1080p clip of up to 15 seconds in roughly 38 seconds on a single Nvidia H100 GPU, per Alibaba's own published benchmark.
  • Anonymous Arena Debut: Surfaced under no vendor name on a public leaderboard on April 7, 2026 and reached the top rank within days, before Alibaba confirmed authorship three days later.

Pros

  • Reached the top of Artificial Analysis's video arena for text-to-video and image-to-video without audio, within days of an anonymous debut.
  • Generates audio and video together in one pass, avoiding the lip-sync drift that separate-pipeline video and dubbing tools introduce.
  • Bills per second of rendered video with no subscription commitment, unlike consumer-credit competitors.

Cons

  • Clips are capped at 15 seconds, well short of Wan 3.0's longer single-shot ceiling.
  • Full native audio was not exposed on every 1.0 production endpoint at launch; a later HappyHorse 1.1 update completed that rollout.
  • No independently downloadable weights exist despite Alibaba's stated Apache 2.0 open-source plan; the public HuggingFace repository still returns an authentication error.

Benchmarks

  • LMArena Elo: 1392 cited: Artificial Analysis · 26 Sep 2026 — Rating from blind human votes on which answer is better.
  • LMArena rank: #1 cited: Artificial Analysis · 26 Sep 2026 — Position on the blind human-preference leaderboard; #1 is best.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

What are HappyHorse 1.0's pricing plans in 2026?

Alibaba's official API partner fal.ai bills HappyHorse 1.0 per second of rendered video: $0.14 for 720p and $0.28 for 1080p, with no subscription or minimum spend. Enterprise volume pricing is available on request through fal.ai or directly via Alibaba Cloud Model Studio's console.

Is HappyHorse 1.0 free to use?

No published free tier exists for HappyHorse 1.0's API access; generation is billed from the first second on every official channel. Some third-party resale sites advertise trial credits, but Alibaba's own channels start billing immediately.

What should you use instead of HappyHorse 1.0?

Kling 3.0 and Veo 3.1 are the closest rivals on the same Artificial Analysis video arena, both shipping native audio at launch. Alibaba's own Wan 3.0 is the pick for longer single-shot clips, and MiniMax H3 is the option for teams that need open, downloadable weights.

What separates HappyHorse 1.0 from Kling 3.0?

HappyHorse 1.0 topped Artificial Analysis's video arena in April 2026 for text-to-video and image-to-video without audio (Elo 1333 and 1392), a category Kling 3.0's own June 2026 arena snapshot did not lead (Elo 1248). Kling 3.0 counters with a higher native resolution ceiling and a broader launch-day audio language set.

How long does it take to get going with HappyHorse 1.0?

Sign up for an API key on fal.ai or provision access through Alibaba Cloud Model Studio, then call the text-to-video, image-to-video or reference-to-video endpoint with a prompt and optional reference images. A first working call can return a finished clip in well under a minute once the key is issued.

Top Alternatives

  • Kling 3.0: Pick Kling 3.0 if you need a higher native resolution ceiling; pick HappyHorse 1.0 if you want the model that topped Artificial Analysis's no-audio video arena categories at launch.
  • Veo 3.1: Pick Veo 3.1 if richer conversational native audio and Google Cloud integration matter more than HappyHorse 1.0's per-second fal.ai pricing.
  • Wan 3.0: Pick Wan 3.0 if you need longer single-shot clips than HappyHorse 1.0 supports; HappyHorse 1.0 answers back with native lip-synced dialogue Wan 3.0 does not build in.
  • MiniMax H3: Pick MiniMax H3 if open weights and self-hosting outrank HappyHorse 1.0's arena-topping motion quality and managed API convenience.

More AI Models on HokAI

Visit HappyHorse 1.0 Official Page