Last updated: 2026-09-17
Vidu, from ShengShu Technology, has over 10 million users worldwide generating AI video. Its standout feature is Reference-to-Video: feed it several images and it keeps the same face, object, or setting consistent across a whole clip. The Q3 model can also lay in synced dialogue and sound as it renders.
About Vidu
Vidu is an AI video generation platform developed by ShengShu Technology, a Beijing-based startup founded in March 2023 by researchers from Tsinghua University. It runs entirely in the browser at vidu.com, with no local GPU required, and offers three ways to create a clip: from a text prompt, from a still image, or from a set of reference images.
ShengShu Technology's April 2026 Series B raised $293M (¥2B), led by Alibaba Cloud with TAL Education Group, Baidu Ventures, and other investors participating: the largest single round in China's video-generation sector to date. Total funding now tops $987M, and the company counts 3,000+ enterprise customers with ARR past $20M as of 2025.
On 15 September 2026, ShengShu launched Vidu S2: a real-time interactive digital-character model (S2-Avatar, output resolution raised from 540p to 720p) that holds a live conversation over text or voice and can take in a new reference image mid-interaction, plus S2-Editing, which applies style, outfit, subject and background changes to an incoming video stream in real time. It puts Vidu head-to-head with talking-head specialists like Synthesia, HeyGen and D-ID, though those target scripted corporate video rather than live interaction. On straight clip generation Vidu still competes most directly with Kling AI, Runway, Pika and Luma's Dream Machine in the wider media generation category.
Vidu rarely sits alone in a creator's pipeline. Because Reference-to-Video and Image-to-Video both start from a still, creators commonly generate the source image first in Midjourney, Adobe Firefly or Freepik, then feed it into Vidu for motion. Vidu's own audio layer covers dialogue and sound effects but not music, so a soundtrack typically comes from a separate generator such as Suno, and voice work outside Vidu's built-in track often runs through ElevenLabs. The rendered clip then gets trimmed for social formats in CapCut or captioned in Descript before publishing. Alibaba Cloud, which led ShengShu's April 2026 Series B, distributes Vidu's models through its own Model Studio for enterprise customers already on that stack.
Screenshots


Pricing
Free tier with bonus starter credits and unlimited off-peak generation. The $8/$28/$79 monthly prices for Standard, Premium and Ultimate are the annual-billing rate (20% off); paying month-to-month runs $10, $35 and $99 for the same 800/4,000/8,000 credits (about 200/1,000/2,000 videos). Standard clears the watermark and licenses output for commercial use; Premium adds priority generation; Ultimate raises the throughput ceiling.
035/sec off-peak; enterprise volume pricing is quote-based rather than a fixed monthly fee. Kling AI and Pika sell console-style monthly plans in the same $8-$95 band.
| Tier | Monthly price | What it includes |
|---|---|---|
| Free | Free | |
| API Pay-As-You-Go | Free | |
| Standard | $8/mo | |
| Premium | $28/mo | |
| Ultimate | $79/mo |
Feature Comparison by Tier
| Feature | Free | Premium | Standard | Ultimate |
|---|---|---|---|---|
| Credits per month | Starter bonus | 4,000 | 800 | 8,000 |
| Approx. videos per month | — | ~1,000 | ~200 | ~2,000 |
| Watermark-free output | — | ✓ | ✓ | ✓ |
| Commercial use license | — | ✓ | ✓ | ✓ |
| Priority generation | — | ✓ | — | ✓ (max) |
Key Features
- Native Audio-Video Co-Generation: Vidu Q3 model generates synchronized audio and video together in a single pass up to 16 seconds: including dialogue with lip-sync, ambient sound, and music.
- Real-Time Interactive Avatars and Editing (Vidu S2): Launched 15 September 2026: S2-Avatar holds a live text or voice conversation at up to 720p and can take a new reference image mid-interaction, while S2-Editing restyles a live camera feed, swaps outfits or the background, or replaces the subject as the stream plays.
- Multi-Subject Consistency: Reference-to-Video mode blends up to 7 reference images into a single consistent video: maintaining character appearances, object shapes, and scene elements across clips.
- Three Creation Modes: Text-to-Video from natural language prompts; Image-to-Video with first-frame and last-frame control; Reference-to-Video blending up to 7 reference images into one consistent video.
- My References Library: Save characters, props, and scenes for reuse across generations: enabling series production and brand consistency without re-uploading reference images.
- All-in-One Creative Suite: AI Sound Effects Generator (48 kHz), AI Image Generator, and ViduClaw marketing agent: plan and produce complete video campaigns from storyboard to final output in one platform.
Pros
- Reference-to-Video's 7-image consistency is hard to match: most rival tools only hold one or two reference images before character continuity breaks down across a scene.
- The unlimited off-peak generation lets you judge real output quality at no cost: many competitors cap free use at a handful of low-res previews.
- A typical render finishes in well under a minute, fast enough to try a shot several times before committing paid credits to a final version.
- Native audio-video co-generation means a finished clip already carries synced dialogue and background audio, skipping the separate voiceover-and-sync pass most rival tools still require.
Cons
- Failed renders still consume credits: users report effective output far below advertised credit allowances due to problematic generations.
- Poor customer support and strict no-refund policy: Trustpilot 2.0/5 with widespread complaints about unresponsive ticket handling.
- Struggles with complex multi-element prompts: reviewers explicitly recommend simple, focused prompts; narrative or physically complex scenes often fail.
- Anime/cinematic aesthetic only: unsuitable for UGC-style ads or realistic product demos without heavy post-processing.
Data Handling
- Training-data policy
- Vidu's privacy policy caps operational data retention at six months (longer where law requires it, e.g. tax/fraud records) and does not explicitly state whether prompts or outputs are used for model training; it names Gemini, GLM and Qwen as third-party AI services data may be shared with to fulfill the service. No enterprise data-isolation tier is documented.
- Data retention
- 180 days
- Compliance
- GDPR
Frequently Asked Questions
What does Vidu actually cost?
Vidu's subscription plans run $10, $35 and $99 a month for Standard (800 credits), Premium (4,000 credits) and Ultimate (8,000 credits); paying annually drops those to $8, $28 and $79 a month, a 20% discount. Standard removes the watermark and clears output for commercial use, and Premium adds priority generation. Outside the subscription, API credits cost $0.005 each, and Q3-turbo video generation runs $0.065 a second at 1080p in standard hours or $0.035 off-peak. Enterprise volume pricing is quote-based through Vidu's platform team.
Can you use Vidu without paying?
Yes. Vidu's free tier includes bonus starter credits plus unlimited generation during off-peak hours, so you can test all three creation modes at no cost. Because the paid tiers specifically advertise removing the watermark and unlocking commercial use, free-tier output ships watermarked and is not licensed for commercial work.
What are Vidu's closest competitors?
Kling AI, Runway, and Pika are the closest matches on hokai. Kling AI is the pick if you want a platform already serving 60 million-plus creators; Runway leans into Gen-4.5's physics-aware rendering and a full agent-driven editing suite; Pika fits best if MCP support for agentic workflows matters more than reference consistency.
How does Vidu compare to Kling AI in 2026?
Kling AI is the more direct rival: both platforms generate audio with the video in one pass and offer a real free tier. Kling AI pulls ahead on raw resolution, native 4K against Vidu's 1080p ceiling, and adds lip-synced dialogue across five languages. Vidu answers back with Reference-to-Video: blending several images to hold one character, object, or scene consistent across a clip, a feature Kling AI's spec sheet does not list.
How do you set up Vidu?
Create a free account at vidu.com; there is nothing to install since Vidu runs entirely in the browser. Pick a mode: type a prompt for Text-to-Video, upload a still for Image-to-Video, or add multiple reference images to My References for Reference-to-Video. A typical clip renders in as little as 10 seconds, and the free tier's starter credits cover several test generations before you need to pay.
Top Alternatives
- Kling AI: Pick Vidu for reference-to-video consistency across up to 7 images; pick Kling AI for native 4K output and lip-synced audio in five languages.
- Runway: Pick Vidu for built-in reference-to-video consistency and native audio-video generation; pick Runway for Gen-4.5's physics-aware rendering and its full agent-driven editing suite.
- Pika: Pick Vidu for multi-image reference consistency and native synced audio; pick Pika for Pikaframes aspect control and MCP integration into agentic pipelines.
HokAI guides covering Vidu
- The Best Free AI for Turning a Photo Into a Video in 2026: Kling, CapCut, HeyGen and Fliki compared for free photo-to-video, checked 22 Sept 2026. See which free tiers actually renew, and which one just quietly stopped.
- Best AI Video Generators in 2026: Pick the Category Before You Pick the Tool: Kling AI generates a second of 1080p video for about 13 cents, cheaper than Runway's 23 cents. See how 31 AI video tools split into three buying decisions.