ByteDance Seed shipped SeedRealtime on August 5, 2026 as a new entry in the GenMedia family that also includes Seedance 2.5 and Seed Audio 1.0. It runs free inside the Doubao consumer app, with no developer API, benchmark score, or parameter count published yet.
SeedRealtime, launched by ByteDance's Seed team, is a full-duplex model that fuses audio, video, and text understanding in one system instead of chaining separate speech-recognition, vision, and text-to-speech stages. ByteDance's internal human evaluation found it cuts conversational pacing errors roughly in half versus cascaded pipelines, and it can speak up unprompted when it notices a relevant visual change.
Provider: ByteDance · Family: SeedRealtime
Input modalities: text, audio, video · Output: text, audio, tool-calls
About SeedRealtime
SeedRealtime is ByteDance Seed's native audio-visual full-duplex LLM, built by the Beijing-based Seed research org and fully rolled out on August 5, 2026 inside its GenMedia lineup, alongside Seedance 2.5 and Seed Audio 1.0. Rather than chaining separate speech-recognition, vision-language, and text-to-speech modules the way cascaded assistants do, it fuses perception, understanding, decision-making, and expression inside one end-to-end model. ByteDance has not released SWE-bench, GPQA, MMLU, or any independent benchmark score for SeedRealtime, and no outside lab appears to have evaluated it. The only figure ByteDance has shared is from its own human evaluation: testers found it cuts pacing problems, like overlapping speech and slow replies after a pause, roughly in half against the prior cascaded pipeline, a self-reported UX result rather than a reasoning or coding score. ByteDance has not disclosed a context window, session length, or output token limit, and there is no technical report or model card to derive one from. The product centers on live, continuous camera-and-microphone sessions in the Doubao app rather than long-document processing, which likely explains the missing figure. SeedRealtime takes audio, video, and text as input and produces spoken audio and text output, with tool calls folded into its responses. It resolves ambiguous or homophone speech by checking what is currently in view, such as reading a menu to interpret a spoken order, and acts proactively: touring a museum, it volunteered a reminder the instant a requested exhibit panned back into frame, unasked. It manages its own turn-taking by tracking scene, speaker, pauses, and background chatter rather than an external voice-activity trigger, and was also demoed catching a live espresso-extraction mistake and filtering out unrelated airport chatter. SeedRealtime carries no published per-token, subscription, or API price, since ByteDance has not opened a standalone way to call it. It ships free inside the Doubao app for any iOS or Android user, unlike Seedance and Seed Audio, which both carry published per-generation API pricing. The Doubao app is the only access path. ByteDance has explicitly not shipped a Volcano Engine or BytePlus endpoint for SeedRealtime, unlike the rest of its GenMedia lineup. There is no SDK, self-hosting option, downloadable weights, or quantized variant; a user opens a combined camera-and-microphone view with one tap, and the session runs from there. ByteDance has not published a system card, red-team partner list, or refusal-rate benchmark for SeedRealtime. The Seed team's general privacy policy says it collects only content a user uploads, plus attached file metadata, without a SeedRealtime-specific retention window. As a feature inside a mainstream Chinese consumer app, it runs under Doubao's existing moderation and legal-compliance rules. SeedRealtime fits Doubao users who want an assistant that reacts to what it sees and hears without a repeated prompt, such as monitoring a cooking step or getting a nudge while touring a museum. It is not a fit for developers needing an API, a benchmarked model, or an enterprise deployment, since there is nothing here to integrate, self-host, or place on a leaderboard; look at a published realtime voice API instead. No training data composition, licensing detail, or cutoff date has been published. SeedRealtime sits inside ByteDance's Beijing-based Seed org and inherits Doubao's existing terms of service rather than a model-specific license. The August 5 release is SeedRealtime's first version, with nothing prior to compare it against. It landed about six weeks after ByteDance's Volcano Engine FORCE conference, where the Seed team shipped several other GenMedia models, extending that cadence into live interaction. The lineup also lists Seeduplex, a separate speech-only full-duplex model, so full-duplex interaction now spans parallel research tracks.
Pricing
SeedRealtime has no published per-token or subscription price; ByteDance has not opened a paid or developer tier for it. It is free to any Doubao app user on iOS or Android, with no credit limit or waitlist disclosed.
Key Features
- Unified Audio-Visual-Text Model: Runs perception, understanding, decision-making, and speech generation inside a single system, skipping the separate speech-recognition, vision, and text-to-speech modules most voice assistants chain together.
- Self-Directed Turn-Taking: Decides when to speak by tracking scene, speaker, pauses, and background noise itself, rather than relying on an external voice-activity-detection trigger.
- Unprompted Proactive Alerts: Speaks up on its own when it notices a relevant visual change, such as flagging a tracked museum exhibit the moment it re-enters the camera's view.
- Joint Audio-Visual Disambiguation: Cross-references what it currently sees to resolve ambiguous or homophone speech, for example reading a menu to correctly interpret a spoken order.
- One-Tap Live Sessions in Doubao: Starts a combined camera-and-microphone session inside the free Doubao app on iOS and Android with a single tap and no setup.
Pros
- Runs audio, video, and text through one model instead of a cascaded pipeline, which ByteDance's own testers rated as roughly halving conversational pacing problems.
- Speaks up unprompted on relevant visual changes instead of waiting to be addressed, demoed catching an espresso mistake mid-pour and flagging a museum piece as it came into frame.
- Live for any Doubao user today with no waitlist, signup form, or developer key required.
Cons
- No developer API: ByteDance has not opened a Volcano Engine or BytePlus endpoint, so it cannot be built into a third-party product yet.
- No parameter count, technical report, or independent benchmark score has been published, which makes it hard to size up against other frontier models.
- Locked to the Doubao consumer app, with no self-hosted, open-weights, or enterprise deployment path.
Frequently Asked Questions
How much does SeedRealtime cost to use?
SeedRealtime costs nothing to use today: ByteDance has not released a developer API, so anyone simply opens the Doubao app on iOS or Android and starts a live session. There is no credit system, subscription tier, or waitlist, and ByteDance has not said whether that changes once an API ships.
How does SeedRealtime compare on benchmarks to models like GPT-5 or Gemini 2.5?
ByteDance has not put out an SWE-bench, GPQA, or MMLU score for SeedRealtime, so there is no direct leaderboard comparison to reasoning-focused models yet. The only figure ByteDance has shared is an internal human-evaluation result showing roughly half as many conversational pacing errors as its earlier cascaded pipeline, a self-reported UX metric rather than an independent benchmark.
Is SeedRealtime open source or proprietary?
SeedRealtime is closed and proprietary. ByteDance has released no weights, no parameter count, and no research paper, and the only way to use it is through the free Doubao app rather than a download or self-hosted deployment.
Does SeedRealtime train on the camera and microphone data it sees?
ByteDance has not published a training or retention policy specific to SeedRealtime. Its general Seed privacy policy states it collects the content a user uploads, such as text, image, audio, and video, along with associated file metadata, but does not name a dedicated opt-out for live camera and microphone sessions.
Who should use SeedRealtime, and who should look elsewhere?
It suits everyday Doubao users who like a voice-and-camera assistant that reacts to their surroundings in real time, such as flagging a museum exhibit or catching a cooking mistake unprompted. Anyone who needs a callable API, an independently benchmarked model, or a self-hosted option should pick Seed 2.1 for text and agent work instead, since SeedRealtime offers none of those today.