FLUX 3

One multimodal model for video, image, and native-audio generation.

FLUX 3 launched July 23, 2026 as the first model from Black Forest Labs, the 2024-founded lab behind FLUX.1 and FLUX.2, to jointly generate video, image, and audio from one Self-Flow architecture. It suits teams needing synchronized dialogue baked into the video itself, rather than dubbed on afterward.

FLUX 3 is Black Forest Labs' unified multimodal model, built on the Self-Flow architecture, generating up to 20 seconds of HD or Full HD video with native multilingual audio, dialogue, and sound effects from text, images, or existing clips, alongside a robotics leg called FLUX-mimic.

Provider: Black Forest Labs · Family: FLUX 3

More about Black Forest Labs on HokAI

Input modalities: text, image, video · Output: video, audio

About FLUX 3

FLUX 3 is Black Forest Labs' first multimodal foundation model, announced July 23, 2026 and built on a new architecture the company calls Self-Flow, which jointly trains on images, video, and audio inside one shared system rather than bolting separate generators together. Black Forest Labs is the German AI lab founded in 2024 by former Stability AI researchers behind the original Stable Diffusion and FLUX.1 image line, and FLUX 3 is their move from still images into a single model that also produces motion and sound. The design goal, per the company's own launch material, is a model that learns how objects hold together, how things move, and how events sound, instead of learning each modality in isolation. Only the video generation leg of FLUX 3 is generally available so far, launched August 4, 2026. It generates continuous clips lasting up to twenty seconds in a single pass, at HD (up to 1 megapixel per frame) or Full HD (up to 2 megapixels per frame), from text, a still image, first-and-last frames, keyframes, or an existing clip. It supports text-to-video, image-to-video, video-to-video, video continuation, and controlled scene transitions using keyframes, plus optional native audio: multilingual dialogue with lip-sync, sound effects, and environmental ambience generated alongside the visuals rather than dubbed on afterward. FLUX 3 Image, the still-image generation leg, was announced alongside video but has not shipped a public release date as of this writing. An open-weight FLUX 3 Dev variant is confirmed on the roadmap for later in 2026 but has no license terms, parameter count, or Hugging Face repository published yet. No independent third-party benchmark suite exists for FLUX 3 the way SWE-bench or MMLU exists for language models. Black Forest Labs published its own pre-release preference-test win rates against two named rivals ahead of launch, but those figures come from an early checkpoint and were not independently reproduced, so no benchmark figures are recorded here until a third party publishes one. Pricing for the shipped video leg runs per second of output rather than per token, tiered across Draft, Standard HD, and Standard Full HD modes, each with a lower rate for text-or-image-to-video than for video-to-video. There is no subscription or seat fee, only metered usage against a prepaid balance; exact per-second rates are listed in the pricing table below. The still-image and robotics legs of the product line carry no announced rate card yet. Access today is the direct Black Forest Labs API (docs.bfl.ai) plus third-party inference platforms including fal.ai, which mirrors the Draft, Standard HD, and Standard FHD tiers at the same rates. No AWS Bedrock, Google Vertex, or Azure listing exists yet for FLUX 3; the company's prior FLUX.1/FLUX.2 image lines are more broadly distributed across those clouds. For the open-weight FLUX 3 Dev variant, expect the same non-commercial open-weights license structure BFL used for FLUX.1-dev and FLUX.2-dev rather than a permissive open-source license, based on the licensing pattern the company has followed for every prior Dev release. The same Self-Flow backbone also powers FLUX-mimic, a robotics Video-Action Model built with mimic robotics that predicts physical robot actions from the same joint representation, currently in early access and already being tested by Audi on production-line manipulation tasks; BFL states it can be fine-tuned on as little as thirty minutes of robot demonstration data and deployed on a single on-prem GPU. This is the clearest signal of what Self-Flow is actually for: a shared world model that outputs pixels, sound, or motor commands depending on the head attached to it. No public system card, red-team partner disclosure, or safety benchmark has been released for FLUX 3. Black Forest Labs has not published a HarmBench score, jailbreak-resistance rating, or content-restriction policy specific to FLUX 3, unlike frontier LLM vendors who publish these alongside launch. Buyers evaluating FLUX 3 for production should treat the safety posture as unverified rather than assume parity with the company's FLUX.1/FLUX.2 image-model documentation. FLUX 3 is best suited today for teams building short-form video with synchronized dialogue or sound baked into the generation itself, rather than as a separate TTS or foley pass, and for teams that specifically need the video-continuation and keyframe-transition features. It is not yet a fit for teams needing still-image generation from this exact model line (FLUX.1/FLUX.2 remain the shipped image products), for teams needing a documented safety posture for regulated deployments, or for teams needing self-hosted/open-weight access before the FLUX 3 Dev release lands. Competing shipped video models with longer track records and independent evaluation, such as Google's Veo 3.1 and OpenAI's Sora 2, are worth comparing directly since FLUX 3's own preference-test wins have not yet been reproduced by a third party.

Pricing

Video generation is billed per second, not per token. Draft Mode: $0.06/s (text/image-to-video), $0.12/s (video-to-video). Standard HD: $0.17/s (text/image-to-video), $0.41/s (video-to-video). Standard Full HD: $0.29/s (text/image-to-video), $0.53/s (video-to-video). Pay-as-you-go, no subscription. FLUX 3 Image and FLUX-mimic pricing not yet published.

Key Features

  • Twenty-second single-pass video generation: Generates continuous clips up to 20 seconds long in one generation at HD (1MP/frame) or Full HD (2MP/frame), longer than most shipped competitors' single-generation cap.
  • Native multilingual audio: Generates lip-synced dialogue in multiple languages, sound effects, and environmental ambience inside the same pass as the video, not as a separate dubbing step.
  • Keyframe-controlled transitions: Accepts first-and-last frames or intermediate keyframes to control scene changes and camera moves within one continuous generation.
  • Video-to-video and continuation: Extends or transforms an existing clip, not just text/image-to-video, letting a workflow build on prior generations.
  • Self-Flow shared architecture: The same backbone that generates pixels and audio also powers FLUX-mimic, a robotics Video-Action Model already piloted by Audi.

Pros

  • Longer single-generation video length than most shipped competitors' per-generation cap.
  • Native audio (dialogue, effects, ambience) is generated with the video, removing a separate TTS/dubbing pass from the pipeline.
  • Pay-as-you-go per-second pricing with a cheap Draft tier for fast iteration before rendering a final Standard HD/FHD pass.

Cons

  • FLUX 3 Image has no public release date, so the model line cannot yet replace FLUX.1/FLUX.2 for still images.
  • No independent third-party benchmark exists; the only quality comparisons are vendor-reported from an early checkpoint.
  • No public system card or safety benchmark, unlike the documentation Black Forest Labs ships for its image models.

Frequently Asked Questions

How much does FLUX 3 cost?

FLUX 3's video generation charges per second of output, not per token, across three tiers: Draft ($0.06/s or $0.12/s for video-to-video), Standard HD ($0.17/s or $0.41/s), and Standard Full HD ($0.29/s or $0.53/s). It is pay-as-you-go through the Black Forest Labs API or fal.ai, with no subscription fee. FLUX 3 Image and FLUX-mimic pricing have not been published yet.

How does FLUX 3 compare to Veo 3.1, Sora 2, or Kling?

Independent, third-party benchmarks for FLUX 3 have not been published yet. Black Forest Labs' internal preference tests claimed wins over Kling v3 Pro and Runway Gen-4.5, but those results come from an early checkpoint and have not been independently reproduced, so treat them as a vendor claim rather than a settled comparison against Veo 3.1 or Sora 2.

Is FLUX 3 open source or open weights?

No. FLUX 3 is proprietary and API-only today. Black Forest Labs has confirmed an open-weight FLUX 3 Dev variant is planned for later in 2026, following the non-commercial open-weights license pattern used for FLUX.1-dev and FLUX.2-dev, but no license terms, parameter count, or release date have been published yet.

Does FLUX 3 train on user data?

Black Forest Labs has not publicly disclosed a data retention or training-on-inputs policy specific to FLUX 3. Teams with strict data-handling requirements should confirm the current policy directly with Black Forest Labs before sending production or sensitive prompts.

Who is FLUX 3 best for and who should avoid it?

FLUX 3 fits teams making short-form video that needs synchronized dialogue or sound generated with the clip itself, and robotics teams evaluating FLUX-mimic for manipulation tasks. It is not yet the right choice for still-image generation (use FLUX.1 or FLUX.2 instead), regulated deployments needing a published safety posture, or anyone needing self-hosted access before FLUX 3 Dev ships.

More AI Models on HokAI

Visit FLUX 3 Official Page