FLUX 3 Video Explained: 20-Second AI Clips With Audio
FLUX 3 Video is Black Forest Labs' video model, generally available since August 4, 2026. Here are the specs, the per-second pricing, the benchmark claims, and what 20 second clips with native audio change for short-form marketing.

Key takeaways
- Black Forest Labs made FLUX 3 Video generally available on August 4, 2026, through the BFL API, dashboard, and select partners.
- FLUX 3 Video generates clips of up to 20 seconds in HD or Full HD at 24 frames per second, with native audio included at no extra cost.
- Pricing is $0.06 per second in draft mode, $0.17 per second for a full HD render, and $0.29 per second at Full HD, so a 20 second Full HD clip costs $5.80.
- Video continuation is the expensive mode at $0.43 per second in HD and $0.54 per second at Full HD.
- Black Forest Labs reports internal Elo scores of 1,135 on text-to-video and 1,051 on image-to-video, ahead of Gemini Omni Flash, MiniMax H3, and Seedance 2.0, with no independent arena results published at launch.
FLUX 3 Video is Black Forest Labs' text-to-video, image-to-video, and video continuation model, and it went generally available on August 4, 2026 through the BFL API, the BFL dashboard, and a handful of partner platforms. It generates clips of up to 20 seconds in HD or Full HD at 24 frames per second, and it generates the sound at the same time as the picture: dialogue with lip sync in more than a dozen languages, sound effects, and room ambience, all included in the per-second price. Rendering starts at $0.06 per second for a draft and runs to $0.29 per second for a full Full HD render. The combination is what makes this a release worth reading about, not any single spec: 20 second length, audio that comes out of the same pass as the video, and a draft tier cheap enough to explore with.
What is FLUX 3 Video?
FLUX 3 Video is a generative video model that turns text prompts, still images, or existing footage into short clips with synchronized audio. It is one part of FLUX 3, which Black Forest Labs describes as a single multimodal model spanning image, video, audio, and action prediction. The video half shipped first, on a pay-as-you-go basis through the BFL dashboard and API. Black Forest Labs has said that open-weight access is part of the wider FLUX 3 rollout, but at launch FLUX 3 Video is a hosted product you call, not a checkpoint you run.
The company is best known for the FLUX image models, which became the default open-weight choice for image generation across a large slice of the AI tooling ecosystem. FLUX 3 Video is its move from stills into motion, and it arrived into a market that got crowded very fast. MiniMax H3 landed on July 31, 2026, Seedance 2.5 followed on August 8, and Google Gemini Omni Flash had been holding the top of the public video leaderboards through the summer. Four serious releases inside two weeks is the actual story of this quarter.
What are the specs of FLUX 3 Video?
- Clip length of 5 to 20 seconds for text-to-video and image-to-video, and 5 to 15 seconds for continuation built from up to 4 seconds of input footage.
- HD output at up to 1 megapixel per frame, or FHD at up to 2 megapixels per frame.
- A fixed 24 frames per second, across seven aspect ratios that include 9:16 for vertical short-form.
- Native audio generated alongside the frames: lip-synced multilingual dialogue, sound effects, and environmental ambience, at no extra charge.
- Three working modes: text-to-video, image-to-video that accepts 1 to 10 keyframe images which can be pinned to timestamps, and video continuation that extends existing footage while holding camera movement and dialogue.
- A draft tier that renders at HD for roughly a third of the full-render price.
The keyframe handling deserves a second look, because it is the part that turns a slot machine into a tool. Pinning images to timestamps inside a clip means you can specify where a shot starts, what it passes through in the middle, and where it lands, rather than writing a paragraph of prompt and hoping. For product video, that is the difference between a model that occasionally shows your packaging and a model you can direct.
How much does FLUX 3 Video cost per clip?
Black Forest Labs prices FLUX 3 Video by the second, and the rate depends on mode, resolution, and whether you are drafting or rendering final. Draft mode costs $0.06 per second for every input type. A full-quality render costs $0.17 per second in HD and $0.29 per second at Full HD for text-to-video and image-to-video. Video continuation is the premium mode at $0.43 per second in HD and $0.54 per second at Full HD, which reflects the extra work of matching existing footage.
Turning that into clip prices: an 8 second HD hook costs $1.36 at full quality and $0.48 as a draft. A full 20 second Full HD clip costs $5.80. Extending an existing 15 second clip at Full HD costs $8.10. The spread between the cheapest and most expensive second is close to nine times, so the mode you pick matters more to your monthly bill than the number of clips you make.
What does draft mode change about the way you work?
Draft mode is the quiet feature that changes the economics. Because a draft costs about a third of a full render and comes back faster, the sane workflow is to generate many cheap variants, look at them, and only pay full price for the one shot you actually intend to publish. Ten drafts of an 8 second concept cost $4.80 total, which is less than one 20 second Full HD final. Explore wide, commit narrow.
That maps neatly onto how short-form actually works. The variable that decides whether a video performs is almost never render quality. It is the hook, the framing, and the first two seconds, and the only reliable way to find the right one is to make a lot of them and let the feed vote. A cheap draft tier is a hook-testing budget in disguise.
Why does native audio matter more than resolution?
Sound is where AI video usually gives itself away. When dialogue is dubbed onto finished footage, lips drift out of sync, room tone does not match the room, and sound effects sit on top of the picture rather than inside it. Generating audio in the same pass as the frames means the model is composing one result, so a hand hitting a counter makes the right noise at the right moment, and a spoken line is shaped by the same take that produced the mouth saying it.
Resolution, by contrast, is the least interesting number on the spec sheet for short-form. TikTok, Reels, and Shorts all recompress hard on upload. A Full HD master mostly buys you headroom for cropping and for on-screen text that survives the platform encoder. If you are choosing between spending your budget on Full HD or on more variants, more variants wins almost every time.
How does FLUX 3 Video compare to Seedance and Gemini?
Black Forest Labs published internal Elo scores of 1,135 on text-to-video and 1,051 on image-to-video, placing FLUX 3 Video ahead of Gemini Omni Flash, MiniMax H3, and Seedance 2.0. Read those with the usual caution. They are vendor-run preference tests, the image-to-video margin is narrow enough that coverage described it as roughly level with Seedance 2.0, and no independent arena scores had been published at the time of the launch. Vendor benchmarks tell you what a model is proud of, not how it behaves on your brief.
The honest way to evaluate a video model is to run your own product, your own script, and your own hook through it, then look at retention on the posts. Everything else is a leaderboard.
Where FLUX 3 Video differentiates most clearly is length. Twenty seconds in a single generation is longer than most of its direct competition, which means a hook and a payoff can live in one continuous shot instead of being stitched from two generations with a visible seam in the middle. For demo and explainer formats, that is worth more than a benchmark position.
Does a better video model mean better marketing results?
Not on its own, and this is the trap every model launch sets. A better model raises the ceiling on how good a single clip can look. It does nothing about the parts of the job that actually decide outcomes: how many clips you post, how consistently you post them, whether the accounts you post from have any standing with the algorithm, and whether you can tell which post produced a signup. Plenty of brands now sit on a folder of beautiful renders and post twice a month.
The platforms have also moved against volume without substance. TikTok and Instagram both down-rank content that reads as mass-produced, watermarked, or reposted, and reward completion rate and rewatches. That rewards the brand making many genuinely different, well-hooked posts, not the one making the same post at higher resolution.
How do you turn AI video into an actual posting engine?
This is the gap Fastlane fills. You give it your website URL, it learns your product and positioning, and it generates the content: hyper-realistic AI UGC videos from a library of more than 1,000 AI characters, slideshows, hook plus demo videos, and memes, remixed against live trends rather than against whatever was popular last quarter. Blitz mode lets you swipe through what it made and approve in minutes rather than sitting through a review queue.
From there it publishes natively to TikTok, Instagram Reels, and YouTube Shorts, schedules weeks ahead, and reports unified analytics with signups and sales attributed back to the individual post that produced them. If your accounts are the bottleneck rather than your content, Fastlane also sells human-warmed TikTok and Instagram accounts, created from scratch by real people in up to 12 countries and warmed for 5 days in your niche, from $80 per account per month at the launch offer plus $1.50 per post upload. There is a developer API and an MCP server if you would rather drive all of it from your own stack.
FLUX 3 Video, Seedance 2.5, and MiniMax H3 are all good news for anyone making short-form, and the per-second prices keep falling. But the model was never the hard part. Distribution is. Start free at usefastlane.ai, with no credit card, and paid plans from $29 per month when you want the volume.
Frequently asked questions
What is FLUX 3 Video?
FLUX 3 Video is Black Forest Labs' generative video model, which turns text prompts, still images, or existing footage into clips of up to 20 seconds with synchronized native audio.
When was FLUX 3 Video released?
Black Forest Labs opened general availability on August 4, 2026, through the BFL API and dashboard on a pay-as-you-go basis, after an earlier limited-access period.
How long can a FLUX 3 Video clip be?
Text-to-video and image-to-video run 5 to 20 seconds in a single generation, and video continuation runs 5 to 15 seconds from up to 4 seconds of input footage.
How much does FLUX 3 Video cost?
Draft renders cost $0.06 per second, full quality HD costs $0.17 per second, Full HD costs $0.29 per second, and video continuation costs $0.43 per second in HD or $0.54 per second at Full HD.
Does FLUX 3 Video generate audio?
Yes. Audio is generated with the frames rather than dubbed on afterwards, covering lip-synced multilingual dialogue, sound effects, and environmental ambience, and it is included in the per-second price.
Is FLUX 3 Video open source?
Not at launch. FLUX 3 Video is a hosted API product, and Black Forest Labs has said open-weight access for parts of the FLUX 3 rollout is coming without committing to a date.
Is FLUX 3 Video better than Seedance 2.0 or Gemini Omni Flash?
Black Forest Labs' own preference testing puts it ahead on text-to-video and roughly level with Seedance 2.0 on image-to-video, but those are vendor-run numbers and no independent arena scores had been published at launch.
How do you post AI video like this to TikTok and Instagram at scale?
Fastlane turns a website URL into AI UGC videos, slideshows, hook and demo clips, and memes, then publishes them natively to TikTok, Instagram Reels, and YouTube Shorts on a schedule.
