← All posts

An AI video model is now moving robot hands at Audi

2026-07-26 · 4 min read

On July 23, Black Forest Labs announced FLUX 3. If the name means nothing to you, the lab's earlier FLUX models are what a lot of design and marketing tools quietly run on when they generate an image. @bfl_ai introduced the new one as "One multi-modal model for Image, Video, Audio and Action-Prediction."

The video part got the headlines. The last item on that list is the one worth reading twice.

What actually shipped

FLUX 3 Video is in early access now. It generates clips up to 20 seconds with audio produced alongside the picture rather than bolted on afterward, including dialogue, sound effects, and ambient noise. It handles text to video, image to video, video to video, and keyframe transitions. Image generation opens in early access in the following weeks. The open-weight release, FLUX 3 Dev, comes last.

The lab's own preference tests, run on 10-second 720p clips, put FLUX 3 ahead of Luma Ray 3.2 in 93 percent of comparisons and Runway Gen-4.5 in 77 percent. Against ByteDance's Seedance 2.0 and Google's Gemini Omni Flash, it won 52 percent, which is a coin flip. Black Forest Labs says the results are preliminary and that no independent testing exists yet. Give vendor-run preference tests the same weight you would give a restaurant reviewing its own food.

The part that is not about video

The second post in the thread is the story. @bfl_ai wrote that "An early version of FLUX 3 is now running on robots." Working with the robotics company mimic, they built FLUX-mimic, a video-action model sitting on the same FLUX 3 backbone, and it has been tested and deployed at Audi.

The tasks are not demo-reel tasks. Kitting parts into structured trays. Inserting electronic control units into tight-fitting fixtures. Handling soft, flexible materials like seals and cables. That last category is the tell. Traditional factory automation is excellent at rigid, repeatable motion and genuinely bad at a rubber door seal that never lies the same way twice.

The speed numbers are small enough to matter. The backbone runs from camera input to world representation in under 80 milliseconds on a single NVIDIA RTX 5090, with a full system reaction time of 101 milliseconds. That is one desktop-class GPU, not a server rack.

Why a video model can drive a robot

The reasoning is simple once someone says it out loud. To generate video that looks real, a model has no choice but to learn contact, weight, momentum, and cause and effect. It has to know the glass tips before it spills, the cable sags under its own weight, the fingers close before the object lifts. Predicting the next frame and predicting the next action turn out to be close to the same problem, because a robot's actions are a low-dimensional shadow of what the camera is already watching.

Which means the years of compute poured into making prettier marketing clips produced something else as a side effect: a working intuition for physics. Now it is being pointed at hardware.

The honest read

Stay skeptical about the specifics. This is early access with one robotics partner at one automaker, on numbers the vendors published themselves. "Tested and deployed" covers a wide range of realities. And the open weights that would let anyone verify this are not out. @multimodalart of Hugging Face, who follows open-weight image models as closely as anyone, posted that FLUX.3 Dev covering image, video, audio, and robotic action prediction is coming soon, and called the underlying self-flow architecture "a cool evolution of flow matching." Coming soon is not shipped.

The direction, though, is not in doubt, because it is the same direction everything else is moving: one general model absorbing jobs that used to need separate specialized systems underneath.

What this signals if you run a business

Two things, and neither one requires you to care about robots.

  • Capabilities keep arriving sideways. Nobody set out to build a factory controller. They built a video model, and the physics came free. The useful tool for your business next quarter is probably a side effect of something being built for a different reason right now.
  • The general model keeps eating the specialized one. The pattern that already played out in writing, support, and code is now reaching physical work.

The practical takeaway is the same one we keep landing on. Do not build your operation around a specific tool, because the tool underneath you will be replaced within a year and you will have to rebuild. Build it around the process: the quoting, the scheduling, the intake, the follow-up that eats your week regardless of which model is winning benchmarks in July.

Those processes are stable. The models are not. If you want an outside read on which parts of your week are worth handing to software, New Face Design does a free process audit for businesses in the Fox Valley and beyond.

A company known for making pictures spent this week showing off a robot fitting car door seals. That is roughly the pace we are working at.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week actually goes and identify the first process worth automating. You keep the map either way. No pitch deck, no pressure.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere