- Frivolous use of the term World Model.
- Claims 20 seconds of video, shows only jumpcuts.
Coming soon!
> Over the next few weeks and months, we will make the following capabilities available
> Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
> We will also release more technical details on the underlying approach.
But then again I heard the downers have always been the first to leave their dung comments here so let's see...
I'm confused, videos contain images and audio ...?
Isn’t video + audio all you need?
I honestly hope they put an unrealistic amount of wilhelm scream into the learning process, just for fun.
- Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
- Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
- Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
- Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)> our mission to develop real-world visual intelligence
Visual is mono-modal, isn't it?
Of course there will be feuds from robots of different family groups but they will be minimal as it quickly becomes symmetrical robot conflict with high casualties as they learn too fast from each other, it's likely those will be avoided, it will be after all much easier to confront humans for any given resources.
Truly a pinnacle for technology, albeit perhaps not for mankind.
It does not seem as grating as the slop I typically see in README.md files or generated docs.
Hiring in university towns is a pretty standard practice for startups outside SF.
Maybe they presume that after a series of "good enough to some" they may be getting near the Real Thing?
The term "world model" as it was once used in model-based RL can now apparently refer to anything as silly as linear regression. Then again, the RL folks probably borrowed the term from behavioral scientists before them. It's probably best to simply accept this :/
A similar thing happened to "object oriented" which has been misused by philosophers and visual artists alike.
I see nothing here but them TELLING us how great it is. Not showing us.
At the same time, I can see social media is being flooded with absolutely horrid AI slop images and video. I'm much more pessimistic about the practical beneficial real world use of totally artificial image and video generators. It seems the uses that these things are put to when they get into the hands of millions of people are detrimental to society and not a benefit.
FLUX 3 is now available in Early Access.
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.
No single modality provides a complete description. Each is a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process. Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.
Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.
FLUX 3 is our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments. Early results in content creation and physical AI suggest it is the right path.

FLUX 3 builds on Self-Flow, our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, we significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time.

Self-Flow vs. Flow Matching (FM). Left: generation error (Fréchet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better).
As a result, FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video. We are highlighting a few of the model’s key capabilities below.
FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation.
Its core capabilities include the following (all outputs come with native audio generation):
For the preliminary analysis below, we generated 10-second text-to-video clips in 720p with audio.

Evaluations are early and we expect further improvements
As the model and the harness around it are still in development, these results are preliminary, and we expect further improvements during the early access phase. Across early evaluations, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons.
While still in development, FLUX 3 Video is already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities. Furthermore, these capabilities can be combined to create sequences lasting several minutes, where visual references help ensure that the characters remain consistent across all scenes.
FLUX 3 Video is now available in Early Access here
FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. In preliminary evaluations conducted during midtraining, FLUX 3 already shows a significant improvement over earlier versions of FLUX: its ability to handle complex prompts and text generation has improved significantly. The model produces a wide range of output styles (see the following samples), and is able to render high-accuracy text in multiple languages.

As with video evaluations, these are preliminary results, and we expect further improvements before release. We will open up an early access phase for FLUX 3 Image in the following weeks.
FLUX 3's world understanding extends to action prediction. We have taken two routes to it: integrating native action prediction into FLUX 3 directly, scaling up our initial work in Self-Flow; and using the pretrained video backbone as a dynamics-aware foundation that specialized action models can be finetuned from with limited task-specific data.
For the second, mimic robotics was one of the first partners to gain early access to FLUX 3. Together we developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic's expertise in robot learning for dexterous manipulation and production deployment. Read our thesis on why physical AI and content creation run on the same foundation, and how it's being tested on real production tasks at Audi.
Over the next few weeks and months, we will make the following capabilities available, each after an early access phase for ensuring smooth rollout, collecting feedback and rigorous safety-testing. All capabilities are built from the same underlying multimodal flow matching model. These capabilities and models include:
We will also release more technical details on the underlying approach.
We are only beginning to scratch the surface of versatile, capable, unified multimodal models, and what they will enable. From interactive image & video editing, simulation to computer use and physical AI, the frontier is wide open. While we gradually roll out these new capabilities, we are already working on the next generation models. Our goal is to unify perceptual, action and language prediction in the same unified model.
If you are interested in exploring and building with FLUX 3, get in touch here. If you are interested in contributing to our mission, join us! We are hiring in Germany and the US.
https://arxiv.org/abs/1805.06485 is enough to make any half decent game dev sit around wondering what was going on, as well as the general "maybe we're wasting a lot of space with all these floats?" At least now LLMs can tell them what SoTA is in other fields before they try to rediscover it.
For me this is like announcing a slightly more efficient form of coal-powered steam engine. Great, but it's still running on coal. I'm excited about moving beyond coal to cleaner and more sustainable energy sources. Current "AI", based on machine learning, is just recycling existing human works: books, films, code, forum posts etc. But those works, like the coal, is going to run out. We've decided to stop making new things and just burn the coal that's already been deposited.
Have you watched the 46s video full screen on a monitor and not marvelled at the incredible 4K detail of the FPV motorcycle racing clip?
• It initially required significantly more VRAM and was much slower than alternatives released around the same time (like Z‑Image Turbo).
• The license felt overly restrictive.
Depending on the members, there certainly is. Put aside the "dismissers", those who have a habit or a hormonal reliance to cast a "meh". Those who objectively assess according to the input that the development of facts provide may bend their "apparent mood" accordingly. This may be more evident here because in brighter times we may be more inclined to post and submit about more idle intellectual beauty ("complications in ancient clocks"), and in darker times it makes sense that we are more focused on the problems.
Your post would have been just fine without that first bit.
> No nudity, lewdness, or sexually suggestive imagery. If it wouldn’t be appropriate for a workplace or younger audience, don’t post it here. NSFW tags are not an exception, stay classy.
Emphasis on the second sentence; where do these people work?!
/s
It's from epistemology. It is not there to refer to the subjective but to the objective.
> A similar thing happened to "object oriented" which has been misused by philosophers and visual artists alike
For instance?
Heh. On the other hand, we are seeing so much negativity around "AI", while we have access to something as close as possible to the "universal communicator" from trek. There are now several CotS things you can buy that basically serve this purpose. You can walk into a mall, buy one, and travel to the most remote place on Earth, speak in your own language, and almost anyone else can speak in their own language, and through technomagic the two of you can communicate. And yet there's so much pessimism...
Pessimistic opinions on the labor market, views about society and politics that border on the dystopian, as well as exaggerated concerns about datacenter environmental impacts, all seem like a poor fit for this website.
Otherwise a Gen Ai has to check this subreddit
These are the last good ol’ days you will have before it’s all over. Worrying about it won’t change the future. Just try to focus on all the good AI does.
There are urgent matters, important matters, and nice matters - all relevant to us.
[0]: https://en.wikipedia.org/wiki/Youth_word_of_the_year_(German...
Klein 9b is decent for image-to-image, but when used for pure generative purposes brings back SDXL levels of body horror (have some Gattica pianists).
If you can get past the annoying JSON structuring, Ideogram 4 is probably the best option in the open‑weights world right now for text2image purposes, and even scored higher than the original Nano-Banana.
For ref: Flux.2 scored 5, Ideogram4 scored 8, and gpt-image-2 scored 12 out of 15.
Comparison of Flux.2 [dev], Ideogram4, NB Pro, and gpt-image-2.
Anyone ever taken a CS course that includes extensive discussion of the Therac-25?
I use flux2.dev on my 5090 with a prompt upsampler (that I host on my 3080ti, it's been so long since I looked at it, its either some qwen or Mistral model). I use the upsampler to produce json structure based on this prompting guide: https://docs.bfl.ml/guides/prompting_guide_flux2
For being able to run on my 5090, it produces incredibly consistent and stable results that work for my use case so well, I find myself reaching for it over proprietary models.
I'm always interested to see objective comparisons though. It's clear from yours that flux lags. But the fact that I can use it locally over closed source sota models is very compelling (to me).
You've convinced me to checkout ideogram4 though!
My response to unironic OOO believers is that we must wage war on objects:
but we could be curious on how and why you saw misuse.
What substance is there beyond "I don't like it, and I think it is a negative for society", typically delivered with what I perceive as a self-righteous attitude?
We live in a time where artificial intelligence begins to rival humans in some areas.
We can now really automate things.
How interesting is that? There is so much to talk about.
I personally use Qwen3 8b LLM for this purpose, but I think that the KJ prompt builder node handles this out of the box in ComfyUI.
After AI: 99% of everything is crap
... but there will be >10x more of it, so we'll ultimately end up with more good stuff.
To quote the movie, that’s just your opinion, man. To me, the impact of technology on society is 100X as interesting and important than the technology itself. Tech does not exist in a vacuum. And we don’t invent it for its own sake. The major point of AI (maybe the only point of it) is what humans will do to each other with it. And like it or not, not all of those things are positive.
You will not see me leaving comments such as this, for example, on the release of a "strong in coding tasks" tool like GLM5.2.
I actively go out of my way to avoid patronizing businesses that advertise with AI slop generated images now, AI slop restaurant menus, and so forth.
Whether you want to interpret a desire for authenticity as some sort of self-righteous attitude is up to you. There's a lot of people that share my opinion, and a lot of them that don't. There is also clearly a lot of money behind pushing AI slop images everywhere. Facebook and the various 'pages' and 'groups' that are near 90% AI slop content are a fine example of that. I'm sure a great many advertising impressions and click-throughs have been served, much revenue has been earned. Great success.
Is there something in particular that you think would be interesting to discuss here?
That must be the poster child of the most boring and meaningless accusation thrown around new technologies (after the data center water scare, which was just plain bogus). We've heard this non stop since Cambridge Analytica nothingburger. It's all noise distracting from the one meaningful aspect of it: marketing to people is electoral manipulation, and since that's unquestionably allowed, the horses have left the barn long time ago, and futzing over the AI barn door control makes no sense at this point.
If anything, the problem that this is effective in the first place is the one to address - this translates to people still believing anything politicians say, despite decades of continued proof all the campaign promises are just plain bullshit, and stated beliefs are situational and not principled.
For a more immediate example, look at Russian weaponization of social media in Mali for a pro-russia, anti-everyone-else narrative. Often deployed against a population that has a much greater level of credulity of anything they see on social media, and lack of inoculation against it by multiple years of seeing artificially generates nonsense.