I’m the Co-founder and CTO of Krea. We’re excited because we wanted to release the weights for our model and share it with the HN community for a long time.
My team and I will try to be online and try to answer any questions you may have throughout the day.
- GitHub repository: https://github.com/krea-ai/flux-krea
- Model Technical Report: https://www.krea.ai/blog/flux-krea-open-source-release
- Huggingface model card: https://huggingface.co/black-forest-labs/FLUX.1-Krea-dev
https://genai-showdown.specr.net
On another note, there seem to be some indication that Wan 2.2+ future models might end up becoming significant players in the T2I space though you'll probably need a metric ton of LoRAs to cover some of the lack of image diversity.
we prepared a blogpost about how we trained FLUX Krea if you're interested in learning more: https://www.krea.ai/blog/flux-krea-open-source-release
and
Cannot access gated repo for url https://huggingface.co/black-forest-labs/FLUX.1-Krea-dev/res.... Access to model black-forest-labs/FLUX.1-Krea-dev is restricted. You must have access to it and be authenticated to access it. Please log in.
Does this suggest that Flux Krea requires more training to achieve strong text-to-image alignment compared to Flux Dev? Or is it possible that Krea is optimized differently (e.g. for style, detail, or artistic variation rather than strict prompt adherence)?
Curious if anyone else has experienced this or has any insight into the differences between these two. Would love to hear your thoughts
[1]: https://www.reddit.com/r/StableDiffusion/comments/1mec2dw/te...
My reasoning: If the user types in "a cat reading a book" then it seems obvious that the result should look like a real cat which is actually reading a book. So it obviously shouldn't have an "AI style", but it also shouldn't produce something that looks like an illustration or painting or otherwise unrealistic. Without further context, a "cat" is a photorealistic cat, not an illustration or painting or cartoon of a cat.
In short, it seems that users who want something other than realism should be expected to mention it in the prompt. Or am I missing some other nuances here?
Imagine one of these: https://imgur.com/a/DiAOTzJ but with two spouts at the top dropping different colored balls
Its attempts: https://imgur.com/undefined https://imgur.com/a/uecXDzI
"Octopus DJ spinning the turntables at a rave."
The human like hands the DJ sprouts are interesting, and no amount of prompting seems to stop them.
Opinionated, as the paper says.
- cost per image - latency per image
Hope you guys can add it somewhere!
Does this have any application for generating realistic scenes for robotics training?
I noticed that the URL for this submission is wrong: I tried to submit the correct URL (https://www.krea.ai/blog/flux-krea-open-source-release) but, for some reason, the submission gets flagged as duplicated and then I can only find this item which has a URL to our old blog post.
In the mean time, I'll setup a server-side redirect from the old blog post to our new one, but it would be nice to fix the link and I don't think I can do it on my side.
From the article it doesn’t seem as though photorealism per se was a goal in training; was that just emergent from human preferences, or did it take some specific dataset construction mojo?
Regarding this part: > Since flux-dev-raw is a guidance distilled model, we devise a custom loss to finetune the model directly on a classifier-free guided distribution.
Could you go more into detail on the specific loss used for this and any other possible tips for finetuning this that you might have? I remember the general open source ai art community had a hard time with finetuning the original distilled flux-dev so I'm very curious about that.
From a business point of view, there are many use-cases. Here's a list in no particular order:
- You can quickly generate assets that can be used _alongside_ more traditional tools such as Adobe Photoshop, After Effects, or Maya/Blender/3ds Max. I've seen people creating diffuse maps for 3D using a mix of diffusion models and manual tweaking with Photoshop.
- Because this model is compatible with the FLUX architecture, we've also seen people personalizing the model to keep products or characters consistent across shots. This is useful in e-commerce and fashion industry. We allow easy training in our website — we labeled it Krea 1 — to do this, but the idea with this release is to encourage people with local rigs and more powerful GPUs to be able to tweak with LoRAs themselves too.
- Then I've seen fascinating use-cases such as UI/UX designers who prompt the model to create icons, illustrations, and sometimes even whole layouts that then they use as a reference (like Pinterest) to refine their designs on Figma. This reminds me of people who have a raster image and then vectorize it manually using the pen tool in Adobe Illustrator.
We also have seen big companies using it for both internal presentations and external ads across marketing teams and big agencies like Publicis.
EDIT: Then there's a more speculative use-case that I have in mind: Generating realistic pictures of food.
While many restaurants have people who either make illustrations of their menu items and others have photographers, the big tail of restaurants do not have the means/expertise to do this. The idea we have from the company perspective is to make it as easy as snapping a few pictures of all your dishes and being able to turn all your menu (in this case) into a set of professional-looking pictures that accurately represent your menu.
Check this out: https://github.com/krea-ai/flux-krea
Let me see if we can add more details on the blog post and thanks for the flag!
Though we wanted to keep this technical blogpost free from marketing fluff, but maybe we over-did it.
However, sometimes it's hard to give an exact price per image, as it depends on resolution, number of steps, whether a LoRA is being used or not, etc.
It’s simple: hackability and recruiting!
The open-source community hacking around it and playing with it PLUS talented engineers who may be interested in working with us already makes this release worth it. A single talented distributed systems engineer has a lot of impact here.
Also, the company ethos is around AI hackability/controllability, high-bar for talent, and AI for creatives - so this aligns perfectly.
The fact that Krea serves both in-house and 3rd-Party models tells you that we are not that bullish on models being a moat.
One interesting use case would be if you are focusing on a robotics task that would require perception of realistic scenes.
What stood out to me was that Flux Dev followed the text prompts more accurately, whereas Krea’s generations were more loosely aligned or "off" in terms of prompt fidelity with deformations in body type and the architecture.
Does this suggest that Flux Krea requires more training to achieve strong text-to-image alignment compared to Flux Dev? Or is it possible that Krea is optimized differently (e.g. for style, detail, or artistic variation rather than strict prompt adherence)?
Curious if anyone else has experienced this or has any insight into the differences between these two. Would love to hear your thoughts
In a nutshell, it follows the same license as BFL Flux-dev model.
.scrollbar-hide {
-ms-overflow-style: none;
scrollbar-width: none;
}Also, FWIW, this model focus was around aesthetics rather than strict prompt adherence. Not to excuse the bad samples, but to emphasize what was one of the research goals.
It’s a thorny trade-off, but an important one if one wants to get rid of what’s sometimes known as “the flux look”.
Re: Wan 2.2 I’ve also been reading of people commenting about using Wan 2.2 for base generation and Krea for the refiner pass which I thought was interesting.
<link rel="canonical" href="https://www.krea.ai/blog/new-krea">
Our software follows canonical links when it finds them.I've fixed the link above now (and rolled back the clock on the submission, to make up for lost time) but you might want to fix this for future pages.
I guess my point being: do you have any (real) experienced media production people working with you? People that have experience working in actual feature film VFX, animated commercial, and multi-million dollar budget productions?
If you really want to make your efforts a wild success, simply support traditional media production. None of the other AI image/video/audio providers seem to understand this, and it is gargantuan: if your tools plugged into traditional media production, it will be adopted immediately. Currently, they are tentatively and not adopted because they do not integrate with production tools or expectations at all.
what does " designed to be compatible with FLUX architecture" mean and why is that important?
We used two types of datasets for post-training. Supervised finetuning data and preference data used for RLHF stage. You can actually use less than < 1M samples to significantly boost the aesthetics. Quality matters A LOT. Quantity helps with generalisation and stability of the checkpoints though.
Often, the training is done in FP16 then quantized down to FP8 or FP4 for distribution.
Maybe you got a lucky roll :)
> FWIW, this model focus was around aesthetics
Agreed - whereas these tests are really focused on various GenAI image models ability to follow complicated prompts and are not as concerned with overall visual fidelity.
Regarding the "flux look" I'd be interested to see if Krea addresses both the waxy skin look AND the omnipresent shallow depth of field.
Thank you so much! I knew that HN software was advanced, but I didn’t know you guys used Canonical URLs like Google does. Smart and thanks for helping us with this slip!!!
Another note about preference optimisation and RL is that it has really high quality ceiling but needs to be very carefully tuned. It's easy to get perfect anatomy and structure if you decide to completely "collapse" the model. For instance, ChatGPT images are collapsed to have slight yellow color palette. FLUX images always have this glossy, plastic texture with overly blurry background. It's similar to reward hacking behavior you see in LLMs where they sound overly nice and chatty.
I had to make a few compromises to balance between "stable, collapsed, boring model" and "unstable, diverse, explorative" model.
i asked chat for an explanation and it said bfloat has a higher range (like fp32) but less precision.
what does that mean for image generation and why was bfloat chosen over fp?
Krea / Blog / Releasing Open Weights for FLUX.1 Krea
Sangwu Lee, Erwann Millon 31 July, 2025
Today, we're releasing an open version of Krea 1, our first image model trained in collaboration with Black Forest Labs to offer superior aesthetic control and image quality. This checkpoint is a guidance distilled model fully compatible with FLUX.1-dev allowing seamless integration with the existing ecosystem. FLUX.1-Krea [dev] has been distilled to match the quality of Krea 1 with a focus on preserving aesthetics and photorealism.
Unlike most image models, FLUX.1 Krea has been created with opinionated aesthetics in mind. We focused on creating a model that truly fits our specific aesthetic preferences. In this technical report, we'll share the process and learnings from developing this model, including insights on pre-training and post-training, as well as future research directions.
"When a measure becomes a target, it ceases to be a good measure"
—Charles Goodhart
Image generation has come a long way since the early days of generating cats and flowers with GANs. Today's models can generate coherent human faces, limbs, and hands. They understand exact quantities, render complex typography, and make an astronaut ride a horse.
However, a clear trend when working with AI generated images is their unique look: overly-blurry backgrounds, waxy skin textures, boring composition, and more. Together, these problems constitute what is now known as the "AI look".

Some examples of the "AI look" in human faces
People often focus on how "smart" a model is. We often see users testing complex prompts. Can it make a horse ride an astronaut? Does it fill up the wine glass to the brim? Can it render text properly? Over the years, we have devised various benchmarks to formalize these questions into concrete metrics. The research community has done a remarkable job advancing generative models. However, in this pursuit of technical capabilities and benchmark optimization, the messy genuine look, stylistic diversity, and creative blend of early image models took a backseat.

Generations from DALLE-2, an imperfect model which produced flawed but interesting outputs. Source
Our goal from the beginning was simple: "Make AI images that don't look AI." As users of generative AI ourselves, we wanted to create a model that addressed these issues. Unfortunately, many of the academic benchmarks and metrics are misaligned with what users actually want.
For the pre-training stage, metrics such as Fréchet inception distance (FID) and CLIP Score are useful for measuring general performance of the model since most images at this stage are incoherent. Beyond the pre-training stage, evaluation benchmarks such as DPG, GenEval, T2I-Compbench, and GenAI-Bench are widely used to benchmark academic and industry models. But, these benchmarks are limited to measuring prompt adherence with focus on spatial relationships, attribute binding, object counts, etc.
For evaluating aesthetics, models such as LAION-Aesthetics, Pickscore, ImageReward, HPSv2 are commonly used, but many of these models are finetuned variants of CLIP, which processes low-resolution images (224×224 pixels) with limited parameter count. As the capability of image generation models has increased, these older aesthetic score models are no longer good enough to evaluate them.
For instance, we find that LAION Aesthetics — a model commonly used to obtain high quality training images — to be highly biased towards depicting women, blurry backgrounds, overly soft textures, and bright images. While these aesthetic scorers and image quality filters are useful for filtering out bad images, relying on these models to obtain high quality training images adds implicit biases to the model's priors.

Examples of images that are in the top 5% of LAION aesthetics
While better aesthetic scorers based on vision language models ([1], [2]) are emerging, the issue remains that human preference and aesthetics are highly personal. They cannot be easily reduced to a single number. Advancing model capabilities without regressing towards the "AI look" requires careful data curation and thorough calibration of model outputs.
"The sculpture is already complete within the marble block, before I start my work. It is already there, I just have to chisel away the superfluous material."
—Michelangelo
Training an image generation model can be roughly divided into two stages: pre-training and post-training. Most of the aesthetics of a model are learned during the post training stage, but before explaining our post-training methodology, let's go over some intuition on how we think about these training stages.
Good mode coverage
Bad mode coverage
Pre-training is all about mode coverage, post-training is all about mode collapsing.
The focus of the pre-training stage should be all about "mode coverage" and "world understanding." During this stage, we give the model rich knowledge about the visual world: styles, objects, places, people. The goal here is to maximize diversity.
We would even argue that the pre-trained model should be trained on "bad" data, as long as the undesirable aspects of the data are accurately captured in its conditioning. Indeed, in addition to telling the model what we want, we often want to tell it what we don't want.
Many image generation workflows use negative prompts like "too many fingers, deformed faces, blurry, oversaturated" to improve image quality. For the negative prompt to steer the model away from undesirable parts of the data distribution, it must first have learned what these undesirable parts look like. Negative prompting would not be effective if the model never saw examples of "bad images."
While post-training has the highest impact on the final quality of the model, it's important to remember that the quality ceiling of the model and the stylistic diversity comes from the pre-trained model.
During post-training, the focus should be on shifting and chipping away the undesirable part of the distribution. A pre-trained model can output diverse images and understands a wide range of concepts, but struggles to reliably output high quality images as it's not biased enough towards producing aesthetic outputs. This is where mode collapsing comes into play: we want to start biasing the model towards the part of the distribution we find desirable.
To start post-training, we need a "raw" model. We want a malleable base model with a diverse output distribution that we can easily reshape towards a more opinionated aesthetic. Unfortunately, many existing open weights models have been already heavily finetuned and post-trained. In other words, they are too "baked" to use as a base model.
To be able to fully focus on aesthetics, we partnered with a world-class foundation model lab, Black Forest Labs, who provided us with flux-dev-raw, a pre-trained and guidance-distilled 12B parameter diffusion transformer model.

Samples from flux-dev-raw generations
As a pre-trained base model, flux-dev-raw does not achieve image quality anywhere near that of state-of-the-art foundation models. However, it is a strong base for post-training for three reasons:
Our post-training pipeline is split into two stages. A Supervised Finetuning (SFT) stage and Reinforcement Learning from Human Feedback (RLHF) stage. During the supervised finetuning stage, we hand curate a dataset of the highest quality of images that match our aesthetic standards. For training FLUX.1 Krea [dev], we also incorporate high quality synthetic samples from Krea-1 during SFT stage. We find that synthetic images to be beneficial for stabilizing the performance of the checkpoint.
Since flux-dev-raw is a guidance distilled model, we devise a custom loss to finetune the model directly on a classifier-free guided (CFG) distribution. After the SFT stage, the model's image quality output is significantly improved. However, further work is needed to make the model more robust and nail the aesthetics we are looking for. This is where RLHF comes in.
Pre-training

SFT

RLHF

During RLHF, we apply a variant of preference optimization technique which we call TPO to further boost the aesthetics and stylization of our model. We use high quality internal preference data which has been rigorously filtered to ensure data quality. In many cases, we applied multiple rounds of preference optimization to further calibrate the model's outputs.
While exploring various post-training techniques, we discovered a few key findings we would like to share.
You need a surprisingly small amount of data (< 1M) to do good post-training. Quantity helps with stability and mitigating biases, but the quality of the data matters the most. This observation is in line with previous literature that report the effectiveness of training on small set of carefully curated data ([3], [4], [5])
Our preference labels were carefully collected from labelers who were acutely aware of the current model's limitation, areas of improvement, strengths, and weaknesses. In particular, we ensured that the images in our preference annotation interface contained a diverse set to obtain a focused annotation.
There are many open source preference datasets ([6], [7], [8]) that have been used to benchmark preference finetuning techniques. During exploration stages, these datasets were useful for testing various techniques. However, we found that training on existing datasets led to unintended behaviors such as:
It's our belief that a model that has been finetuned on "global" user preference is suboptimal. For goals like text rendering, anatomy, structure, and prompt adherence where there's an objective ground truth, data diversity and scale are helpful. However, for subjective goals such as aesthetics, it's almost adversarial to mix different aesthetic tastes together.
High fashion photographyMinimalism lover
Pre-training is all about mode coverage, post-training is all about mode collapsing.
For example, consider a case where one user loves high fashion photography and another user is into minimalist style drawings. Given a focused annotation from the respective users, it would be easy to align the model to excel at respective styles. But, when you merge the two distributions together, we get a marginal preference distribution which is not biased enough to make either party happy. This limitation can be partially addressed by prompting, but it's not a satisfactory solution. Most people often end up relying on LoRAs to get the level of stylization they want out of the model because prompting is insufficient for their use case. Furthermore, users often want reasonable defaults without extensive prompting and adding modifiers to get aesthetic outputs from the model.
Nobody's happy here
Global preference will make both parties unsatisfied.
Motivated by this intuition, we decided to collect our preference data in a very opinionated manner which aligns with our aesthetic taste with a clear art direction. It's often better and easier to overfit a model towards a certain style.
As a product-focused company, we focus on building intuitive and delightful user experiences for interacting with generative models. We see Krea 1 as our first step to offering a model that meets the aesthetic standard and quality that creatives have been craving for. With the open release of FLUX.1 Krea [dev], we are excited to see what the open source community will build on top of it.
We plan to improve core capabilities of the model as well as expanding to more visual domains to allow our users to explore, blend, and mix diverse set of visuals.
This work was our first step into aesthetics research. We have built a model that to provide an opinionated aesthetic, but we want to build something that is more personal and tailored to your sense of aesthetics. In future works, through personalization, aesthetics, and controllability research, we hope to provide you a model that clicks with your taste and the ability to refine your work.
We would like to thank the Black Forest Labs team for providing us with their base model weights. None of this would be possible without their contribution. Additionally, we thank our data, infrastructure, and product teams, whose hard work was key to building a foundation for our post-training pipeline.
@misc{flux1kreadev2025, author={Sangwu Lee, Titus Ebbecke, Erwann Millon, Will Beddow, Le Zhuo, Iker García-Ferrero, Liam Esparraguera, Mihai Petrescu, Gian Saß, Gabriel Menezes, Victor Perez}, title={FLUX.1 Krea [dev]}, year={2025}, howpublished={\url{https://github.com/krea-ai/flux-krea}} }
Want to contribute to work like this?
Join us — we're hiring
Then optimise for max (Quality + A*R)
Arguably amplitude of A should do R but I think the AI-ness and the AI-ness-relevance are distinct concepts (It could be highly relevant but it can't tell what it should be).
Not the real reason. The real reason is that training has moved to FP/BF16 over the years as NVIDIA made that more efficient in their hardware, the same reason you're starting to see some models being released in 8bit formats (deepseek).
Of course people can always quantize the weights to smaller sizes, but the master versions of the weights is usually 16bit.
I don't see a difference.
The vast majority of humans (>99.99999%) can tell that those things are different. The fact that you can't should be somewhat concerning.
> The vast majority of humans (>99.99999%) can tell that those things are different.
Citation needed.
That's clearly not what I said, or what you claimed. You claimed that looking at real things while living life is the same thing as viewing millions of artificial images made by humans:
>> Looking at real things while living life is categorically different from viewing millions of artworks made by other artists
> I don't see a difference.
You're now actively lying about your claims (and mine) and it's clear that you aren't interested in actually debating your point, just performative statements.
I'm not going to debate this further, just going to call out your lies and fallacies for the record.
It's pretty interesting that everyone I've talked to who wants to steal the work of tens/hundreds of thousands of artists to train ML on resorts to lying and clinically insane statements to try to justify their beliefs.