How do we post-train a text-to-image model? While recent advances have made reinforcement learning increasingly practical for diffusion-based models, improving frontier open-source models remains surprisingly difficult. The key question is: what reward should we optimize?
Human preference is an obvious starting point. If people consistently prefer one image over another, a reward model trained on those choices can provide a powerful learning signal. But preference alone does not capture everything users care about. An image can look great while missing an object from the prompt, adding things that were never requested, or drifting away from the requested style.
In our latest work, we explore a simple idea: combine human preference with explicit rubric-based rewards. We train an Arena reward model on roughly 5 million pairwise human preference votes, and use vision-language models to automatically construct rubrics that evaluate whether an image follows the prompt, respects relevant constraints, and avoids known reward-hacking behaviors.
The result is a post-training recipe that improves two already-strong text-to-image models under live human evaluation. Our post-trained FLUX.2-dev gains 69 Arena points in our live T2I leaderboard, while our post-trained Ideogram 4 reaches a score of 1224 - and surpasses every publicly listed open-source model by Sep 04, 2026.

Reward Design for T2I Post-Training
Recent methods such as Flow-GRPO and DiffusionNFT have made reinforcement learning increasingly practical for diffusion- and flow-based image generation models. But what reward actually represents a better image?
Learning Human Preference from Pairwise Comparisons
Our first component is a reward model trained on real-world human preferences. Using roughly 5 million pairwise human votes collected from Text-to-Image Arena across more than 100 models, we train a Bradley–Terry reward model. Given a prompt and an image, the model outputs a scalar preference score that captures broad human judgments such as visual quality, composition, and aesthetics. The following figure shows that our image reward model outperforms other state-of-the-art models on the MMRB2 benchmark. Further, the post training performance improves with more reward model training data.

Faithfulness Reward through Auto-Rubrics
Large-scale preference data provides a strong and general signal, but it does not explicitly check whether every detail in the prompt has been correctly followed. To explicitly measure prompt following, we introduce a faithfulness reward based on automatically generated, prompt-specific rubrics.
For each training prompt, we use a language model to decompose the prompt into a tree-structured checklist of concrete yes/no questions. Each question evaluates one aspect of the requested image, such as whether an object is present, whether it has the correct attribute, whether two objects have the requested spatial relationship, or whether a specified style is preserved. Dependencies between questions are also captured—for example, checking the color of an object only makes sense if that object is present.
A vision-language model then evaluates the generated image against these questions, and the fraction of satisfied criteria becomes the faithfulness reward. More details can be found in our paper [1].
Other Rubric Rewards and Ensembling
Faithfulness autorubric methods can construct questions that check basic object presence, attributes, and relationships. However, they are less suited to checking constraints, such as whether the user wants to avoid an undesired style, especially when these constraints are implicit. More importantly, these autorubrics are derived from the prompt before training and therefore cannot anticipate reward-hacking behaviors that emerge during RL training.
To complement this, we introduce rubric-based rewards for constraint satisfaction and reward-hacking prevention. The constraint reward checks whether the model introduces content that conflicts with the user’s intent—for example, adding unnecessary objects, changing a requested style, or violating negative instructions. Because some prompts are open-ended and allow creative freedom, we apply this reward only when the prompt calls for relatively strict adherence.
We also use rubric-based rewards for recurring reward-hacking behaviors that emerge during RL, such as adding garbled text or drifting toward photorealism when a non-photographic style is requested.
Rather than treating these anti-reward hack rewards as independent reward objectives, we use them to shape the preference reward through a simple gating function. Specifically, let denote the preference reward and we define a binary indicator indicating whether a reward-hacking behavior is detected. We use the reward hack prevention reward to modulate the normalized preference reward:
Thus, the preference reward is preserved when no reward-hacking behavior is detected, while it is canceled when the anti-reward-hacking signal is flagged. This allows the detector to suppress known exploits without introducing another independent objective that the policy can optimize or game.
Finally, we found that models trained with different reward configurations often develop complementary strengths. We therefore ensemble independently optimized policies directly in weight space by averaging their weight updates [2]. Together with independent reward normalization and prompt-dependent gating, this gives us a simple reward-composition recipe that works consistently without extensive reward-weight tuning.
Offline Results: Why the Reward Composition Matters
We train on 10K real user prompts sampled from Arena and evaluate on a held-out set of 1K Arena prompts. For each checkpoint, we compare the post-trained model against the frozen base model using the MMRBv2 pairwise evaluation protocol, with Gemini-3.5-Flash as the judge. Each image pair is evaluated in both presentation orders to reduce position bias, and we report the resulting win rate against the base model.
The results show a clear progression. Optimizing preference alone is not sufficient; Adding the faithfulness reward substantially improves performance, and introducing the intent-gated constraint reward further raises the win rate to 64.2%. Finally, ensembling policies trained with complementary reward configurations gives the strongest result, reaching 66.0% win rate with Arena RM. The same recipe also holds if using the open source reward model such as Pickscore.
The following visual examples compare images generated with and without anti-reward-hack reward. We observe a photorealistic hacking mode during training.

Read More
For more details on our methods, experiments, and analysis, please see our papers:
References
[1] Ban, Y., Xie, T., An, S., Hong, Y., Frick, E., Hsu, I., ... & Hsieh, C. J. (2026). Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist. In NeurIPS, 2026 (Spotlight).
[2] Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., ... & Schmidt. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In ICML 2022.










