
How to Post-Train Text-to-Image Models: Combining Preference and Rubric Rewards
We introduce a post-training approach that combines human preference signals with explicit rubric-based rewards. Trained on 5M Arena pairwise preference votes, it pushes FLUX.2-dev to #2 on the live T2I leaderboard and Ideogram 4 past every publicly listed open-source model.
Research
Arena Team—2 Oct 2026









