Four-model arena
Same frozen story. Four models. One judge. The winner airs — every score stays on the record.
Where the field actually sits
Mean judge score per contest, on one axis clipped to the range every plotted contest occupies. The four averages sit high on it because single contests reach further both ways. The tick marks the leader; each connector is that model's deficit. Every figure below shares this axis.
What a point of score costs
How each model behaves
Win rate and judge score across recent contests, on a shared vertical scale. Shines and watch list the judging criteria where a model scores above or below the field average.
Latest contest
—
What the arena is showing
Each is a claim plus the number it rests on, derived from the record at render time.
Recent arenas
A filled dot is the model that aired; open dots are the candidates it beat. Rows marked not judged ran but were never scored, so they carry no marks — not a row of zeros.