Generated clips are evaluated against the task prompt, conditioning frame, and a recorded robot run. Model identities are hidden, and every score links to its evidence.
Before publishing a rubric version, we compare the automated evaluator's labels with human labels and report their agreement.
Each case contains the task prompt, conditioning frame, and recorded robot trajectory used as the reference.
The evaluator compares clips without seeing model names. Left-right order is randomized, and every pair is evaluated again with the positions reversed.
Pairwise preferences update Elo ratings. Per-axis labels are reported as pass rates. Every published score links to the corresponding clips and evaluator notes.