Updated 31 Jul 2026

World-model benchmarks for embodied tasks.

Generated clips are evaluated against the task prompt, conditioning frame, and a recorded robot run. Model identities are hidden, and every score links to its evidence.

Cycle
World model leaderboard
0
Evaluation decisions
0
Models evaluated
0
Evaluation axes
0
Agreement with human labels
0
Position-swap consistency
Method

How scores are produced

Before publishing a rubric version, we compare the automated evaluator's labels with human labels and report their agreement.

Inputs for each case

Each case contains the task prompt, conditioning frame, and recorded robot trajectory used as the reference.

Blind comparison with position swap

The evaluator compares clips without seeing model names. Left-right order is randomized, and every pair is evaluated again with the positions reversed.

Ratings, pass rates, and evidence

Pairwise preferences update Elo ratings. Per-axis labels are reported as pass rates. Every published score links to the corresponding clips and evaluator notes.