How Embodied Analysis evaluates world models and action models, and what the evaluations show.
How we evaluate action models. VLAs and VLMs run the same tasks in simulation and on real robots, and human pairwise judgments of each run are aggregated into Elo.
How we evaluate world models. Humans compare generated robot videos pairwise on six dimensions, from physical plausibility to goal completion, and the results are aggregated into Elo.