Leaderboard

PRM-as-a-Judge metrics on 42 simulation tasks and 18 real-world tasks, covering 6,076 published rollouts.

BenchmarkRoboDojo — Simulation Last UpdatedAugust 18, 2026 EvaluatorRobo-Dopamine (GRM-2.0-8B-Preview) Tasks42 Models16 Published Rollouts6,076

The RoboDojo ranking is computed from process metrics produced by Robo-Dopamine (GRM-2.0-8B-Preview). For the complete evaluation setup, model coverage, and detailed analysis, see our technical report.

Methodology and metric definitions are described in our technical blog, and we will keep this leaderboard synchronized as new benchmark collaborations go live.

Community Call For Transparent Evaluation

We encourage benchmark organizers and model developers to release transparent execution videos together with leaderboard submissions. Open rollout evidence makes it possible to inspect not only whether a policy succeeded, but also how it succeeded or failed.

PRM-as-a-Judge supports transparent process evaluation across benchmarks, including RoboChallenge and RoboDojo. We welcome benchmark teams and researchers to collaborate with us on broader evaluation coverage, shared trajectory auditing, and stronger PRM evaluators.

We welcome collaboration and feedback from the robotics community. If you have questions or would like to work with us on dense trajectory evaluation, please reach out at liuyuyang2025@ia.ac.cn.

RoboDojo Simulation — Model Ranking

Models are independently ranked by SR within the selected environment.

Loading...

RoboDojo Simulation — Per-Task Ranking

Select a task to compare models by SR and process metrics.

Loading...

Citation

If this leaderboard or evaluation pipeline helps your work, please cite both PRM-as-a-Judge papers:

@article{ji2026prmjudge,
  title   = {PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing},
  author  = {Ji, Yuheng and Liu, Yuyang and Tan, Huajie and Huang, Xuchuan and Huang, Fanding and Xu, Yijie and Chi, Cheng and Zhao, Yuting and Lyu, Huaihai and Co, Peterson and Cao, Mingyu and Zhang, Qiongyu and Li, Zhe and Zhou, Enshen and Wang, Pengwei and Wang, Zhongyuan and Zhang, Shanghang and Zheng, Xiaolong},
  journal = {arXiv preprint arXiv:2603.21669},
  year    = {2026},
  url     = {https://arxiv.org/abs/2603.21669}
}

@article{liu2026prmjudge15,
  title   = {PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment},
  author  = {Liu, Yuyang and Shen, Yanqing and Chen, Ruike and Zhao, Jifan and Tian, Yuxuan and Zhang, Yichi and Long, Tianfeng and Yin, Zixuan and Wang, Yipu and Qin, Ziheng and Tan, Wenxing and Shi, Yang and Cao, Mingyu and Xiao, Runze and Wang, Ziqi and Yin, Zhixin and Chu, Shiwei and Zhang, Yi-Fan and Mu, Yao and Ji, Yuheng and Wang, Yihao and Yan, Jun and Wang, Zhongyuan and Wang, Pengwei and Zheng, Xiaolong},
  journal = {arXiv preprint arXiv:2608.14284},
  year    = {2026},
  url     = {https://arxiv.org/abs/2608.14284}
}