RoboLab Leaderboard

Overall Difficulty Competency Axes

RoboLab-120 Overall

Performance of task generalist models across three levels of language specificity.

# Policy Type N SR% Score EE Speed EE SPARC Robot Cams Obs Action
1 OASIS WAM VLM + WAM 468 / 1200
39.0% 53.7 6.91 cm/s −5.26 DROID EGO-L/RWrist RGBP Joint Pos
2 Cosmos3-Nano-Policy WAM 441 / 1200
36.8% 51.9 7.13 cm/s −5.99 DROID EGO-L/RWrist RGBP Joint Pos
3 Phoenix TAMP+FM 413 / 1200
34.4% 45.9 6.13 cm/s −5.55 DROID EGO-LWrist RGBDP Joint Pos
4 BiMind v0.1 VLA 400 / 1200
33.3% 47.2 5.26 cm/s −5.88 DROID EGO-LWrist RGBP Joint Pos
5 π0.5 VLA 336 / 1200
28.0% 43.4 5.35 cm/s −8.34 DROID EGO-LWrist RGBP Joint Pos
6 DreamZero WAM 308 / 1200
25.7% 39.8 3.64 cm/s −6.41 DROID EGO-L/RWrist RGBP Joint Pos
7 Cosmos3-Edge-Policy WAM 275 / 1200
22.9% 7.04 cm/s −5.46 DROID EGO-L/RWrist RGBP Joint Pos
8 π0-FAST VLA 186 / 1200
15.5% 26.9 4.60 cm/s −9.63 DROID EGO-LWrist RGBP Joint Pos
9 GR00T N1.6 VLA 87 / 1200
7.2% 17.1 4.30 cm/s −6.87 DROID EGO-LWrist RGBP Joint Pos
10 π0 VLA 60 / 1200
5.0% 12.2 4.18 cm/s −9.49 DROID EGO-LWrist RGBP Joint Pos
11 paligemma-binning VLA 41 / 1200
3.4% 9.9 1.93 cm/s −16.52 DROID EGO-LWrist RGBP Joint Pos
Verified RoboLab/IsaacSim/IsaacLab data in fine-tuning RoboLab/IsaacSim/IsaacLab data in training Other simulation data in training Closed-source weights

Results by Task Difficulty

Difficulty reflects the number of subtasks and reasoning complexity per task, grouped into 3 difficulty levels.

Simple Moderate Complex
# Policy N SR% Score N SR% Score N SR% Score
1 OASIS WAM 265 / 640
41.4% 50.7 148 / 390
37.9% 55.6 55 / 170
32.4% 60.5
2 Cosmos3-Nano-Policy 260 / 640
40.6% 51.0 138 / 390
35.4% 52.3 43 / 170
25.3% 54.4
3 Phoenix 274 / 640
42.8% 51.9 100 / 390
25.6% 40.0 39 / 170
22.9% 37.0
4 BiMind v0.1 248 / 640
38.8% 46.2 128 / 390
32.8% 50.9 24 / 170
14.1% 42.5
5 π0.5 190 / 640
29.7% 38.9 123 / 390
31.5% 50.5 23 / 170
13.5% 44.1
6 DreamZero 167 / 640
26.1% 35.5 117 / 390
30.0% 48.4 24 / 170
14.1% 35.9
7 Cosmos3-Edge-Policy 164 / 640
25.6% 91 / 390
23.3% 20 / 170
11.8%
8 π0-FAST 129 / 640
20.2% 26.7 52 / 390
13.3% 29.2 5 / 170
2.9% 21.9
9 GR00T N1.6 56 / 640
8.8% 13.6 31 / 390
7.9% 25.1 0 / 170
0.0% 11.9
10 π0 46 / 640
7.2% 9.6 14 / 390
3.6% 18.1 0 / 170
0.0% 8.7
11 paligemma-binning 22 / 640
3.4% 6.2 19 / 390
4.9% 16.2 0 / 170
0.0% 9.0

Results by Competency Axes

Performance across competency axes: Visual (color, semantics, size), Procedural (action-oriented reasoning across affordances, reorientation, and stacking), and Relational (interpreting conjunctions, counting, and spatial relationships in multi-object instructions).

Relational Visual Procedural
# Policy N SR% Score N SR% Score N SR% Score
1 OASIS WAM 196 / 420
46.7% 59.3 248 / 840
29.5% 45.3 67 / 220
30.5% 50.9
2 Cosmos3-Nano-Policy 207 / 420
49.3% 60.9 219 / 840
26.1% 42.9 60 / 220
27.3% 46.8
3 Phoenix 105 / 420
25.0% 37.3 282 / 840
33.6% 45.4 48 / 220
21.8% 30.3
4 BiMind v0.1 150 / 420
35.7% 51.7 247 / 840
29.4% 43.0 47 / 220
21.4% 33.3
5 π0.5 142 / 420
33.8% 46.8 200 / 840
23.8% 38.8 48 / 220
21.8% 41.6
6 DreamZero 140 / 420
33.3% 45.5 147 / 840
17.5% 31.9 58 / 220
26.4% 42.3
7 Cosmos3-Edge-Policy 112 / 420
26.7% 141 / 840
16.8% 33 / 220
15.0%
8 π0-FAST 98 / 420
23.3% 36.7 74 / 840
8.8% 19.5 14 / 220
6.4% 16.4
9 GR00T N1.6 42 / 420
10.0% 23.2 42 / 840
5.0% 15.1 6 / 220
2.7% 11.1
10 π0 27 / 420
6.4% 18.8 16 / 840
1.9% 8.2 1 / 220
0.5% 6.6
11 paligemma-binning 24 / 420
5.7% 15.6 23 / 840
2.7% 9.3 1 / 220
0.5% 4.0

Sim-Real Correlation

Benchmark-level Correlation with Real World Experiments

Performance ranking correlate strongly with RoboArena Elo scores, an open-source real-world benchmark for generalist robot policies. This means that at a benchmark level, policies rank in our benchmark as they do in the real world.

Spearman ρ

Adding Your Model

If you have a model that you'd like to see featured, fill out the form below:


Questions? Get in touch.