RoboLab Leaderboard

Overall Difficulty Competency Axes

RoboLab-120 Overall

Performance of task generalist models across three levels of language specificity.

Showing all policies
Type
Inputs
Includes policies declared open source, even when a download link has not yet been published. Shows policies with no reported RoboLab or other simulation data in training or fine-tuning.
# Policy Type N SR% Score EE Speed EE SPARC Size VRAM Obs
1 Lakeside LWM 1531 / 3000
51.0% 65.3 7.10 cm/s −4.84 5B 27.2 GB RGBP
2 FLUX 3 Action WAM 515 / 1200
42.9% — 7.36 cm/s −5.37 7B 69 GB RGBP
3-TIE HiDream-O1-Embodied VLA 479 / 1200
39.9% 50.5 5.32 cm/s −4.95 6B 16 GB RGBP
3-TIE Atomic-WAM VLM+WAM 475 / 1200
39.6% 55.5 7.69 cm/s −4.77 16.2B 40 GB RGBP
3-TIE OASIS WAM VLM+WAM 468 / 1200
39.0% 53.7 6.91 cm/s −5.26 16B + VLM 40 GB RGBP
6 Cosmos3-Nano-Policy WAM 441 / 1200
36.8% 51.9 7.13 cm/s −5.99 16B 40 GB RGBP
7 Phoenix TAMP+FM 413 / 1200
34.4% 45.9 6.13 cm/s −5.55 ~250M 14 GB RGBDP
8 BiMind v0.1 VLA 400 / 1200
33.3% 47.2 5.26 cm/s −5.88 7B 14 GiB RGBP
9 VoLo Agent 339 / 1200
28.2% 44.1 5.50 cm/s −8.92 — 80 GB RGBDP
10 π0.5 VLA 336 / 1200
28.0% 43.4 5.35 cm/s −8.34 3.3B >8 GB RGBP
11 DreamZero WAM 308 / 1200
25.7% 39.8 3.64 cm/s −6.41 14B 2×80 GB RGBP
12 Cosmos3-Edge-Policy WAM 275 / 1200
22.9% — 7.04 cm/s −5.46 4B 9.20 GB RGBP
13 π0-FAST VLA 186 / 1200
15.5% 26.9 4.60 cm/s −9.63 3B >8 GB RGBP
14 GR00T N1.6 VLA 87 / 1200
7.2% 17.1 4.30 cm/s −6.87 3B 24 GB RGBP
15 π0 VLA 60 / 1200
5.0% 12.2 4.18 cm/s −9.49 3.3B >8 GB RGBP
16 paligemma-binning VLA 41 / 1200
3.4% 9.9 1.93 cm/s −16.52 3B >8 GB RGBP
Verified ◆RoboLab/IsaacSim/IsaacLab data in fine-tuning ◈RoboLab/IsaacSim/IsaacLab data in training ◇Other simulation data in training Closed-source weights

Results by Task Difficulty

Difficulty reflects the number of subtasks and reasoning complexity per task, grouped into 3 difficulty levels.

Simple Moderate Complex
# Policy N SR% Score N SR% Score N SR% Score
1 Lakeside 921 / 1600
57.6% 65.2 430 / 975
44.1% 63.5 180 / 425
42.4% 69.5
2 FLUX 3 Action 314 / 640
49.1% — 153 / 390
39.2% — 48 / 170
28.2% —
3 HiDream-O1-Embodied 276 / 640
43.1% 47.7 147 / 390
37.7% 54.3 56 / 170
32.9% 52.2
4 Atomic-WAM 284 / 640
44.4% 54.6 158 / 390
40.5% 58.1 33 / 170
19.4% 53.1
5 OASIS WAM 265 / 640
41.4% 50.7 148 / 390
37.9% 55.6 55 / 170
32.4% 60.5
6 Cosmos3-Nano-Policy 260 / 640
40.6% 51.0 138 / 390
35.4% 52.3 43 / 170
25.3% 54.4
7 Phoenix 274 / 640
42.8% 51.9 100 / 390
25.6% 40.0 39 / 170
22.9% 37.0
8 BiMind v0.1 248 / 640
38.8% 46.2 128 / 390
32.8% 50.9 24 / 170
14.1% 42.5
9 VoLo 209 / 640
32.7% 43.7 100 / 390
25.6% 46.7 30 / 170
17.6% 39.8
10 π0.5 190 / 640
29.7% 38.9 123 / 390
31.5% 50.5 23 / 170
13.5% 44.1
11 DreamZero 167 / 640
26.1% 35.5 117 / 390
30.0% 48.4 24 / 170
14.1% 35.9
12 Cosmos3-Edge-Policy 164 / 640
25.6% — 91 / 390
23.3% — 20 / 170
11.8% —
13 π0-FAST 129 / 640
20.2% 26.7 52 / 390
13.3% 29.2 5 / 170
2.9% 21.9
14 GR00T N1.6 56 / 640
8.8% 13.6 31 / 390
7.9% 25.1 0 / 170
0.0% 11.9
15 π0 46 / 640
7.2% 9.6 14 / 390
3.6% 18.1 0 / 170
0.0% 8.7
16 paligemma-binning 22 / 640
3.4% 6.2 19 / 390
4.9% 16.2 0 / 170
0.0% 9.0

Results by Competency Axes

Performance across competency axes: Visual (color, semantics, size), Procedural (action-oriented reasoning across affordances, reorientation, and stacking), and Relational (interpreting conjunctions, counting, and spatial relationships in multi-object instructions).

Relational Visual Procedural
# Policy N SR% Score N SR% Score N SR% Score
1 Lakeside 564 / 1050
53.7% 67.4 993 / 2100
47.3% 61.8 169 / 550
30.7% 50.2
2 FLUX 3 Action 208 / 420
49.5% — 299 / 840
35.6% — 79 / 220
35.9% —
3 HiDream-O1-Embodied 202 / 420
48.1% 56.4 224 / 840
26.7% 39.1 65 / 220
29.5% 46.7
4 Atomic-WAM 189 / 420
45.0% 57.9 286 / 840
34.0% 49.9 57 / 220
25.9% 46.1
5 OASIS WAM 196 / 420
46.7% 59.3 248 / 840
29.5% 45.3 67 / 220
30.5% 50.9
6 Cosmos3-Nano-Policy 207 / 420
49.3% 60.9 219 / 840
26.1% 42.9 60 / 220
27.3% 46.8
7 Phoenix 105 / 420
25.0% 37.3 282 / 840
33.6% 45.4 48 / 220
21.8% 30.3
8 BiMind v0.1 150 / 420
35.7% 51.7 247 / 840
29.4% 43.0 47 / 220
21.4% 33.3
9 VoLo 134 / 420
31.9% 44.1 196 / 840
23.3% 39.3 36 / 220
16.4% 40.4
10 π0.5 142 / 420
33.8% 46.8 200 / 840
23.8% 38.8 48 / 220
21.8% 41.6
11 DreamZero 140 / 420
33.3% 45.5 147 / 840
17.5% 31.9 58 / 220
26.4% 42.3
12 Cosmos3-Edge-Policy 112 / 420
26.7% — 141 / 840
16.8% — 33 / 220
15.0% —
13 π0-FAST 98 / 420
23.3% 36.7 74 / 840
8.8% 19.5 14 / 220
6.4% 16.4
14 GR00T N1.6 42 / 420
10.0% 23.2 42 / 840
5.0% 15.1 6 / 220
2.7% 11.1
15 π0 27 / 420
6.4% 18.8 16 / 840
1.9% 8.2 1 / 220
0.5% 6.6
16 paligemma-binning 24 / 420
5.7% 15.6 23 / 840
2.7% 9.3 1 / 220
0.5% 4.0

Sim-Real Correlation

Benchmark-level Correlation with Real World Experiments

Performance ranking correlate strongly with RoboArena Elo scores, an open-source real-world benchmark for generalist robot policies. This means that at a benchmark level, policies rank in our benchmark as they do in the real world.

Spearman ρ —

Adding Your Model

If you have a model that you'd like to see featured, fill out the form below:


Questions? Get in touch.