EgoWM Benchmark
Long-term leaderboard and evaluation protocol for grounded egocentric hand + camera + depth prediction on cached-latent video foundation models.
Follows the modular pattern from Large Video Planner (registered datasets + registered models + protocol + JSON results), adapted for the per-hand-per-frame metric-3D surface EgoWM optimizes.
Repository: benchmark/
Full training arc — v10 through v33
The absolute-MPJPE convergence across 20+ model iterations, stacked on a cumulative training-step axis:

Every iter, every metric (6-panel: abs / rr / PA / wrist / pixel / depth across all iters on cumulative x-axis):

The gold star marks v33 s3000 (current SOTA: 62.91 mm abs).
Convergence, v10 → v12 (earlier zoom)
Focused view of the first arc through v12:


Multi-model side-by-side video
Six-column comparison across model versions on 3 ARCTIC val clips (v8d, v9b, v10, v11, v12, v13 side-by-side; skeleton overlay + per-frame MPJPE ticker):
5-column (v8d → v12)
4-column (v8d → v11)
3-column (v8d → v10)
4-panel per-clip galleries
Each gallery: RGB / GT+pred skeleton / depth heatmap / 3D trajectory, per-clip video, per-clip metrics under each card.
- 🎞️ v33 champion gallery (62.91 mm) — 6 arctic clips
- Best clip s08 box_use_01 · abs 47 mm · pixel 11 px — genuinely at Iter 206 oracle territory!
- 🎞️ v13 gallery
- 🎞️ v12 gallery
- 🎞️ v11 gallery
- 🎞️ v10 12-clip gallery
Per-source generalization (v12 step 3000 FINAL vs v11, 25 clips/source)
Cross-source eval. Reveals the training-time hand_valid fix helps arctic massively but hurts ho3dv3 — confirming ho3dv3 GT is noisier than the model:
| Source | v11 s3000 abs | v12 s3000 abs | Δ | v11 pix | v12 pix | Δ |
|---|---|---|---|---|---|---|
| arctic 🥇 | 161 mm | 118 mm | −27 % | 88 px | 36 px | −59 % |
| h2o | 151 mm | 148 mm | −2 % | 80 px | 78 px | −3 % |
| ho3dv3 | 255 mm | 269 mm | +5 % | 196 px | 196 px | 0 % |
| asmhand | — | — | — | — | ||
| dexycb | — | — | — | — | ||
| hot3d_aria | — | — | — | — | ||
| hot3d_quest3 | — | — | — | — |
Big win — arctic: 161 → 118 mm (−43 mm, −27%) and 88 → 36 px (−52 px, −59%). Training-time fix compounded huge on arctic where the mislabeled off-frame frames were concentrated.
Modest spillover — h2o: 151 → 148 mm. h2o still slightly better on pixel (78 vs arctic's 36 is on N=25 vs arctic's arctic-lite N=40 = 34 — apples-apples on N=25 arctic is 36 px). h2o benefit is smaller because it has fewer off-frame mislabels to begin with.
Regression — ho3dv3: 255 → 269 mm (+14 mm). Confirms Iter 204/206 flag: ho3dv3 GT labels are noisier than the model. Any improvement that better fits the ARCTIC-side signal makes the model less compatible with ho3dv3's noisy labels. Recommendation: drop ho3dv3 from the training pool.
Data engine debt (blocks 3 of 7 sources): - asmhand: Iter 200 fix for NaN global_orient labels masks ~100% of val frames - dexycb / hot3d_aria / hot3d_quest3: manifest builder OOB (some clips reference frame indices past zarr end)
Both bugs are known (docs/20_depth_fusion_analysis Iter 204/205), fix is straightforward but hasn't been prioritized. When fixed, we can add proper per-source rows to the leaderboard.
Live leaderboard — ARCTIC-lite
Protocol: 40 ARCTIC val clips from the v8d all7 split. Same seed across models → fair comparison.
| Model | Steps | abs MPJPE ↓ | rr MPJPE | PA MPJPE | wrist trans | pixel err | depth log-res |
|---|---|---|---|---|---|---|---|
| oracle-iter204 (GT-orient + GT-trans) | oracle | 17 mm | — | 13 mm | — | — | — |
| oracle-iter206 (GT-2D + per-clip s*) | oracle | 54 mm | — | — | — | 0 px | — |
| v13_stable (FINAL) 🥇 | 3000 | 90 mm | 45 mm | 13 mm | 67 mm | 35 px | 0.21 |
| v13_stable | 2500 | 91 mm | 46 mm | 13 mm | 69 mm | 36 px | 0.25 |
| v13_stable | 2000 | 94 mm | 45 mm | 13 mm | 73 mm | 39 px | 0.18 |
| v13_stable | 1500 | 91 mm | 46 mm | 14 mm | 69 mm | 37 px | 0.22 |
| v13_stable | 1000 | 101 mm | 48 mm | 13 mm | 79 mm | 42 px | 0.25 |
| v13_stable | 500 | 105 mm | 55 mm | 15 mm | 77 mm | 43 px | 0.24 |
| v12_gtmask (previous champ) | 3000 | 116 mm | 78 mm | 14 mm | 72 mm | 34 px | 0.27 |
| v12_gtmask | 2500 | 116 mm | 80 mm | 14 mm | 72 mm | 36 px | 0.28 |
| v12_gtmask | 1500 | 116 mm | 80 mm | 14 mm | 72 mm | 38 px | 0.29 |
| v12_gtmask | 2000 | 119 mm | 80 mm | 14 mm | 73 mm | 36 px | 0.27 |
| v12_gtmask | 1000 | 119 mm | 82 mm | 14 mm | 72 mm | 38 px | 0.31 |
| v12_gtmask | 500 | 122 mm | 82 mm | 14 mm | 75 mm | 41 px | 0.31 |
| v11_pushpix (previous leader) | 3000 | 149 mm | 87 mm | 14 mm | 103 mm | 93 px | 0.28 |
| v11_pushpix (best abs+wrist) | 2000 | 149 mm | 89 mm | 14 mm | 100 mm | 94 px | 0.31 |
| v11_pushpix | 2500 | 152 mm | 89 mm | 14 mm | 107 mm | 96 px | 0.30 |
| v11_pushpix | 500 | 154 mm | 87 mm | 14 mm | 115 mm | 103 px | 0.30 |
| v11_pushpix | 1000 | 157 mm | 90 mm | 14 mm | 108 mm | 99 px | 0.33 |
| v11_pushpix | 1500 | 164 mm | 88 mm | 14 mm | 129 mm | 98 px | 0.31 |
| v10_scale (FINAL) 🏁 | 3000 | 159 mm | 89 mm | 15 mm | 121 mm | 105 px | 0.24 |
| v10_scale | 2500 | 159 mm | 91 mm | 15 mm | 123 mm | 106 px | 0.32 |
| v10_scale | 2000 | 160 mm | 93 mm | 15 mm | 117 mm | 107 px | 0.37 |
| v10_scale | 1500 | 183 mm | 96 mm | 15 mm | 149 mm | 110 px | 0.45 |
| v10_scale | 1000 | 165 mm | 96 mm | 15 mm | 131 mm | 115 px | 0.42 |
| v10_scale | 500 | 189 mm | 87 mm | 15 mm | 162 mm | 120 px | 0.60 |
| v10_scale | 2000 | ||||||
| v10_scale | 2500 | ||||||
| v10_scale | 3000 (final) | ||||||
| v8d step 3000 | 3000 | 211 mm | 102 mm | 13 mm | 110 mm | — | — |
| v9b step 200 | 200 | 222 mm | 92 mm | 13 mm | 130 mm | 327 px | — |
| v9a step 200 | 200 | 226 mm | 94 mm | 14 mm | 132 mm | 335 px | — |
Sorted by absolute MPJPE ascending. Oracle rows are theoretical bounds achievable with GT injection at inference; no model achieves them.
Metrics
Full definitions in benchmark/protocols/metrics.py.
Hand pose
abs_MPJPE(mm) — mean per-joint 3D distance in camera frame.rr_MPJPE(mm) — root-relative (wrist-subtracted): isolates orient + shape.PA_MPJPE(mm) — Procrustes-aligned: pure hand shape.wrist_MPJPE(mm) —‖P_wrist - Q_wrist‖: isolates trans error.PCK@0.10— % joints within 10% of hand bbox diagonal.pixel_err(px) —‖(u_pred, v_pred) - (u_gt, v_gt)‖at output res.
Depth / pointmap
AbsRel— mean(|d_pred - d_gt| / d_gt).δ<1.25— fraction of pixels within 1.25× ratio.SSI_log— scale-shift-invariant log-depth L1.
Camera
ATE(m),RPE_trans(m),RPE_rot(deg).
System-level
present_acc— binary accuracy on presence classification.vis_AP— per-joint visibility average precision.
Protocols
Full spec in benchmark/protocols/protocol.py.
| ID | Manifest | Split | Source | N clips | Purpose |
|---|---|---|---|---|---|
arctic-lite |
all7.jsonl | v8d all7 split (val) | arctic | 40 | Fast iter |
arctic-standard |
all7.jsonl | same | arctic | 150 | Paper-level |
all7-stratified |
all7.jsonl | same | (all 7) | 25/src × 7 = 175 | Full-pool |
Adding a model
- Register in
benchmark/scripts/run_eval.py'sMODEL_REGISTRY. - Run:
bash python benchmark/scripts/run_eval.py \ --model v10_scale_step1000 \ --protocol arctic-lite \ --out benchmark/results/v10_scale_step1000.arctic-lite.json - Rebuild leaderboard:
bash python benchmark/scripts/leaderboard.py - Commit
benchmark/results/*.json+benchmark/leaderboard/*→ row appears.
Reference points (external / oracle)
| Reference | abs MPJPE | Source |
|---|---|---|
| Iter 206 oracle (GT-2D + per-clip s*) | 54 mm | Iter 206 diagnostic, ARCTIC N=150 |
| Iter 204 oracle-orient + oracle-trans | 17 mm | Iter 204 diagnostic, ARCTIC N=150 |
| PA-MPJPE floor (shape only) | 13 mm | Iter 201 diagnostic |
| SomantisModel v14d (2D pixel) | 13.4 px | External champion (published) |
Meta
- Repository:
benchmark/ - Auto-regenerated: from
benchmark/results/*.jsonviabenchmark/scripts/leaderboard.py - Contributions: PRs adding new model results welcome; must include one JSON per (model × protocol) run.