EgoWM Benchmark

Long-term leaderboard and evaluation protocol for grounded egocentric hand + camera + depth prediction on cached-latent video foundation models.

Follows the modular pattern from Large Video Planner (registered datasets + registered models + protocol + JSON results), adapted for the per-hand-per-frame metric-3D surface EgoWM optimizes.

Repository: benchmark/


Full training arc — v10 through v33

The absolute-MPJPE convergence across 20+ model iterations, stacked on a cumulative training-step axis:

Hero convergence — full arc

Every iter, every metric (6-panel: abs / rr / PA / wrist / pixel / depth across all iters on cumulative x-axis):

Full training arc — 6 panels

The gold star marks v33 s3000 (current SOTA: 62.91 mm abs).

Convergence, v10 → v12 (earlier zoom)

Focused view of the first arc through v12:

Hero convergence — v10-v12 zoom

Training arc — v10-v12 6 panel


Multi-model side-by-side video

Six-column comparison across model versions on 3 ARCTIC val clips (v8d, v9b, v10, v11, v12, v13 side-by-side; skeleton overlay + per-frame MPJPE ticker):

5-column (v8d → v12)
4-column (v8d → v11)
3-column (v8d → v10)

4-panel per-clip galleries

Each gallery: RGB / GT+pred skeleton / depth heatmap / 3D trajectory, per-clip video, per-clip metrics under each card.


Per-source generalization (v12 step 3000 FINAL vs v11, 25 clips/source)

Cross-source eval. Reveals the training-time hand_valid fix helps arctic massively but hurts ho3dv3 — confirming ho3dv3 GT is noisier than the model:

Source v11 s3000 abs v12 s3000 abs Δ v11 pix v12 pix Δ
arctic 🥇 161 mm 118 mm −27 % 88 px 36 px −59 %
h2o 151 mm 148 mm −2 % 80 px 78 px −3 %
ho3dv3 255 mm 269 mm +5 % 196 px 196 px 0 %
asmhand
dexycb
hot3d_aria
hot3d_quest3

Big win — arctic: 161 → 118 mm (−43 mm, −27%) and 88 → 36 px (−52 px, −59%). Training-time fix compounded huge on arctic where the mislabeled off-frame frames were concentrated.

Modest spillover — h2o: 151 → 148 mm. h2o still slightly better on pixel (78 vs arctic's 36 is on N=25 vs arctic's arctic-lite N=40 = 34 — apples-apples on N=25 arctic is 36 px). h2o benefit is smaller because it has fewer off-frame mislabels to begin with.

Regression — ho3dv3: 255 → 269 mm (+14 mm). Confirms Iter 204/206 flag: ho3dv3 GT labels are noisier than the model. Any improvement that better fits the ARCTIC-side signal makes the model less compatible with ho3dv3's noisy labels. Recommendation: drop ho3dv3 from the training pool.

Data engine debt (blocks 3 of 7 sources): - asmhand: Iter 200 fix for NaN global_orient labels masks ~100% of val frames - dexycb / hot3d_aria / hot3d_quest3: manifest builder OOB (some clips reference frame indices past zarr end)

Both bugs are known (docs/20_depth_fusion_analysis Iter 204/205), fix is straightforward but hasn't been prioritized. When fixed, we can add proper per-source rows to the leaderboard.


Live leaderboard — ARCTIC-lite

Protocol: 40 ARCTIC val clips from the v8d all7 split. Same seed across models → fair comparison.

Model Steps abs MPJPE ↓ rr MPJPE PA MPJPE wrist trans pixel err depth log-res
oracle-iter204 (GT-orient + GT-trans) oracle 17 mm 13 mm
oracle-iter206 (GT-2D + per-clip s*) oracle 54 mm 0 px
v13_stable (FINAL) 🥇 3000 90 mm 45 mm 13 mm 67 mm 35 px 0.21
v13_stable 2500 91 mm 46 mm 13 mm 69 mm 36 px 0.25
v13_stable 2000 94 mm 45 mm 13 mm 73 mm 39 px 0.18
v13_stable 1500 91 mm 46 mm 14 mm 69 mm 37 px 0.22
v13_stable 1000 101 mm 48 mm 13 mm 79 mm 42 px 0.25
v13_stable 500 105 mm 55 mm 15 mm 77 mm 43 px 0.24
v12_gtmask (previous champ) 3000 116 mm 78 mm 14 mm 72 mm 34 px 0.27
v12_gtmask 2500 116 mm 80 mm 14 mm 72 mm 36 px 0.28
v12_gtmask 1500 116 mm 80 mm 14 mm 72 mm 38 px 0.29
v12_gtmask 2000 119 mm 80 mm 14 mm 73 mm 36 px 0.27
v12_gtmask 1000 119 mm 82 mm 14 mm 72 mm 38 px 0.31
v12_gtmask 500 122 mm 82 mm 14 mm 75 mm 41 px 0.31
v11_pushpix (previous leader) 3000 149 mm 87 mm 14 mm 103 mm 93 px 0.28
v11_pushpix (best abs+wrist) 2000 149 mm 89 mm 14 mm 100 mm 94 px 0.31
v11_pushpix 2500 152 mm 89 mm 14 mm 107 mm 96 px 0.30
v11_pushpix 500 154 mm 87 mm 14 mm 115 mm 103 px 0.30
v11_pushpix 1000 157 mm 90 mm 14 mm 108 mm 99 px 0.33
v11_pushpix 1500 164 mm 88 mm 14 mm 129 mm 98 px 0.31
v10_scale (FINAL) 🏁 3000 159 mm 89 mm 15 mm 121 mm 105 px 0.24
v10_scale 2500 159 mm 91 mm 15 mm 123 mm 106 px 0.32
v10_scale 2000 160 mm 93 mm 15 mm 117 mm 107 px 0.37
v10_scale 1500 183 mm 96 mm 15 mm 149 mm 110 px 0.45
v10_scale 1000 165 mm 96 mm 15 mm 131 mm 115 px 0.42
v10_scale 500 189 mm 87 mm 15 mm 162 mm 120 px 0.60
v10_scale 2000
v10_scale 2500
v10_scale 3000 (final)
v8d step 3000 3000 211 mm 102 mm 13 mm 110 mm
v9b step 200 200 222 mm 92 mm 13 mm 130 mm 327 px
v9a step 200 200 226 mm 94 mm 14 mm 132 mm 335 px

Sorted by absolute MPJPE ascending. Oracle rows are theoretical bounds achievable with GT injection at inference; no model achieves them.


Metrics

Full definitions in benchmark/protocols/metrics.py.

Hand pose

Depth / pointmap

Camera

System-level


Protocols

Full spec in benchmark/protocols/protocol.py.

ID Manifest Split Source N clips Purpose
arctic-lite all7.jsonl v8d all7 split (val) arctic 40 Fast iter
arctic-standard all7.jsonl same arctic 150 Paper-level
all7-stratified all7.jsonl same (all 7) 25/src × 7 = 175 Full-pool

Adding a model

  1. Register in benchmark/scripts/run_eval.py's MODEL_REGISTRY.
  2. Run: bash python benchmark/scripts/run_eval.py \ --model v10_scale_step1000 \ --protocol arctic-lite \ --out benchmark/results/v10_scale_step1000.arctic-lite.json
  3. Rebuild leaderboard: bash python benchmark/scripts/leaderboard.py
  4. Commit benchmark/results/*.json + benchmark/leaderboard/* → row appears.

Reference points (external / oracle)

Reference abs MPJPE Source
Iter 206 oracle (GT-2D + per-clip s*) 54 mm Iter 206 diagnostic, ARCTIC N=150
Iter 204 oracle-orient + oracle-trans 17 mm Iter 204 diagnostic, ARCTIC N=150
PA-MPJPE floor (shape only) 13 mm Iter 201 diagnostic
SomantisModel v14d (2D pixel) 13.4 px External champion (published)

Meta