EgoWM — Grounded egocentric hand + camera + depth on Wan 2.1

Fine-tuning a video foundation model (Wan 2.1-1.3B) with metric-grounded heads for hand pose, camera pose, and depth on 33k egocentric hand-object interaction clips.

Latest (2026-07-13): v36_sgdr13 step 3000 = 61.00 mm abs / rr 35.9 / PA 10.55 / wrist 43.4 / pix 20.81 / depth 0.151 on ARCTIC-lite (N=40). On 150-clip: 60.04 mm (first sub-61!). Thirteenth stacked SGDR delivered −0.62 mm on 40-clip / −1.09 mm on 150-clip. Wrist broke 44 mm floor. Notebook 132→66.0 (−50% — first below 50% of baseline!), s05 121→70.2 (−42%). v13 → v36: −32% abs on 150-clip. Gap to Iter 206 oracle: 7.0 mm on 40-clip / 6.0 mm on 150-clip.


Showcase

Training convergence — v10_scale on ARCTIC

Convergence

v10_scale converged from 189 mm at step 500 to 159 mm at step 3000. Depth head continued improving through the end (SSI residual 0.60 → 0.24). Path to the Iter 206 oracle bound (54 mm) still requires either a bigger model or smarter architecture; the current 159 mm is a 28% improvement over v9b and 25% over v8d at 1/8 the training compute.

Full training arc — all 6 metrics × 6 ckpts

Training arc

Six panels showing every metric across the v10_scale checkpoint series (step 500 → 3000), with v9b baseline (red dashes), v8d step 3000 (gray dots), and reference oracles (dashed lines) for context. Best v10 ckpt marked with gold star.

Multi-model comparison — v8d → v9b → v10 → v11 → v12 → v13 on same clips (6 columns)

Three ARCTIC val clips, six models predicting on the same frames. Columns left → right: v8d step 3000 (yellow) · v9b step 200 (orange) · v10 step 3000 (red-orange) · v11 step 3000 (dark red) · v12_gtmask step 3000 (purple) · v13_stable step 3000 (teal, FINAL 90mm). Per-column MPJPE ticker updates each frame. GT skeleton overlaid in green/cyan on all columns. The training arc is visible left-to-right: v8d predictions are stiffest (mean-collapsed), v9b begins to track, v10 refines, v11 refines further after pixel-focused finetune, v12 becomes tightest after the training-time GT off-frame mislabel fix, and v13 crisply follows GT after rotation-matrix loss + trunk freeze killed the axis-angle mean-collapse.

Earlier versions (for reference)

5-column (v8d → v12)
4-column (v8d → v11)
3-column (v8d → v10)
[**🎞️ v33 CHAMPION gallery (62.91 mm) →**](gallery_v33/gallery.html) · [**🎞️ v13 gallery →**](gallery_v13/gallery.html) · [**🎞️ v12 gallery →**](gallery_v12/gallery.html) · [**🎞️ v11 gallery →**](gallery_v11/gallery.html) · [**🎞️ v10 12-clip gallery →**](gallery/gallery.html)

Each card shows one clip with a 2×2 grid video: 1. Raw RGB (input) 2. GT + v10 skeleton overlay (green/cyan GT, red/salmon pred) 3. Predicted depth heatmap (turbo colormap on pmap Z channel) 4. 3D camera path + right-hand wrist trajectory (rotating view)

Per-clip metrics (abs MPJPE, rr, PA, pixel err) shown below each video. Best single clip is s08 capsulemachine_use_01 · abs 106 mm · pixel 27 px (essentially 2× oracle territory on the good clips).


Reports

Date Report One-liner
2026-07-12 📖 v22→v33 SGDR arc 74.9→62.91 mm push (−12 mm) via 10 stacked SGDR cycles. Warm-start compounding with cosine warm restarts. Wrist broke 7 whole-mm floors (54→44 mm). Gap to Iter 206 oracle now 7.9 mm on 150-clip.
2026-07-12 📖 v17→v21 arc summary 78.9→75.0 push via layered adaptive weights. Big win: v18 layered per-clip + keyword boost. Nulls: v19 re-scoring, v20 low-LR, temporal smoothing. h2o cross-source improved -11 mm.
2026-07-11 📖 v13→v16 arc summary 12 mm push via diagnostic-driven training. What worked: v15d rebalance, continued specialist training. What didn't: v14 PnP, v15c 6D-rot, v15a multi-res.
2026-07-11 🥇 v13 failure by category Notebook cluster + s05 dominate the tail — informed v15d/v16 rebalance strategy.
2026-07-11 🔍 v13 gap diagnostic Where do the remaining 36 mm live? Per-clip distribution now compact (median 85, min 29, no bimodality); worst clips are "notebook" static-hand typing; h2o PA regressed 22→28 (shape overfit); v14 direction: Somantis-style 2D + s* + PnP lift.
2026-07-11 🥇 v13 final report v13_stable s3000 = 90/45/13/67/35/0.21. Beats v12 champion by −26 mm abs, −33 mm rr (−42%).
2026-07-10 🚧 v13 diagnostic (initial) Why v12's 116 mm was unacceptable: axis-angle MSE mean-collapse, trunk drift, HO3Dv3 pollution, val NaN.
2026-07-10 🥇 v12 final report v12_gtmask s3000 = 116 mm / 34 px / 0.27. Training-time fix compounds better than eval-time filter. Superseded by v13.
2026-07-10 🏆 Live leaderboard Long-term multi-model comparison. v12 step 3000 currently leads at 116 mm / 34 px.
2026-07-10 🎞️ Gallery (12 clips) Per-clip 4-panel videos with metric annotations.
2026-07-10 📖 12-hour arc summary Full 12h autonomous exploration recap: v11 launch → per-source → best/worst → depth → GT filter → v12 retrain. Champion: v11 + GT filter at 124 mm / 39 px.
2026-07-10 🏆 GT off-frame filter finding 25 mm FREE improvement (149 → 124 mm) by filtering frames where GT wrist projects outside camera view. Pixel error 93 → 39 px (58%!). No retraining.
2026-07-10 🌊 Depth quality analysis v11 pred depth vs DA3 GT: AbsRel 0.33, δ<1.25 0.17 across 4 clips. Depth is uniform across best/worst clips — failure mode of worst clips is not depth quality.
2026-07-10 🔬 Failure mode analysis Per-clip metrics on 150 clips: BEST clips at oracle (64 mm / 14 px), WORST at total failure (500 mm / 2761 px). Median performance ~130 mm — mean dragged by 3% tail.
2026-07-10 🧭 Iter 209 Direction (post-v10) What v10 achieved + Gap analysis + v11-v15 backlog with expected impact per recipe.
2026-07-10 v10 Scale Training log Full-scale 7-head training on 33k pool with per-head GT masking.
2026-07-10 v10 Overfit Sanity Test 2-clip overfit: every loss ≥60% drop. Pipeline confirmed.
2026-07-10 3-way Benchmark Fair v8d/v9a/v9b eval that motivated Iter 208.
2026-07-10 Iter 208 Blueprint Architecture + parameter-reuse plan reusing VGGT's 32.65 M point_head.

Headline numbers (ARCTIC-lite protocol, N=40 clips)

Model Steps abs MPJPE ↓ rr MPJPE PA MPJPE wrist trans pixel err depth log-res
v36_sgdr13 (LEADER) 🥇 3000 61.00 mm 35.9 mm 10.55 mm 43.4 mm 20.81 px 0.151
v35_sgdr12 3000 61.62 mm 36.0 mm 10.59 mm 44.0 mm 20.82 px 0.160
v34_sgdr11 3000 62.61 mm 36.5 mm 10.70 mm 44.6 mm 21.9 px 0.155
v33_sgdr10 3000 62.91 mm 36.9 mm 10.72 mm 44.5 mm 21.7 px 0.152
v32_sgdr9 3000 63.70 mm 37.2 mm 10.83 mm 44.8 mm 21.8 px 0.151
v31_sgdr8 3000 64.67 mm 37.5 mm 10.9 mm 45.4 mm 22.7 px 0.156
v30_sgdr7 3000 65.26 mm 37.6 mm 10.95 mm 45.9 mm 22.8 px 0.152
v29_sgdr6 3000 66.00 mm 37.8 mm 11.0 mm 46.4 mm 22.9 px 0.149
v28_sgdr5 3000 66.04 mm 38.0 mm 11.1 mm 46.3 mm 23.4 px 0.150
v27_sgdr4 3000 67.18 mm 38.4 mm 11.1 mm 47.2 mm 24.1 px 0.146
v27_sgdr4 2500 68.7 mm 38.6 mm 11.1 mm 48.4 mm 24.4 px 0.153
v26_sgdr3 3000 68.95 mm 38.8 mm 11.2 mm 48.7 mm 24.9 px 0.148
v26_sgdr3 3500 70.18 mm 38.5 mm 11.1 mm 50.2 mm 24.7 px 0.156
v26_sgdr3 500 70.30 mm 39.1 mm 11.1 mm 50.2 mm 25.3 px 0.163
v25_sgdr2 3000 70.67 mm 39.3 mm 11.2 mm 50.2 mm 25.7 px 0.152
v25_sgdr2 3500 70.91 mm 38.9 mm 11.2 mm 50.6 mm 25.2 px 0.160
v24_sgdr 3500 72.02 mm 39.3 mm 11.2 mm 51.6 mm 26.4 px 0.162
v24_sgdr 3000 72.6 mm 39.6 mm 11.3 mm 51.9 mm 26.8 px 0.156
v24_sgdr 500 73.3 mm 40.4 mm 11.4 mm 52.6 mm 26.7 px 0.168
v22_bigscore 2500 74.89 mm 40 mm 11.4 mm 54 mm 27 px 0.162
v22_bigscore 2000 75.2 mm 40 mm 11.4 mm 55 mm 27 px 0.161
v18_adaptive 2500 75.0 mm 40 mm 11.2 mm 55 mm 28 px 0.174
v21_h2o (h2o boost) 1000 75.6 mm 40 mm 11.3 mm 55 mm 28 px 0.176
v18_adaptive 1500 75.4 mm 41 mm 11.4 mm 55 mm 29 px 0.178
v18_adaptive 1000 75.3 mm 40 mm 11.3 mm 56 mm 29 px 0.184
v18_adaptive 2000 76.2 mm 40 mm 11.3 mm 56 mm 28 px 0.170
v18_adaptive 500 81 mm 41 mm 11.5 mm 61 mm 29 px 0.172
v17_cosine 1500 78.3 mm 40 mm 11.3 mm 59 mm 28 px 0.159
v17_cosine 1000 78.6 mm 41 mm 11.4 mm 59 mm 30 px 0.164
v17_cosine 500 82 mm 40 mm 11.3 mm 63 mm 30 px 0.172
v16_more_boost 1000 78.9 mm 41 mm 11.5 mm 60 mm 31 px 0.161
v16_more_boost 500 79 mm 41 mm 11 mm 60 mm 33 px 0.18
v15c_rot6d (null) 500 79 mm 40 mm 11 mm 60 mm 32 px 0.18
v15d_rebalance 1000 80.5 mm 41 mm 11.5 mm 61 mm 32 px 0.19
v15d_rebalance 1500 80.7 mm 41 mm 11 mm 62 mm 31 px 0.18
v15d_rebalance 500 82 mm 40 mm 12 mm 62 mm 33 px 0.19
v15a_multires (v13 head) 1000 83 mm 42 mm 12 mm 64 mm 36 px 0.17
v15a_multires (v13 head) 3000 86 mm 42 mm 12 mm 64 mm 31 px 0.18
v13_stable (FINAL) 3000 90 mm 45 mm 13 mm 67 mm 35 px 0.21
v13_stable 2500 91 mm 46 mm 13 mm 69 mm 36 px 0.25
v13_stable 2000 94 mm 45 mm 13 mm 73 mm 39 px 0.18
v13_stable 1500 91 mm 46 mm 14 mm 69 mm 37 px 0.22
v13_stable 1000 101 mm 48 mm 13 mm 79 mm 42 px 0.25
v13_stable 500 105 mm 55 mm 15 mm 77 mm 43 px 0.24
v12_gtmask (previous champ) 3000 116 mm 78 mm 14 mm 72 mm 34 px 0.27
v12_gtmask 2500 116 mm 80 mm 14 mm 72 mm 36 px 0.28
v12_gtmask 1500 116 mm 80 mm 14 mm 72 mm 38 px 0.29
v12_gtmask 2000 119 mm 80 mm 14 mm 73 mm 36 px 0.27
v12_gtmask 1000 119 mm 82 mm 14 mm 72 mm 38 px 0.31
v12_gtmask 500 122 mm 82 mm 14 mm 75 mm 41 px 0.31
v11_pushpix (previous) 3000 149 mm 87 mm 14 mm 103 mm 93 px 0.28
v11_pushpix (best wrist) 2000 149 mm 89 mm 14 mm 100 mm 94 px 0.31
v11_pushpix 2500 152 mm 89 mm 14 mm 107 mm 96 px 0.30
v11_pushpix 1500 164 mm 88 mm 14 mm 129 mm 98 px 0.31
v11_pushpix 1000 157 mm 90 mm 14 mm 108 mm 99 px 0.33
v11_pushpix 500 154 mm 87 mm 14 mm 115 mm 103 px 0.30
v10_scale (FINAL) 🏁 3000 159 mm 89 mm 15 mm 121 mm 105 px 0.24
v10_scale 2500 159 91 15 123 106 0.32
v10_scale 2000 160 93 15 117 107 0.37
v10_scale 1500 183 96 15 149 110 0.45
v10_scale 1000 165 96 15 131 115 0.42
v10_scale 500 189 87 15 162 120 0.60
v8d step 3000 (baseline) 3000 211 102 13 110
v9b step 200 (baseline) 200 222 92 13 130 327
v9a step 200 (baseline) 200 226 94 14 132 335
oracle Iter 206 (GT-2D + s*) oracle 54 0
oracle Iter 204 (GT-orient + GT-trans) oracle 17 13
PA-MPJPE floor (shape only) oracle 13

Sorted with best model (lowest absolute MPJPE) at top of each group.


What's in this repo

EgoWM/
├── egowm/              — Wan-DiT trunk + MANO + heads (production ~2400 LoC)
├── data_engine/        — trainers, evals, viz, manifests
├── benchmark/          — long-term leaderboard (LVP-style)
│   ├── protocols/      — metrics + protocol registry
│   ├── results/        — one JSON per (model, protocol) run
│   ├── scripts/        — leaderboard, gallery, arc builders
│   └── gallery/        — rendered videos + posters
├── report/             — public deployable (this folder)
│   ├── README.md       — the page you are reading
│   ├── iters/*.md      — per-iteration reports
│   ├── gallery/        — same assets as benchmark/gallery/
│   ├── build.py        — markdown → HTML
│   └── *.html          — pre-built for Cloudflare Workers
└── docs/               — internal engineering log (private, not deployed)

How to add a model to the leaderboard

  1. Train your model somewhere (preserving the head interface — see egowm/models/ and data_engine/train_egowm_*.py).
  2. Run the shared eval: bash python data_engine/eval_v10_scale.py \ --ckpt path/to/your.pt \ --n-clips 40 --source-filter arctic \ --out benchmark/results/your_model.arctic-lite.json
  3. Tag with model, step, protocol fields in the JSON.
  4. Rebuild: bash python benchmark/scripts/leaderboard.py # leaderboard python benchmark/scripts/build_gallery.py --ckpt path/to/your.pt \ --model-name "Your Model" # gallery python benchmark/scripts/build_training_arc.py # training arc
  5. Commit benchmark/results/*.json, benchmark/gallery/*, and the built HTML.

Contributor

Reports and code authored during 2026-06-25 → 2026-07-10 by Haoran Geng (@geng-haoran) with Claude Opus 4.7 (1M context) as pair-programmer.