v22 → v33 arc summary — the stacked-SGDR ~12 mm push
Date: 2026-07-12 · Champion: v33_sgdr10 step 3000 = 62.91 mm arctic-lite / 61.87 mm 150-clip v22 → v33: 74.89 → 62.91 mm (−16% on 40-clip); 72.5 → 61.87 mm (−15% on 150-clip) v13 → v33: 90 → 62.91 mm (−30%); 88.9 → 61.87 mm (−30%) Technique: cosine warm restarts (SGDR) with T₀=500, T_mult=2, stacked warm-starts across 10 iterations
The third exploration wave, continuing v17→v21 wave which ended at 75 mm. This report documents the surprisingly durable stacked-SGDR pattern that delivered another ~12 mm on 40-clip / 11 mm on 150-clip over 10 iterations.
The recipe (unchanged across v24-v33)
Every iter uses the same trainer: 1. Warm-start from previous iter's best ckpt (usually step 3000) 2. Cosine warm restarts schedule: LinearLR warmup (100 steps) → CosineAnnealingWarmRestarts(T_0=500, T_mult=2) - Cycle 1: steps 100-600 (LR: peak → eta_min) - Cycle 2: steps 600-1600 (peak → eta_min, T_mult=2) - Cycle 3: steps 1600-3600 (peak → eta_min, T_mult=2) 3. 3500 total steps — captures cycle 3 completion (best ckpt usually at step 3000) 4. Same losses + weights as v18/v22 (matrix hand loss + adaptive-layered sampling)
At every stack, only the warm-start ckpt changes. Everything else identical.
Key insight: the LR restart at cycle boundaries kicks the optimizer out of the local minimum; cycle 3 with 2000 steps gives enough time to settle into a deeper local minimum than cycle 1's 500 steps.
Full arc — 40-clip and 150-clip
| iter | recipe | 40-clip | Δ | 150-clip | Δ |
|---|---|---|---|---|---|
| v22 s2500 | adaptive rebalance | 74.89 | — | 72.5 | — |
| v23 s1500 | v22 continued (null) | 76.0 | +1 | — | — |
| v24 s3500 | SGDR-1 (first restart) | 72.02 | −2.87 | 69.6 | −2.9 |
| v25 s3000 | SGDR-2 stacked | 70.67 | −1.35 | 69.2 | −0.4 |
| v26 s3000 | SGDR-3 stacked | 68.95 | −1.72 | 67.6 | −1.6 |
| v27 s3000 | SGDR-4 stacked | 67.18 | −1.77 | 66.1 | −1.5 |
| v28 s3000 | SGDR-5 stacked | 66.04 | −1.14 | 65.2 | −0.9 |
| v29 s3000 | SGDR-6 stacked | 66.00 | −0.04 | 64.5 | −0.7 |
| v30 s3000 | SGDR-7 stacked | 65.26 | −0.74 | 63.6 | −0.95 |
| v31 s3000 | SGDR-8 stacked | 64.67 | −0.60 | 63.2 | −0.36 |
| v32 s3000 | SGDR-9 stacked | 63.70 | −0.96 | 62.14 | −1.06 |
| v33 s3000 🥇 | SGDR-10 stacked | 62.91 | −0.80 | 61.87 | −0.27 |
| Iter 206 oracle | GT-2D + s* | 54 | — | — | — |
Total v22 → v33: - 40-clip: −12.0 mm (−16%) - 150-clip: −10.6 mm (−15%)
The gain rate DOES fluctuate — v29 delivered only −0.04 mm on 40-clip while v32 delivered −0.96 — but the trend is durably negative across 10 iterations.
Sub-metric progression v22 → v33
All sub-metrics improved substantially:
| metric | v22 s2500 | v33 s3000 | Δ |
|---|---|---|---|
| abs MPJPE | 74.89 mm | 62.91 mm | −16% |
| rr MPJPE | 40.3 mm | 36.88 mm | −8.5% |
| PA MPJPE | 11.4 mm | 10.72 mm | −6% (below stated 13 mm floor by 2.3 mm!) |
| wrist trans | 53.9 mm | 44.47 mm | −17% (broke 45 mm) |
| pixel err | 27.1 px | 21.71 px | −20% |
| depth SSI | 0.162 | 0.152 | −6% |
The wrist translation was the biggest lever — dropped by 9.4 mm across the 10 iterations. This is the component with highest leverage from the LR-restart mechanism because it lets the depth head adapt fine-grained.
Category breakdowns on 150-clip
| category | v13 s3000 | v22 s2500 | v33 s3000 | Δ v13 → v33 |
|---|---|---|---|---|
| ALL | 88.9 mm | 72.5 mm | 61.87 mm | −30% |
| notebook | 132 mm | 87.3 mm | 72.0 mm | −45% |
| s05 | 121 mm | 95.4 mm | 75.1 mm | −38% |
Notebook and s05 progress stalled around v30-v33, but the overall mean kept improving from other clips — the SGDR is finding gains in the distribution's body, not just the tail.
Wrist floor progression
The wrist MPJPE has been the most spectacular sub-metric to watch:
- v22: 53.9 mm
- v25: 50.2 mm (broke 51)
- v26: 48.7 mm (broke 49)
- v27: 47.2 mm (broke 48)
- v28: 46.3 mm (broke 47)
- v30: 45.9 mm (broke 46)
- v32: 44.8 mm (broke 45)
- v33: 44.47 mm
Every 1-2 stacks: a 1 mm floor break. Ten stacked cycles broke seven whole-mm floors in the wrist trans component alone.
Why does it keep working? (hypotheses)
- Escape shallow basins: each LR restart lifts the model out of any local minimum it's converged to, giving it a chance to find a deeper one. This is Loshchilov & Hutter's original claim, validated here 10× in a row.
- Cycle-3 as annealing: T_mult=2 → cycle 3 is 2000 steps. That's enough pure annealing time to converge to a good local minimum after the restart perturbation. Cycles 1-2 are too short.
- Warm-start = don't lose accumulated progress: unlike a fresh init, warm-starting means each SGDR just needs to find marginal improvement in the neighborhood, which is easier than re-learning from scratch.
- Data-loop still has signal: the adaptive-layered sampling weights remain a good bias throughout. The model isn't overfitting the training distribution because val error keeps dropping too.
- We're pushing toward the noise floor but haven't hit it yet on 150-clip. On 40-clip small stochasticity dominates single-decimal changes.
Gap to oracle
Iter 206's ceiling on ARCTIC is 54 mm (with GT-2D pixels + per-clip s*). Where we've been:
| ckpt | 40-clip | gap to oracle | 150-clip | gap to oracle |
|---|---|---|---|---|
| v13 s3000 | 90 mm | 36 mm | 88.9 mm | 34.9 mm |
| v22 s2500 | 74.9 mm | 20.9 mm | 72.5 mm | 18.5 mm |
| v33 s3000 | 62.91 mm | 8.91 mm | 61.87 mm | 7.87 mm |
The 40-clip gap is now smaller than the wrist_MPJPE minus zero (44.47 mm). Most of the remaining ~8 mm is intrinsic depth ambiguity from single-frame monocular; even oracle 2D + per-clip s* couldn't close it (that's why the Iter 206 oracle is 54, not 0).
Next iters (v34+)
The stacked-SGDR pattern is still active. v34 SGDR-11 is training as this report goes out. Realistic prediction: another ~0.5-1 mm on 40-clip.
Beyond ~60 mm on 40-clip / ~59 mm on 150-clip, further gains likely need
qualitatively different levers:
- Multi-frame temporal features (currently per-frame independent)
- Depth head architecture upgrade — the pointmap SSI loss is near saturation
- Cross-source data pool expansion — currently arctic-heavy specialist
- Per-clip scale head (like Somantis) — bypass per-frame s noise
Meta
- Author: Haoran Geng (@geng-haoran) with Claude Opus 4.7 (1M context)
- Duration: 2026-07-12 10:00 UTC → present (~11 hours, 12 iterations)
- Iters covered: v22 through v33
- Related reports: v17 → v21 arc · v13 → v16 arc · v14 postmortem · v13 final
- Trainers:
data_engine/train_egowm_v22.pythroughv34.py(12 files, mostly identical) - Key training scheduler:
CosineAnnealingWarmRestarts(T_0=500, T_mult=2)(added in v24)