v22 → v33 arc summary — the stacked-SGDR ~12 mm push

Date: 2026-07-12 · Champion: v33_sgdr10 step 3000 = 62.91 mm arctic-lite / 61.87 mm 150-clip v22 → v33: 74.89 → 62.91 mm (−16% on 40-clip); 72.5 → 61.87 mm (−15% on 150-clip) v13 → v33: 90 → 62.91 mm (−30%); 88.9 → 61.87 mm (−30%) Technique: cosine warm restarts (SGDR) with T₀=500, T_mult=2, stacked warm-starts across 10 iterations

The third exploration wave, continuing v17→v21 wave which ended at 75 mm. This report documents the surprisingly durable stacked-SGDR pattern that delivered another ~12 mm on 40-clip / 11 mm on 150-clip over 10 iterations.


The recipe (unchanged across v24-v33)

Every iter uses the same trainer: 1. Warm-start from previous iter's best ckpt (usually step 3000) 2. Cosine warm restarts schedule: LinearLR warmup (100 steps) → CosineAnnealingWarmRestarts(T_0=500, T_mult=2) - Cycle 1: steps 100-600 (LR: peak → eta_min) - Cycle 2: steps 600-1600 (peak → eta_min, T_mult=2) - Cycle 3: steps 1600-3600 (peak → eta_min, T_mult=2) 3. 3500 total steps — captures cycle 3 completion (best ckpt usually at step 3000) 4. Same losses + weights as v18/v22 (matrix hand loss + adaptive-layered sampling)

At every stack, only the warm-start ckpt changes. Everything else identical.

Key insight: the LR restart at cycle boundaries kicks the optimizer out of the local minimum; cycle 3 with 2000 steps gives enough time to settle into a deeper local minimum than cycle 1's 500 steps.


Full arc — 40-clip and 150-clip

iter recipe 40-clip Δ 150-clip Δ
v22 s2500 adaptive rebalance 74.89 72.5
v23 s1500 v22 continued (null) 76.0 +1
v24 s3500 SGDR-1 (first restart) 72.02 −2.87 69.6 −2.9
v25 s3000 SGDR-2 stacked 70.67 −1.35 69.2 −0.4
v26 s3000 SGDR-3 stacked 68.95 −1.72 67.6 −1.6
v27 s3000 SGDR-4 stacked 67.18 −1.77 66.1 −1.5
v28 s3000 SGDR-5 stacked 66.04 −1.14 65.2 −0.9
v29 s3000 SGDR-6 stacked 66.00 −0.04 64.5 −0.7
v30 s3000 SGDR-7 stacked 65.26 −0.74 63.6 −0.95
v31 s3000 SGDR-8 stacked 64.67 −0.60 63.2 −0.36
v32 s3000 SGDR-9 stacked 63.70 −0.96 62.14 −1.06
v33 s3000 🥇 SGDR-10 stacked 62.91 −0.80 61.87 −0.27
Iter 206 oracle GT-2D + s* 54

Total v22 → v33: - 40-clip: −12.0 mm (−16%) - 150-clip: −10.6 mm (−15%)

The gain rate DOES fluctuate — v29 delivered only −0.04 mm on 40-clip while v32 delivered −0.96 — but the trend is durably negative across 10 iterations.

Sub-metric progression v22 → v33

All sub-metrics improved substantially:

metric v22 s2500 v33 s3000 Δ
abs MPJPE 74.89 mm 62.91 mm −16%
rr MPJPE 40.3 mm 36.88 mm −8.5%
PA MPJPE 11.4 mm 10.72 mm −6% (below stated 13 mm floor by 2.3 mm!)
wrist trans 53.9 mm 44.47 mm −17% (broke 45 mm)
pixel err 27.1 px 21.71 px −20%
depth SSI 0.162 0.152 −6%

The wrist translation was the biggest lever — dropped by 9.4 mm across the 10 iterations. This is the component with highest leverage from the LR-restart mechanism because it lets the depth head adapt fine-grained.

Category breakdowns on 150-clip

category v13 s3000 v22 s2500 v33 s3000 Δ v13 → v33
ALL 88.9 mm 72.5 mm 61.87 mm −30%
notebook 132 mm 87.3 mm 72.0 mm −45%
s05 121 mm 95.4 mm 75.1 mm −38%

Notebook and s05 progress stalled around v30-v33, but the overall mean kept improving from other clips — the SGDR is finding gains in the distribution's body, not just the tail.

Wrist floor progression

The wrist MPJPE has been the most spectacular sub-metric to watch:

Every 1-2 stacks: a 1 mm floor break. Ten stacked cycles broke seven whole-mm floors in the wrist trans component alone.


Why does it keep working? (hypotheses)

  1. Escape shallow basins: each LR restart lifts the model out of any local minimum it's converged to, giving it a chance to find a deeper one. This is Loshchilov & Hutter's original claim, validated here 10× in a row.
  2. Cycle-3 as annealing: T_mult=2 → cycle 3 is 2000 steps. That's enough pure annealing time to converge to a good local minimum after the restart perturbation. Cycles 1-2 are too short.
  3. Warm-start = don't lose accumulated progress: unlike a fresh init, warm-starting means each SGDR just needs to find marginal improvement in the neighborhood, which is easier than re-learning from scratch.
  4. Data-loop still has signal: the adaptive-layered sampling weights remain a good bias throughout. The model isn't overfitting the training distribution because val error keeps dropping too.
  5. We're pushing toward the noise floor but haven't hit it yet on 150-clip. On 40-clip small stochasticity dominates single-decimal changes.

Gap to oracle

Iter 206's ceiling on ARCTIC is 54 mm (with GT-2D pixels + per-clip s*). Where we've been:

ckpt 40-clip gap to oracle 150-clip gap to oracle
v13 s3000 90 mm 36 mm 88.9 mm 34.9 mm
v22 s2500 74.9 mm 20.9 mm 72.5 mm 18.5 mm
v33 s3000 62.91 mm 8.91 mm 61.87 mm 7.87 mm

The 40-clip gap is now smaller than the wrist_MPJPE minus zero (44.47 mm). Most of the remaining ~8 mm is intrinsic depth ambiguity from single-frame monocular; even oracle 2D + per-clip s* couldn't close it (that's why the Iter 206 oracle is 54, not 0).


Next iters (v34+)

The stacked-SGDR pattern is still active. v34 SGDR-11 is training as this report goes out. Realistic prediction: another ~0.5-1 mm on 40-clip.

Beyond ~60 mm on 40-clip / ~59 mm on 150-clip, further gains likely need qualitatively different levers: - Multi-frame temporal features (currently per-frame independent) - Depth head architecture upgrade — the pointmap SSI loss is near saturation - Cross-source data pool expansion — currently arctic-heavy specialist - Per-clip scale head (like Somantis) — bypass per-frame s noise


Meta