v13 failure mode by category — notebook cluster dominates the tail

Date: 2026-07-11 · Data: benchmark/results/v13_stable_step3000.per_clip.json (150 arctic val clips) TL;DR: 11% of clips (notebook) contribute 40% of the mean-MPJPE overshoot. Both pixel and depth are worse there.

The v13 gap diagnostic noted that "4 of 5 worst clips are notebook (static-hand typing)". This report drills into the object- and subject-level pattern to inform v15+ direction.


Object category breakdown

category n mean abs median mean pix
notebook 16 132 mm 122 49 px
box 6 104 mm 94 58
phone 18 94 mm 85 31
waffleiron 22 93 mm 96 34
laptop 9 89 mm 80 34
mixer 17 84 mm 81 36
capsulemachine 10 84 mm 88 33
ketchup 10 82 mm 76 39
espressomachine 11 75 mm 64 29
microwave 11 64 mm 53 27
scissors 11 60 mm 55 32

Notebook is a striking outlier — 58% worse abs, 44% worse pixel than the non-notebook median of 84 mm / 34 px. Microwave and scissors sit at the bottom, near or below the Iter 206 oracle (54 mm).

Why notebooks are hard: - Hand rests near-static on flat surface (minimal motion cue) - Fingers close together and mostly out-of-view (partial occlusion by table) - Common with subject s05 which is also the worst-performing subject

Subject breakdown

subject n mean abs
s05 21 124 mm
s04 13 101 mm
s02 13 91 mm
s10 14 88 mm
s09 25 85 mm
s07 13 81 mm
s08 27 75 mm
s06 15 70 mm

s05 is 78% worse than s06. Notebook + s05 together (~30 clips out of 150) account for most of the tail. This suggests either: - s05's hand shape / anatomy differs from the beta-optimized MANO (fits worse) - s05's scenes have more of the hard flat-surface interactions (notebooks) - Both

Attribution: pixel vs. depth in notebook failures

Given the leverage abs_mpjpe ≈ f(pixel_err, depth_err): - If notebook pixel were at non-notebook mean (34 vs 49) → notebook abs ≈ 108 mm - Remaining ~24 mm gap (108 → 132) attributable to depth / pose worse-than-baseline

So notebook's ~48 mm excess splits ~50/50 between pixel and depth. Both matter.

Implications for v15 direction

Pixel-side (v15a multi-res root head): 30% pixel improvement (35→24 px overall) would knock 15-20 mm off notebook AND 10-12 mm off non-notebook. Overall mean: 89 → ~74 mm.

Depth-side: no head-side pure-pixel or pure-rotation fix helps notebook's depth ambiguity. Needs either: - Temporal aggregation (adjacent frames average out static-hand pointmap noise) - Better shape supervision (s05 fits better with tuned betas) - Hard-example mining: upweight notebook + s05 in training

Quick fix candidate: 3× notebook + s05 sampling in training. 10 LoC change, 1 h retrain. Should close half the s05/notebook gap.

Actionable order: 1. v15a (running now) — pixel-focused, helps everyone including notebook. Expected 74-80 mm. 2. v15d — data rebalancing: 3× sample weight on notebook + s05 clips. Try after v15a. 3. v15c (6D-rot on global_orient) — helps rr, not pixel. Independent of category signal. 4. Long-term: temporal / clip-level modeling for static-hand scenarios.


Meta