v13 failure mode by category — notebook cluster dominates the tail
Date: 2026-07-11 · Data:
benchmark/results/v13_stable_step3000.per_clip.json(150 arctic val clips) TL;DR: 11% of clips (notebook) contribute 40% of the mean-MPJPE overshoot. Both pixel and depth are worse there.
The v13 gap diagnostic noted that "4 of 5 worst clips are notebook (static-hand typing)". This report drills into the object- and subject-level pattern to inform v15+ direction.
Object category breakdown
| category | n | mean abs | median | mean pix |
|---|---|---|---|---|
| notebook | 16 | 132 mm | 122 | 49 px |
| box | 6 | 104 mm | 94 | 58 |
| phone | 18 | 94 mm | 85 | 31 |
| waffleiron | 22 | 93 mm | 96 | 34 |
| laptop | 9 | 89 mm | 80 | 34 |
| mixer | 17 | 84 mm | 81 | 36 |
| capsulemachine | 10 | 84 mm | 88 | 33 |
| ketchup | 10 | 82 mm | 76 | 39 |
| espressomachine | 11 | 75 mm | 64 | 29 |
| microwave | 11 | 64 mm | 53 | 27 |
| scissors | 11 | 60 mm | 55 | 32 |
Notebook is a striking outlier — 58% worse abs, 44% worse pixel than the non-notebook median of 84 mm / 34 px. Microwave and scissors sit at the bottom, near or below the Iter 206 oracle (54 mm).
Why notebooks are hard: - Hand rests near-static on flat surface (minimal motion cue) - Fingers close together and mostly out-of-view (partial occlusion by table) - Common with subject s05 which is also the worst-performing subject
Subject breakdown
| subject | n | mean abs |
|---|---|---|
| s05 | 21 | 124 mm |
| s04 | 13 | 101 mm |
| s02 | 13 | 91 mm |
| s10 | 14 | 88 mm |
| s09 | 25 | 85 mm |
| s07 | 13 | 81 mm |
| s08 | 27 | 75 mm |
| s06 | 15 | 70 mm |
s05 is 78% worse than s06. Notebook + s05 together (~30 clips out of 150) account for most of the tail. This suggests either: - s05's hand shape / anatomy differs from the beta-optimized MANO (fits worse) - s05's scenes have more of the hard flat-surface interactions (notebooks) - Both
Attribution: pixel vs. depth in notebook failures
Given the leverage abs_mpjpe ≈ f(pixel_err, depth_err):
- If notebook pixel were at non-notebook mean (34 vs 49) → notebook abs ≈ 108 mm
- Remaining ~24 mm gap (108 → 132) attributable to depth / pose worse-than-baseline
So notebook's ~48 mm excess splits ~50/50 between pixel and depth. Both matter.
Implications for v15 direction
Pixel-side (v15a multi-res root head): 30% pixel improvement (35→24 px overall) would knock 15-20 mm off notebook AND 10-12 mm off non-notebook. Overall mean: 89 → ~74 mm.
Depth-side: no head-side pure-pixel or pure-rotation fix helps notebook's depth ambiguity. Needs either: - Temporal aggregation (adjacent frames average out static-hand pointmap noise) - Better shape supervision (s05 fits better with tuned betas) - Hard-example mining: upweight notebook + s05 in training
Quick fix candidate: 3× notebook + s05 sampling in training. 10 LoC change, 1 h retrain. Should close half the s05/notebook gap.
Actionable order: 1. v15a (running now) — pixel-focused, helps everyone including notebook. Expected 74-80 mm. 2. v15d — data rebalancing: 3× sample weight on notebook + s05 clips. Try after v15a. 3. v15c (6D-rot on global_orient) — helps rr, not pixel. Independent of category signal. 4. Long-term: temporal / clip-level modeling for static-hand scenarios.
Meta
- Author: Haoran Geng (@geng-haoran) with Claude Opus 4.7 (1M context)
- Data source:
benchmark/results/v13_stable_step3000.per_clip.json - Related: v13 gap diagnostic · v14 postmortem