Skip to content

Training recipe

Setting Value
Optimiser AdamW, β = (0.9, 0.999)
Weight decay 0.05 (none on norms and biases)
Gradient clipping global norm 1.0
LR: decoder + heads 3e-4
LR: top encoder block 6e-5, decayed by llrd per block below it
Schedule 5 % linear warmup from 1 % of the LR, then cosine to 1 %
Weight EMA decay 0.9995, warm-up min(d, (1+n)/(10+n)); best.pt stores EMA weights
Crops per epoch 12,000 (36,000 on the final Modal run)
Epochs 40 by default, and a wall-clock cap (max_minutes)

The schedule is driven by stage progress = max(epoch fraction, wall-clock fraction), not by a step count. If a run hits its time cap, it still finishes its anneal instead of stopping with the LR high. It also cannot fall out of sync with the real step count, which happened with v2’s precomputed OneCycleLR.

Learning rate over a run with default settings, for a 26-epoch run with 2 frozen epochs. The encoder LR is zero while it is frozen. v5 also ramps the encoder in linearly over one epoch after unfreezing (not shown).Source: Model_Traning/v5/train.py (lr_scale), config.py
Hardware Precision Micro-batch × accum Notes
H100 80 GB bf16 autocast 16 × 2 (v3, final) or 24 × 1 (v4-modal) no gradient checkpointing
2 × T4 16 GB (Kaggle) fp16 + GradScaler global 32 via DDP encoder gradient checkpointing, about 1.8–1.9 img/s

The H100 batch size was sized from the v3 run’s own measurements. Micro-batch 16 used 42 GB at 44.7 img/s, and micro-batch 32 used 74 GB at 48.0 img/s (7 % faster, and it ran out of memory at epoch 7). The GPU was compute-bound at 99–100 % utilisation either way.

Run GPU Epochs Wall-clock Throughput
v1 Kaggle 2 × T4 12 28 min —
v2 H100 80 GB (Lightning AI)* 20 of 30 (crashed) — —
v3 H100 (Lightning AI), bf16 26 153 min total, peak VRAM 80,345 MiB —
v4-Kaggle 2 × T4, fp16 12 of 40 (time limit) 480 min 1.8–1.9 img/s
v4-2 2 × T4, fp16 13 of 16 (time limit) 480 min 1.8–1.9 img/s
v4-modal H100 (Modal), bf16 30 119 min training —
DAv2 ablation fp16, batch 6 × 2, tile 518 40 674 min —
v5 final H100 (Modal), bf16 7 (best of 7) 106 min training 40.7 img/s, 46.7 GB peak

*The paper lists v2 on an RTX PRO 6000. The run’s own run.log records an H100 80 GB.

Best selection score
2.481m
forested + sparse RMSE on NEON & MVS3DM val
MVS3DM val
1.545m
start checkpoint: 1.615 m
GPU utilisation
88.5%
mean over 459 samples
Training time
106min
two sessions, full-state resume
  1. Warm start from resume-v4-1.6/best.pt, the Kaggle v5 run that first added MVS3DM.
  2. Data: MVS3DM, NEON, GAMUS, US3D and SynRS3D g05/g1. DFC23 and India were dropped because their labels put trees at 0 m.
  3. All 24 encoder blocks trainable from step 0, with the encoder LR ramped in over epoch 1. LRs are half the defaults (1.5e-4 / 3e-5), llrd 0.90.
  4. Sharpness: w_grad 1.0, w_normal 0.5. The balancer stays at β 0.5 / c 5.
  5. Selection: best.pt is chosen on forested + sparse RMSE over the NEON and MVS3DM validation sets (select_on), not on global RMSE, which urban ground dominates.
  6. Budget extended from 90 to 106 min mid-run. The cosine was held at 71 % on resume instead of restarting.

Per-epoch results are on the Benchmarks page. The run is not yet evaluated on held-out test splits.