Training recipe
Optimisation
Section titled “Optimisation”| Setting | Value |
|---|---|
| Optimiser | AdamW, β = (0.9, 0.999) |
| Weight decay | 0.05 (none on norms and biases) |
| Gradient clipping | global norm 1.0 |
| LR: decoder + heads | 3e-4 |
| LR: top encoder block | 6e-5, decayed by llrd per block below it |
| Schedule | 5 % linear warmup from 1 % of the LR, then cosine to 1 % |
| Weight EMA | decay 0.9995, warm-up min(d, (1+n)/(10+n)); best.pt stores EMA weights |
| Crops per epoch | 12,000 (36,000 on the final Modal run) |
| Epochs | 40 by default, and a wall-clock cap (max_minutes) |
Learning-rate schedule
Section titled “Learning-rate schedule”The schedule is driven by stage progress = max(epoch fraction, wall-clock fraction), not by a step count. If a run hits its time cap, it still finishes its anneal instead of stopping with the LR high. It also cannot fall out of sync with the real step count, which happened with v2’s precomputed OneCycleLR.
Model_Traning/v5/train.py (lr_scale), config.pyPrecision and batch sizing
Section titled “Precision and batch sizing”| Hardware | Precision | Micro-batch × accum | Notes |
|---|---|---|---|
| H100 80 GB | bf16 autocast | 16 × 2 (v3, final) or 24 × 1 (v4-modal) | no gradient checkpointing |
| 2 × T4 16 GB (Kaggle) | fp16 + GradScaler | global 32 via DDP | encoder gradient checkpointing, about 1.8–1.9 img/s |
The H100 batch size was sized from the v3 run’s own measurements. Micro-batch 16 used 42 GB at 44.7 img/s, and micro-batch 32 used 74 GB at 48.0 img/s (7 % faster, and it ran out of memory at epoch 7). The GPU was compute-bound at 99–100 % utilisation either way.
Runs and hardware
Section titled “Runs and hardware”| Run | GPU | Epochs | Wall-clock | Throughput |
|---|---|---|---|---|
| v1 | Kaggle 2 × T4 | 12 | 28 min | — |
| v2 | H100 80 GB (Lightning AI)* | 20 of 30 (crashed) | — | — |
| v3 | H100 (Lightning AI), bf16 | 26 | 153 min total, peak VRAM 80,345 MiB | — |
| v4-Kaggle | 2 × T4, fp16 | 12 of 40 (time limit) | 480 min | 1.8–1.9 img/s |
| v4-2 | 2 × T4, fp16 | 13 of 16 (time limit) | 480 min | 1.8–1.9 img/s |
| v4-modal | H100 (Modal), bf16 | 30 | 119 min training | — |
| DAv2 ablation | fp16, batch 6 × 2, tile 518 | 40 | 674 min | — |
| v5 final | H100 (Modal), bf16 | 7 (best of 7) | 106 min training | 40.7 img/s, 46.7 GB peak |
*The paper lists v2 on an RTX PRO 6000. The run’s own run.log records an H100 80 GB.
The v5 final run
Section titled “The v5 final run”- Warm start from
resume-v4-1.6/best.pt, the Kaggle v5 run that first added MVS3DM. - Data: MVS3DM, NEON, GAMUS, US3D and SynRS3D g05/g1. DFC23 and India were dropped because their labels put trees at 0 m.
- All 24 encoder blocks trainable from step 0, with the encoder LR ramped in over epoch 1. LRs are half the defaults (1.5e-4 / 3e-5), llrd 0.90.
- Sharpness:
w_grad1.0,w_normal0.5. The balancer stays at β 0.5 / c 5. - Selection:
best.ptis chosen on forested + sparse RMSE over the NEON and MVS3DM validation sets (select_on), not on global RMSE, which urban ground dominates. - Budget extended from 90 to 106 min mid-run. The cosine was held at 71 % on resume instead of restarting.
Per-epoch results are on the Benchmarks page. The run is not yet evaluated on held-out test splits.