Skip to content

Benchmarks

These are the most trustworthy numbers: v4-modal on data never used for training or checkpoint selection.

GAMUS test · TTA
3.32m
2,861 tiles · MAE 1.51 m · r 0.89
DFC23 · out-of-domain
4.97m
SuperView-1 satellite · r 0.881
India · near-domain
2.77m
35 tiles · r 0.922
MVS3DM · 30 m cells
1.35m
v5 vs LiDAR, 8 held-out sites
Split (tiles)RMSEMAErδ1Balanced RMSE
GAMUS test · plain (2861)3.3731.5490.8860.6274.613
GAMUS test · TTA (2861)3.3211.5110.8900.6294.556
GAMUS test · sliding+TTA (400)3.8732.2570.9110.5644.118
DFC23 val · OOD (246)4.9681.7890.8810.7765.544
India val · near-domain (35)2.7710.8650.9220.8884.447

v4-modal on held-out and cross-domain data. Heights in metres. Source: Model_Traning/V4_modal/Output/metrics.json; Research-Paper/main.tex Table IV

GAMUS validation RMSE per epoch, single pass. v1 and v2 used an earlier validation subset; v4 runs use a harder 400-tile prefix; v5 a seeded random 400. The two Kaggle v4 runs stopped at the 8-hour session limit.Source: src/data/metrics.json ← run metrics.json histories; v1/v2 from Research-Paper/scripts/make_figures.py
Final GAMUS validation RMSE under the three protocols. TTA helps every run; sliding-window inference costs a little because it includes tile seams and scene edges.Source: src/data/metrics.json ← final_plain / final_tta / final_sliding_tta
RunRMSEMAErδ1Bal. RMSERMSE (TTA)RMSE (sliding)
v32.7151.3120.9170.6723.5842.6052.723
v4-Kaggle3.4271.8850.9210.6133.8693.3893.441
v4-23.8032.0650.9020.5954.1673.7903.804
v4-modal3.7542.0010.9050.6014.0603.6313.646
DAv23.5221.9040.9150.6193.9483.4323.394

GAMUS validation, 400 tiles. RMSE through Bal. RMSE are single-pass. Validation sets differ between v3 and the v4-era runs (see Design findings). Source: Model_Traning/{Logs/v3 (2), V4_Kaggle/outputs/*, V4_modal/Output, DAV2_V1/outputs/dav2-v1}/metrics.json

To remove validation-set differences, v3, v4-modal and the DAv2 ablation were scored with one protocol on the same held-out tiles (eval_test.py --compare).

Held-out RMSE under an identical protocol. v4-modal is best on GAMUS test with single-pass and TTA inference, and by far the best on India (v3 never trained on it). DAv2 edges ahead on sliding-window inference.Source: Model_Traning/v3/outputs/v3_vs_v4_vs_dav2.md
Split · protocolv3v4-modalDAv2
GAMUS test · plain3.6283.3733.424
GAMUS test · TTA3.5653.3213.396
GAMUS test · sliding+TTA3.8223.8733.759
India val · plain5.4562.7803.027

Source: Model_Traning/v3/outputs/v3_vs_v4_vs_dav2.md

v5 was trained as a chain of warm-started runs (lineage). None has yet been scored on the held-out test splits. Preliminary · not test-evaluated

Run Primary val RMSE Notes
with_dfc 2.988 m (GAMUS) DFC23 val 5.391 m
without_dfc 2.909 m (GAMUS)
without_dfc / resume 2.863 m (GAMUS)
warm_start / Resume (v5_probe_v4init) 2.804 m (GAMUS) DFC23 4.878 m, India 2.767 m; source of the ONNX export
resume-v2 2.888 m (GAMUS) adds US3D
resume-v3 2.865 m (GAMUS)
resume-v4-1.6 1.615 m (MVS3DM) GAMUS 2.864 m, US3D 3.964 m, gradient ratio 0.23
v5_final_forest 2.481 m selection score MVS3DM 1.545 m, NEON forest 4.11 m
v5 final run, per epoch, on validation sets. The selection score (forested + sparse RMSE on NEON and MVS3DM) fell from 2.717 to 2.481 m, and NEON forest RMSE from 4.61 to 4.11 m. GAMUS urban held within 0.04 m. GAMUS here uses 200 tiles at the corrected 0.25 m GSD, so it is not comparable with earlier GAMUS numbers.Source: Model_Traning/V4_modal/FINAL_RUN_RESULTS.md