Skip to content

Design findings

These findings came from measuring failed or disappointing runs. Each one led to a specific change in a later version.

v2 zoomed random crops to simulate other GSDs, then padded them back to 512 px with zeros.

Samples padded
76.5%
> half padding
49.1%
Real image per sample
58.8%
on average
Supervised px affected
~41%

The labels under the padding were ground, so the network learned that featureless means flat. Fix (v3): choose the crop in source pixels so it never needs padding, and reflect-pad at inference.

On a tile from another sensor (Inria, Austin), v2 predicted a “crumpled mountain range”: only 28.3 % of pixels below 1 m, against 56.4 % in-domain (ground truth 58.2 %), with a median prediction of 4.04 m. The network had tied height to absolute brightness and contrast. Fix (v3): per-scene 2–98 % stretch and photometric jitter in training.

v2’s schedule planned to unfreeze the encoder, but the block lookup failed (cannot locate transformer blocks on the encoder) and the run crashed at epoch 21. Its 3.07 m came entirely from a frozen encoder. Fix (v3): find blocks structurally, probe which blocks receive gradients, and unfreeze all 24. This gave the largest single-version gain: 3.07 → 2.72 m.

v4 raised the stratum balancer’s β from 0.5 to 0.7 to help tall buildings. Once v4-modal’s errors were re-weighted onto v3’s validation height distribution, the like-for-like gap was 0.764 m (the headline gap was 1.026 m, but a quarter of that came from v4’s harder validation set). Almost all of it was flat ground:

Share of the v4-modal vs v3 MSE gap per height band, under v3's stratum mix. The 0–2 m band accounts for 85.5 %; v4 was slightly better on the two tallest bands.Source: Model_Traning/V4_modal/v3VSv4.md §3
Band v3 RMSE v4-modal RMSE v3 bias v4-modal bias
0–2 m 1.565 2.968 +0.415 +0.985
2–5 m 2.651 3.294 +0.572 +0.842
5–10 m 2.837 3.337 −0.359 −0.028
10–20 m 4.795 4.586 −1.375 −1.056
20+ m 5.638 5.553 −2.499 −2.348

Visually, v4 flattened tree canopy to zero height (Gallery). Fix (v5): revert β, the clip, soft bins, the entropy term and top-16 unfreezing together; then add forest data (NEON) and select checkpoints on forested and sparse error.

Both v3 and v4 logged val: 400 tiles from gamus, but from different packings: v3 from the gated HF snapshot (3837 train tiles), v4 from a Kaggle PNG mirror, taking the first 400 of 859. v4’s set had 2.1× the tall-pixel fraction. The distribution-free balanced RMSE agrees with the re-weighted comparison (3.497 vs 3.948 m). Fix (v5): a seeded random validation draw, with the protocol recorded in each run.

The validation split is also 87.5 % urban tiles, against 57.6 % in the test split. Validation numbers therefore favour urban performance.

The obvious way to make an absolute DSM is to add the predicted nDSM to a public DEM. But 30 m DEMs such as SRTM and Copernicus are surface models: they already contain buildings and canopy, blurred over each cell.

DEM + nDSM (naive)30 m DEM already contains the tower+ nDSM 30 m165.0 mtruth 135.0 mDTM + nDSM (DepthWizard)DTM fitted through ground pixels onlynDSM 30 m134.4 mtruth 135.0 mbuilding height counted twiceground-only terrain, building counted once
On a synthetic 30 m tower whose true top is at 135.0 m, DEM + nDSM gives 165.0 m. Fitting a bare-earth DTM through ground pixels only, then adding the nDSM, gives 134.4 m.

Fix (v4, refined in v5): fit terrain through ground pixels only, or anchor so that each 30 m cell averages to the DEM (Absolute DSM).