Loss functions
Source: Model_Traning/v5/models/losses.py. All terms are computed in fp32, even under mixed precision.
Total objective
Section titled “Total objective”The fused prediction carries the main weight. Heads A and B get auxiliary supervision so each stays a useful height estimate for the gate to choose between.
Regression term
Section titled “Regression term”Each height map h is scored against the label y with three parts:
| Term | Definition |
|---|---|
| Balanced L1 | L1 error, weighted per pixel by the stratum balancer (below) |
| SiLog | Scale-invariant log error with a +1 m shift so 0 m is defined: |
| Gradient | L1 on horizontal and vertical first differences, averaged over 4 scales. Keeps edges crisp |
Structural terms
Section titled “Structural terms”| Term | Definition | Purpose |
|---|---|---|
| Normal | 1 − cos between predicted and true surface normals, with slopes scaled by GSD | Keeps roofs planar and walls vertical |
| Flatness | |Laplacian of the prediction| where the true Laplacian is below 0.35 m, on ground / low-veg / water / road or unlabelled pixels | Stops noise on flat ground |
| Bin CE | Cross-entropy on Head B’s bin logits at half resolution. Targets are the nearest bin, or a Gaussian over neighbouring bins when bin_soft_sigma > 0 |
Direct supervision for the bin distribution |
| Bin entropy | Negative entropy of the batch-mean bin distribution | Stops every pixel collapsing into one bin |
| Segmentation | Cross-entropy, ignore index 7 | Trains Head C |
Weights by version
Section titled “Weights by version”| Setting | v3 | v4 | v5 profile | v5 final run |
|---|---|---|---|---|
w_grad |
0.5 | 0.5 | 0.5 | 1.0 |
w_normal |
0.3 | 0.3 | 0.3 | 0.5 |
bin_soft_sigma |
hard | 1.5 | 0.0 | 0.0 |
w_bin_entropy |
— | 0.02 | 0.0 | 0.0 |
| Stratum β / clip | 0.5 / 5 | 0.7 / 8 | 0.5 / 5 | 0.5 / 5 |
v4 changed five of these settings at once (β, clip, soft bins, entropy and top-16 unfreeze) and regressed on flat ground. v5 reverted them. The final run doubled the gradient and normal weights to fight over-smoothing (Design findings).
Stratum balancer
Section titled “Stratum balancer”Height labels are extremely long-tailed: 49–61 % of GAMUS validation pixels are below 2 m, depending on the split version. Plain L1 would spend almost all its gradient on flat ground and under-predict tall structures. The balancer re-weights pixels by how rare their height band is.
With strata s ∈ {0–2, 2–5, 5–10, 10–20, 20+ m} and running frequencies fs (EMA with momentum 0.98):
Weights are then rescaled to a mean of 1 over valid pixels. β controls how strongly rare bands are boosted, and c caps the ratio.
Model_Traning/V4_modal/v3VSv4.md §4.1Coarse-label and consistency terms (v5)
Section titled “Coarse-label and consistency terms (v5)”- Coarse labels. Some sources have labels coarser than their pixels, for example NEON’s 1 m LiDAR rasters on 0.5 m pixels. For sources listed in
coarse_label_sources, the L1, SiLog and gradient terms are computed after k × k average pooling (k = round(label size / GSD)), and those pixels skip the normal, flatness and bin terms. The part is mixed in by its share of valid pixels. An optional vegetation mask (excess-green index > 0.10 with label < 1 m) ignores trees that a label wrongly records at 0 m. - Mean-teacher consistency. The code has an L1 term against an EMA teacher’s prediction on unlabelled imagery, keeping only pixels where the teacher’s σ ≤ 1.5 m. No run has used it, because the unlabelled data store was never built.