Accuracy by landscape
The pooled RMSE, MAE and Pearson columns reproduce the epoch-7 single-pass metrics across MVS3DM, NEON, GAMUS and US3D. The GAMUS RMSE column uses the same epoch-7 checkpoint with D4 averaging and 1.5× input scaling on all 859 GAMUS validation tiles. Raw evidence for both evaluations is linked in the table caption.
The table
Section titled “The table”| Landscape | RMSE (m) ↓ | MAE (m) ↓ | Pearson r ↑ | GAMUS RMSE (m) ↓ | Pooled pixels | GAMUS tiles | Pooled from |
|---|---|---|---|---|---|---|---|
| Urban | 3.85 | 1.69 | 0.90 | 3.06 | 73.6 M | 465.00 | GAMUS, US3D |
| Sparse | 1.32 | 0.32 | 0.63 | 1.25 | 45.5 M | 99.00 | MVS3DM, NEON, GAMUS, US3D |
| Forested | 2.85 | 1.60 | 0.83 | 2.05 | 64.7 M | 295.00 | MVS3DM, NEON, GAMUS |
| Overall | 3.04 | 1.32 | 0.88 | 2.59 | 183.7 M | 859.00 | MVS3DM, NEON, GAMUS, US3D |
RMSE, MAE and r: epoch-7 single-pass pool from four validation sets. GAMUS RMSE: the same checkpoint, D4 + 1.5× input, all 859 GAMUS validation tiles. Source: Model_Traning/v5/outputs/v5/modal/metrics.json · Docs-Site/public/evidence/v5-gamus-d4/probe_metrics.json
The D4 setting lowers full GAMUS RMSE from 2.775 m to 2.590 m (6.64%) while taking 14.2× longer. It worsens RMSE for true heights ≥15 m (5.792 → 5.980 m). See full GAMUS inference comparison for raw metrics and tradeoffs. This is GAMUS validation centre-crop evaluation, not held-out test or absolute DSM scoring.
Where each row comes from
Section titled “Where each row comes from”Each validation tile gets a landscape label from its own ground-truth height map (see the landscape classifier). The four-set columns score all pixels of each landscape together and write n, RMSE, MAE and r per landscape into the epoch-7 metrics.json. Its per-set inputs are:
| Set | Landscape | Tiles | Pixels | RMSE | MAE | r |
|---|---|---|---|---|---|---|
| MVS3DM | sparse | 90.000 | 22,781,911 | 0.861 | 0.321 | 0.801 |
| MVS3DM | forested | 110.000 | 26,440,447 | 1.951 | 1.240 | 0.776 |
| MVS3DM | overall | — | 49,222,358 | 1.545 | 0.815 | 0.816 |
| NEON | sparse | 36.000 | 9,258,211 | 0.685 | 0.254 | 0.933 |
| NEON | forested | 82.000 | 20,916,686 | 4.112 | 2.528 | 0.937 |
| NEON | overall | — | 30,174,897 | 3.444 | 1.830 | 0.953 |
| GAMUS | urban | 111.000 | 28,854,305 | 3.263 | 1.797 | 0.921 |
| GAMUS | sparse | 23.000 | 5,862,495 | 3.146 | 0.800 | 0.331 |
| GAMUS | forested | 66.000 | 17,301,504 | 2.012 | 1.017 | 0.801 |
| GAMUS | overall | — | 52,018,304 | 2.893 | 1.425 | 0.910 |
| US3D | urban | 171.000 | 44,700,003 | 4.187 | 1.614 | 0.881 |
| US3D | sparse | 29.000 | 7,602,176 | 0.266 | 0.023 | -0.040 |
| US3D | overall | — | 52,302,179 | 3.872 | 1.383 | 0.884 |
history[6] in metrics.json, keys val (MVS3DM), val_neon, val_gamus and val_us3d. Overall rows are each set's global block. Source: Model_Traning/v5/outputs/v5/modal/metrics.json
Some landscapes are missing from some sets:
- Urban comes from GAMUS and US3D only. No MVS3DM validation tile is classed urban, and NEON is configured never to produce urban tiles (
landscape_no_urban_sources = neon), because its closed forest canopy was being misread as rooftops. - Forested has no US3D rows, because none of the 200 sampled US3D tiles is classed forested.
- Hilly is absent from the table on purpose. The model’s target is height above ground, so terrain is subtracted from every label and no tile can be classed hilly.
How the sets are pooled
Section titled “How the sets are pooled”Each set reports its own metrics, so the table combines them using the pixel counts:
| Column | Formula | Exact? |
|---|---|---|
| RMSE | √( Σ nᵢ · RMSEᵢ² / Σ nᵢ ) | Yes. Identical to RMSE over all pixels at once. |
| MAE | Σ nᵢ · MAEᵢ / Σ nᵢ | Yes. |
| Pearson r | Σ nᵢ · rᵢ / Σ nᵢ | No. A pixel-weighted mean of per-set r. |
Find it in the run log
Section titled “Find it in the run log”The last evaluation in run.log (lines 757–763) prints the same per-landscape RMSEs, rounded to 0.01 m. In these lines fore, spar and urba are forested, sparse and urban:
eval e7 RMSE=1.545 MAE=0.815 r=0.816 d1=0.645 bal=4.280 tall_bias=-6.11 flat_bias=+0.36 edge=2.281 grad=0.23 [fore=1.95 spar=0.86] eval e7 [neon] RMSE=3.444 MAE=1.830 r=0.953 d1=0.719 bal=3.641 tall_bias=-1.18 flat_bias=+0.43 edge=4.646 grad=0.66 [fore=4.11 spar=0.68] eval e7 [neon @2x pooled] RMSE=3.416 MAE=1.817 r=0.954 d1=0.720 bal=3.608 tall_bias=-1.18 flat_bias=+0.43 edge=4.184 grad=0.63 [fore=4.08 spar=0.67] eval e7 [gamus] RMSE=2.893 MAE=1.425 r=0.910 d1=0.654 bal=3.662 tall_bias=-1.89 flat_bias=+0.53 edge=3.771 grad=0.39 [fore=2.01 spar=3.15 urba=3.26] eval e7 [us3d] RMSE=3.872 MAE=1.383 r=0.884 d1=0.715 bal=6.150 tall_bias=-5.58 flat_bias=+0.27 edge=5.780 grad=0.43 [spar=0.27 urba=4.19] eval e7 select=2.481 m (forested+sparse RMSE, mean over neon,mvs3dm) new best 2.481 m -> best.pt- The unlabelled first line is MVS3DM, the selection set.
[neon @2x pooled]scores NEON again at half resolution. The table does not use it. It uses the full-resolution[neon]line (val_neoninmetrics.json).- The
[gamus]line is the original single-pass GAMUS baseline on 200 tiles. The GAMUS column above uses all 859 validation tiles with D4 + 1.5× input; exact numbers are in the linked probe metrics. - Epoch 7 is the last epoch, and it is also the best. The run hit its 106-minute wall-clock cap during epoch 7 and then evaluated
best.ptagain ([final] evaluating best.pt (epoch 7)). - The log does not print pixel counts. Those are in
metrics.json.
Reproduce it
Section titled “Reproduce it”- Download the two files from the run. They are the files the Modal job wrote to
depthwizard-results:/v5_final_forest/, unchanged:- metrics.json (867 KB). SHA-256
82d361b1a6caf7bf50e20cb485ea9fc57e3b6da52c29bb4b83140036e6a4f5fb - run.log (110 KB). SHA-256
e706a022bbae4ded1379e66160c2502bdc295166692e7ef648b81747a69d9ce2
- metrics.json (867 KB). SHA-256
- Run the script below with the epoch-7 training metrics and the GAMUS D4 probe file. It uses only the Python standard library:
python3 reproduce_landscape_table.py metrics.json probe_metrics.jsonbest.pt = epoch 7 (select score 2.481 m)
landscape pixels RMSE MAE r GAMUS setsurban 73,554,308 3.85 1.69 0.90 3.26 GAMUS, US3Dsparse 45,504,793 1.32 0.32 0.63 3.15 MVS3DM, NEON, GAMUS, US3Dforested 64,658,637 2.85 1.60 0.83 2.01 MVS3DM, NEON, GAMUSoverall 183,717,738 3.04 1.32 0.88 2.89 MVS3DM, NEON, GAMUS, US3D"""Rebuild the v5 "DSM accuracy by landscape" table from the run's metrics.json.
Run from Docs-Site/ (standard library only): python3 scripts/reproduce_landscape_table.py python3 scripts/reproduce_landscape_table.py path/to/metrics.json
The first file is the one the v5 final H100 run wrote on Modal. A copy is served atpublic/evidence/v5-final/metrics.json. An optional second file is the full GAMUSvalidation D4 + 1.5x confirmation for that same epoch-7 checkpoint. The pooledfour-set columns use the training run; the GAMUS-only column uses the optionalall-859-tile confirmation.
Pooling: each validation set reports n, RMSE, MAE and Pearson r per landscape. RMSE = sqrt( sum(n_i * rmse_i^2) / sum(n_i) ) exact: the same as RMSE over all pixels MAE = sum(n_i * mae_i) / sum(n_i) exact r = sum(n_i * r_i) / sum(n_i) approximate: metrics.json has no per-set means or variances, so the exact pooled r cannot be rebuilt"""from __future__ import annotations
import jsonimport mathimport sysfrom pathlib import Path
SETS = {"val": "MVS3DM", "val_neon": "NEON", "val_gamus": "GAMUS", "val_us3d": "US3D"}LANDSCAPES = ["urban", "sparse", "forested"]
def pool(blocks): n = sum(b["n"] for b in blocks) return { "n": n, "rmse": math.sqrt(sum(b["n"] * b["rmse_m"] ** 2 for b in blocks) / n), "mae": sum(b["n"] * b["mae_m"] for b in blocks) / n, "r": sum(b["n"] * b["pearson_r"] for b in blocks) / n, }
def best_epoch(metrics): return min(metrics["history"], key=lambda h: h["select_score"])
def landscape_table(metrics, gamus_readout=None): ep = best_epoch(metrics) per_set = [] # every input row, so the pooling can be checked by hand for key, name in SETS.items(): v = ep[key] for land in LANDSCAPES + ["overall"]: b = v["global"] if land == "overall" else v["per_landscape"].get(land) if b: per_set.append({"set": name, "landscape": land, "n": b["n"], "tiles": b.get("tiles"), "rmse": b["rmse_m"], "mae": b["mae_m"], "r": b["pearson_r"]}) rows = [] for land in LANDSCAPES + ["overall"]: blocks = [ep[k]["global"] if land == "overall" else ep[k]["per_landscape"][land] for k in SETS if land == "overall" or land in ep[k]["per_landscape"]] if gamus_readout is None: g = ep["val_gamus"]["global"] if land == "overall" else ep["val_gamus"]["per_landscape"][land] else: candidates = gamus_readout["sources"]["gamus"]["candidates"] selected = candidates["baseline_d4_zoom150"] g = selected["global"] if land == "overall" else selected["per_landscape"][land] rows.append({"landscape": land, "sets": [SETS[k] for k in SETS if land == "overall" or land in ep[k]["per_landscape"]], **pool(blocks), "gamus_rmse": g["rmse_m"], "gamus_tiles": (859 if land == "overall" else g.get("tiles")), "gamus_protocol": "D4 + 1.5x, all val tiles" if gamus_readout else "single pass, 200 val tiles"}) return {"epoch": ep["epoch"], "select_score": ep["select_score"], "rows": rows, "per_set": per_set}
if __name__ == "__main__": path = Path(sys.argv[1] if len(sys.argv) > 1 else "public/evidence/v5-final/metrics.json") readout_path = Path(sys.argv[2]) if len(sys.argv) > 2 else None readout = json.loads(readout_path.read_text()) if readout_path else None t = landscape_table(json.loads(path.read_text()), readout) print(f"best.pt = epoch {t['epoch']} (select score {t['select_score']:.3f} m)\n") print(f"{'landscape':<10}{'pixels':>13}{'RMSE':>8}{'MAE':>8}{'r':>8}{'GAMUS':>8}{'GAMUS tiles':>13} sets") for r in t["rows"]: print(f"{r['landscape']:<10}{r['n']:>13,}{r['rmse']:>8.2f}{r['mae']:>8.2f}" f"{r['r']:>8.2f}{r['gamus_rmse']:>8.2f}{r['gamus_tiles']:>13} {', '.join(r['sets'])}")scripts/extract_metrics.py imports the same function to build this page’s data, so the page and the script cannot disagree.