Repository navigation
feat(horizon): fixed-range Gower normalization (+ explain site-name) - #400
Merged
Merged
Conversation
…loor) The per-slice horizon Gower distance normalized each numeric feature by the per-slice data spread (max-min across the candidates + pit at that depth), floored at 10% of a fixed plausible span (#377). That made a candidate's distance depend on which OTHER candidates were in the pool and vary by depth: e.g. a 16% vs 0% rock-fragment gap scored 0.727 at 0-20 cm (pool spread 22) but 0.258 at 20-30 cm (a Leptosol 26 km away pushed the spread to 62). Soil properties have known physical spans, so normalize each numeric feature by its fixed plausible range (GLOBAL_HORIZON_PROP_BOUNDS / global_prop_bounds) instead. A candidate's distance is now a stable, local function of (candidate, pit) — independent of the rest of the pool and constant across depths. The fixed span is now the full denominator, not just a floor. Change is isolated to gower_distances (denom = theoretical span when provided); US + global already pass their bounds in. Bound VALUES are unchanged for now (pending soil-scientist review). Result-affecting -> snapshots must be regenerated and validated on the bulk-test suite; version 2.6.0 -> 2.7.0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Follow-up to the fixed-range Gower change (065cb20): - render_explain.py: the horizon caption described the old "range = spread of the candidates at this depth, floored at 10% of the plausible span, varies with depth" behavior. It now normalizes by the fixed plausible span at every depth, so the caption is rewritten to say so (and to note the distance no longer depends on the other candidates in the pool). Display-only; the numbers already flow from gower's denom, and the HTML snapshot test only smoke-checks render. - test_gower_distance.py: the two range tests pinned the floor semantics. One keeps its assertions (comment fixed: denom = fixed span, not floored to 8); the other is replaced with an invariant test that the fixed-range distance is independent of the per-slice spread (the point of the change). No ranking change beyond 065cb20; version stays 2.7.0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
render_html now accepts an optional site_name; when provided it's shown (HTML-escaped) as the report title above the "Soil ID explanation — REGION" heading. Default None keeps existing output byte-identical, so snapshots and callers that don't pass a name are unaffected. The CLI wrapper passes the site/export label it already extracts, so local reports get the name too. (Unrelated to the fixed-range change on this branch; bundled here per request.) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
US + global test_soil_location rankings and the explain trace snapshot move under the fixed-range change (065cb20). Regenerated in the GDAL runner against the pinned soil-id-db image (CI-matching); determinism verified (175 passed twice). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
knipec
approved these changes
Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Normalize each numeric horizon feature in the per-slice Gower distance by its fixed plausible span (
GLOBAL_HORIZON_PROP_BOUNDS/global_prop_bounds) instead of the per-slice data spread. Supersedes the #377 "spread, floored at 10% of the span" hybrid — the fixed span is now the full denominator, not just a floor.Effect: a candidate's horizon distance is now a stable, local function of (candidate, pit) — it no longer depends on which other candidates are in the pool (a distant outlier can't rescale everyone) and no longer varies by depth.
Why
The old data-spread denominator meant the same |Δ| scored very differently by depth and pool. Real example (Humic Alisols, pit rock-fragment 0% vs candidate 16%): Δ 0.727 at 0–20 cm (pool spread 22) but 0.258 at 20–30 cm — where the 62 came from a Leptosol 26 km away. Soil properties have known physical ranges, so normalizing to them is the principled fix.
Bulk test (NAS, paired before/after)
Also in this PR
render_explain.py: horizon caption rewritten to describe fixed-range (the numbers already flow from gower'sdenom); optionalsite_namerenders (HTML-escaped) as the report title, and the CLI wrapper passes its label.test_gower_distance.py: floor tests replaced with the fixed-range invariant.Review notes
site_nametitle change is logically independent of fixed-range (bundled here by request) — easy to split if preferred.🤖 Generated with Claude Code