Skip to content

Scoring tuning: raw-% rock fragment + site weight 0.25 + dis_max consistency (v2.4.0) - #397

Merged
johannesparty merged 7 commits into
mainfrom
feat/soilid-site-weight
Sep 25, 2026
Merged

johannesparty merged 7 commits into
mainfrom
feat/soilid-site-weight

Conversation

@johannesparty

Copy link
Copy Markdown
Contributor

Stacked on #396 (feat/explain-html-renderer). A batch of tuning-driven scoring
changes, each validated against the US bulk test on the NAS (see
SOILID_TUNING.md / JOHANNES_NOTES_2026.md for methodology + numbers).

Changes (net top-1, real US bulk runs)

Change US top-1 notes
baseline (PR 396) 61.0% binned rfv, site_wt 0.5
rock fragment on raw % (c9b168c) 61.56% (+0.56) candidate uses raw c_cfpct_intpl instead of the 5-bucket class midpoint — removes the class-boundary cliff + within-class flattening. Pit still a class midpoint (user only reports a class).
site weight 0.5 → 0.25 (8ac63bf) 61.90% (+0.34) slope/elev were over-weighted vs horizon
global dis_max nanmax → 1.0 (6b3c9c8) ~0 consistency + correctness (fixed, query-independent penalty, matches US path)

missing/recall unchanged throughout — these reorder candidates, they don't
change which soils are retrieved. Version bumped 2.2.0 → 2.4.0; unit snapshots
regenerated + verified deterministic against the pinned soil-id-db.

Not included / notes

  • A uniform-depth-weight change was tried and reverted (3fa7b0c): net ~0 on
    the bulk test (helps surface-heavy pits, slightly hurts deep pits) and
    de-weighting the disturbed topsoil is the more defensible model — so the surface
    weight stays at its original 0.2 (this cancels in the PR's net diff).
  • rfv/LAB-b horizon down-weights an earlier optimization suggested are not
    adopted — the raw-% change superseded them (re-optimizing on the raw-% base
    leaves them neutral).
  • Backend re-pin is intentionally NOT part of this — to be done after review.

🤖 Generated with Claude Code

@johannesparty
johannesparty force-pushed the feat/soilid-site-weight branch from 741d04f to 3d51128 Compare September 25, 2026 18:19
@johannesparty
johannesparty force-pushed the feat/explain-html-renderer branch 2 times, most recently from 21dcf66 to 1d7cb78 Compare September 25, 2026 18:59
Base automatically changed from feat/explain-html-renderer to main September 25, 2026 19:05
johannesparty and others added 7 commits September 25, 2026 12:05
…eight

Three low-risk scoring cleanups (see SOILID_TUNING.md §7 for the analysis):

- global dis_max: np.nanmax(dis_mat_list) -> 1.0, matching the US path. Gower
  distances are ~[0,1] so a fixed, interpretable, query-independent penalty is
  cleaner than the empirical max. Accuracy effect negligible (Δtop1 ≈ -0.04).
- depth weight: surface 0-20cm weight 0.2 -> 1.0 (uniform) in both US and global.
  Tuning showed the surface de-weighting slightly hurt accuracy (~+0.3pt top1);
  kept as a two-part concat so the surface band stays an easy knob.
- cross-reference comments on the duplicated horizon-property bounds tables
  (us_soil.global_prop_bounds <-> global_soil.GLOBAL_HORIZON_PROP_BOUNDS).

Unit-test snapshots regenerated against the pinned soil-id-db image and verified
deterministic.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…s midpoint

Both the US and global horizon comparison fed rock fragment into ranking as a
5-bucket class midpoint (0/8/25/48/80 via getCF), for both the pit and the
candidate. Quantizing the candidate created a cliff at every class boundary —
15% vs 16% scored as 8 vs 25 (near-total mismatch) while 16% vs 35% scored
identical — turning small real differences into large scored ones and erasing
within-class differences.

Use the candidate's RAW rock-fragment % (c_cfpct_intpl, already computed and
already what the displayed profile uses via rf_lyrs) instead of the binned
c_cfpct_intpl_grp, in all three candidate-assembly paths (global, US main, US
OSD merge). The pit stays a class midpoint (getCF_fromClass) since the user only
reports a class; that midpoint is on the same raw-% scale, so the feature bounds
and the #377 range floor are unchanged. Sand/clay are left binned (they derive
from the 2-D texture class, not a 1-D threshold set — separate question).

Result-affecting: horizon/US/global snapshots must be regenerated and the change
validated on the bulk-test suite. Version bumped 2.2.0 -> 2.3.0 so clients flush
cached matches.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
c9b168c switched candidate rock fragment from the binned class midpoint to the
raw %, which is result-affecting, but landed without regenerating the unit
snapshots. Regenerate + verify (deterministic) against the pinned soil-id-db.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The site (slope/elevation) score was over-weighted relative to the horizon score.
Tuning against the US bulk test and a confirmation real run give +0.34 pt top-1
(61.56 -> 61.90 on the raw-% base); top-3/top-5 flat, recall unchanged (pure
reorder — pulls correct soils from rank 3/4 up to rank 1). See SOILID_TUNING.md.

Result-affecting: version 2.3.0 -> 2.4.0; snapshots regenerated + verified.
Note: the rfv/LAB-b horizon down-weights an earlier optimization found are NOT
applied — they were compensating for the class-binning noise that the raw-% rock
fragment change (2.3.0) already fixed; re-optimizing on the raw-% base leaves them
neutral, so site_wt is the only surviving lever.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The uniform-depth change (in the consistency commit) was net ~0 on the raw-% base:
it helps surface-heavy pits (+~1 pt) but slightly hurts deep pits (-0.07), which
cancel. De-weighting the disturbed topsoil (0.2x) is also the more defensible model
for the common deep-data case, so restore it. dis_max=1.0 (the genuinely-more-
correct part of that commit) is kept.

The principled improvement would be data-presence-aware weighting (de-weight the
surface only when there's deeper data), not a constant -- filed under "Potential
future changes" in JOHANNES_NOTES_2026, low priority since the effect is ~0.

Snapshots regenerated + verified.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…CM 50 -> 35

Leptosols/lithosols/rendzinas/rankers are defined at ~25-30 cm to rock; 50 cm was
a provisional placeholder (#375). 35 cm is closer to the taxonomic definition with
a small margin. Result-affecting for pits whose effective bedrock is 35-50 cm and
have a shallow-soil candidate (one global snapshot moved: a shallow candidate's
horizon score is now demoted). Version 2.4.0 -> 2.5.0; snapshots regenerated + verified.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…cally (#398)

process_distance_scores sorted rows by cond_prob desc but then grouped with
sort=True, which re-ordered the groups alphabetically by compname before the
top-12 truncation — so high-probability soils were silently dropped in favor of
lower-probability ones whose names sort earlier (whenever a buffer had >12 groups).
Affects both US and global candidate selection. Fix: groupby(sort=False) so the
top-12 follows the cond_prob-desc order.

Also fixes the following no-op loop that rebound a local `group` (its distance sort
+ min_dist were discarded): write the transformed group back into the list. This
was harmless in practice (both are recomputed downstream) but is now correct.

Adds regression tests. Result-affecting (8 snapshots moved: 4 US + 4 global, all
locations with >12 groups); version 2.5.0 -> 2.6.0; snapshots regenerated + verified.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@johannesparty
johannesparty force-pushed the feat/soilid-site-weight branch from 3d51128 to 6fd3ddb Compare September 25, 2026 19:05
@johannesparty
johannesparty merged commit 7eb7cb9 into main Sep 25, 2026
8 checks passed
@johannesparty
johannesparty deleted the feat/soilid-site-weight branch September 25, 2026 19:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants