Skip to content

feat(api): serve the storage schema plan and apply at the write tier - #1394

Merged
aparajon merged 4 commits into
mainfrom
armand/storage-schema-api
Sep 16, 2026
Merged

aparajon merged 4 commits into
mainfrom
armand/storage-schema-api

Conversation

@aparajon

@aparajon aparajon commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

Why this matters

The storage schema surface exists one layer down as Tern RPCs, and nothing yet reaches it over HTTP. Exposing it is one decision with a real consequence attached: the tier an operator has to hold to read a plan of SchemaBot's own bookkeeping schema.

What it does

Two admin-only routes expose the storage schema surface over HTTP:

POST /api/storage/schema/plan
POST /api/storage/schema/apply

A server answers them from the adapter bound to the storage it booted with, and refuses rather than guessing when a build never resolved one.

Both are POSTs, and that is what carries the authorization gate. The plan reads and writes nothing, so the first instinct is a GET. But the gate that actually bites is the tier: scoped write checks are a pass-through on deployments that leave scoped writes disabled, so tier classification is the whole admin decision there. A GET would have been admitted at the read tier — which is to say SchemaBot's bookkeeping schema would have been readable by everyone holding read access.

  GET /api/storage/schema/plan          POST /api/storage/schema/plan
         │                                     │
         ▼                                     ▼
  TierForRequest: GET                   TierForRequest: non-GET
         │                                     │
         ▼                                     ▼
    read tier                             write tier
         │                                     │
         ▼                                     ▼
  ✗ anyone with read access             admin only, by the default
    can enumerate the internal            rule — no per-path
    bookkeeping schema                    exception to keep in sync

The POST also carries the schema files the plan compares against in its body, so the verb the gate wanted is the verb the payload wanted.

Invariants

  • AZ-2, upholds. The new endpoints are writes by the default rule, so they fail closed with no per-path exception to keep in sync, and the route authorization sweep test covers them. The entry's Enforced: line now also names pkg/auth/tiers.go, where request classification actually lives — the previous pointer resolved cleanly to a file that no longer decides it.

Opened by Claude Code (Opus 5).

@aparajon
aparajon added this pull request to stack #1397 September 11, 2026 20:26
Copilot AI lite review requested due to automatic review settings September 12, 2026 06:53
@aparajon
aparajon force-pushed the armand/storage-schema-api branch from b638831 to 4da6e99 Compare September 12, 2026 06:53

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The route contract and server-version wiring issues remain unresolved, with documentation and test updates also requested.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Adds admin-only HTTP endpoints for inspecting and applying SchemaBot’s storage schema through local or remote adapters, with write-tier authorization, tests, and documentation updates.

Changes:

  • Adds storage-schema request/response types and handlers.
  • Wires storage adapters and API routes.
  • Adds authorization, routing, and handler coverage.
  • Updates AZ-2 and authentication documentation.
File summaries
File Summary
pkg/serve/serve.go Registers the local adapter. Moderate (1 vote): pass the derived version fallback. Nit (1 vote): add Build-to-Service integration coverage.
pkg/auth/tiers.go Documents write-tier classification.
pkg/auth/tiers_test.go Tests storage routes as write-tier requests.
pkg/apitypes/storage_schema_requests.go Defines HTTP payloads.
pkg/api/storage_schema_handlers.go Implements validation, authorization, routing, and responses.
pkg/api/storage_schema_handlers_test.go Tests handler behavior. Nit (1 vote): correct the POST/write-tier test comment.
pkg/api/service.go Registers schema routes. Moderate (3 votes): reconcile the advertised /diff route with the registered /plan route.
pkg/api/route_authorization_sweep_test.go Adds authorization fixtures.
docs/invariants.md Updates AZ-2 enforcement references.
docs/auth.md Documents schema access. Nit (1 vote): update admin-only guidance.
Review details

Suppressed comments (4)

docs/auth.md:657

  • These handlers call authorizeDirectAdminWrite, so storage-schema inspection and convergence are admin-only when scoped authorization is enabled; a database operator group still receives 403. The existing admin-only list in docs/auth.md:759-761 was not updated, so this new table/paragraph can lead operators to grant a scoped group that cannot use the endpoints. Add these operations to that admin-only guidance.
`POST /api/storage/schema/plan` reads without changing anything, and still
requires write access under that default. It reports the internal shape of
SchemaBot's own bookkeeping database, and its sibling route converges that
database, so both belong to the people who operate the server rather than to
everyone who can see the schema changes it runs.

pkg/api/storage_schema_handlers_test.go:339

  • This test comment says the diff is a GET, but the route under test is POST; that reverses the reason the tier gate applies and can mislead future authorization changes. Describe it as being denied because the POST defaults to the write tier.
// database operator is denied on it even though it is a GET, because the tier
// rule admits it as a write — the storage database is SchemaBot's own

pkg/serve/serve.go:519

  • Build derives moduleVersion() for embedded hosts when WithBuildInfo is absent, but the adapter wired here still receives the empty o.version. Normal embedders will therefore return storage reports with an empty version and generic schema attribution, even though the report contract says it identifies the answering binary; pass the same derived fallback into Server.version.
		version:         o.version,

pkg/serve/serve.go:526

  • The production wiring added here is not exercised by the new handler tests: those tests all call SetStorageSchemaService manually. If this Build-to-Service registration is removed or regresses, a real server's local storage-schema routes will refuse with 400 while the handler suite remains green; add an integration assertion that a built server's HTTP handler reaches the server-bound storage adapter.
	svc.SetStorageSchemaService(srv.storageSchemaService())
  • Files reviewed: 10/10 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread pkg/api/service.go
@aparajon
aparajon force-pushed the armand/storage-schema-api branch from 4da6e99 to 7b82a94 Compare September 12, 2026 07:00
@aparajon aparajon changed the title feat(api): serve the storage schema diff and apply at the write tier feat(api): serve the storage schema plan and apply at the write tier Sep 12, 2026
@aparajon
aparajon marked this pull request as ready for review September 12, 2026 07:15
@aparajon
aparajon force-pushed the armand/storage-schema-api branch from 7b82a94 to aa44ec9 Compare September 12, 2026 17:29
@aparajon
aparajon removed this pull request from stack #1397 September 12, 2026 21:00
@aparajon
aparajon added this pull request to stack #1408 September 12, 2026 21:00
@aparajon
aparajon force-pushed the armand/storage-schema-api branch from aa44ec9 to 4c9f197 Compare September 12, 2026 21:41
@Kiran01bm

Copy link
Copy Markdown
Collaborator

🤖 Review findings - created by Kiran's code review agent - for schemabot/pull/1394, 4c9f197.
Verdict: 5 findings — none blocking; 4 non-blocking (metrics label, error mapping, leaked endpoint, docs), 1 suggestion.

Both finder lenses completed. Adversarial verification was capped at 8 candidates, so 1 lower-ranked candidate was never verified either way and is not reported here.

Non-blocking

The two new operation names are missing from the metrics allowlist, so every storage-schema authorization decision is labelled operation="unknown". metrics.go#L583 rewrites any operation absent from knownDirectWriteAuthOperations (metrics.go#L532-L551), and the PR never touches pkg/metrics. A denied apply is therefore indistinguishable from a denied plan or any other unlisted operation, so no alert can be written — and the comment at storage_schema_handlers.go#L283 plus metrics/README.md#L443 both claim otherwise. Nothing pins allowlist/handler parity, so green CI doesn't disprove it.

writeStorageSchemaFailure collapses codes.Unimplemented into a generic 500 "see the answering deployment's logs", discarding the upgrade guidance the client already produced. grpc_client.go#L588 returns "does not support storage schema reads; upgrade that data plane", but storage_schema_handlers.go#L258 only special-cases InvalidArgument and falls through. The operator is sent to read the logs of a data plane whose only problem is version skew the control plane already knows about; log_handlers.go#L311 shows the established UnsupportedCapability / "upgrade it and retry" arm. Same shape for codes.Unavailable.

A gRPC-client resolution failure is returned verbatim with HTTP 400, leaking the configured endpoint address. With tern_deployments.west.production set to a portless west.example, Endpoint() succeeds and tern.NewGRPCClient fails with "split host:port from address west.example", which storage_schema_handlers.go#L111 hands to s.writeError(w, http.StatusBadRequest, err.Error()). That blames the caller for a server misconfiguration and contradicts the redaction policy writeStorageSchemaFailure enforces 150 lines below — whose test asserts db.example never reaches the body. No test covers this path.

docs/auth.md's hand-maintained admin-only enumeration was not updated, so the docs imply a database-operator grant covers the new routes. auth.md#L759 still reads "changing settings, maintaining checks, redriving webhooks, and forcing a lock release" while the Write row at auth.md#L641 gained storage schema. A scoped operator reading this concludes POST /api/storage/schema/plan is theirs and gets a 403 — on the exact grant boundary this PR establishes.

General suggestions

The doc comment on TestStorageSchemaRoutes_DenyScopedOperator says the route "is a GET" when both routes it drives are POSTs — and it's a POST that makes the tier rule classify them as writes. storage_schema_handlers_test.go#L399 states the inverse of tiers.go#L61, where GET is unconditionally TierRead. Someone acting on the comment could "restore" the route to a GET and drop it to the read tier — the outcome the handler comment at storage_schema_handlers.go#L301 explicitly warns about.

The one thing that could have broken, verified

Whether the admin gate can be bypassed on either route. Both are registered POST-only in service.go#L850 and /api/storage/schema/* is absent from readPaths, so TierForRequest's default arm yields TierWrite; the handler-level authorizeDirectAdminWrite then runs before resolveStorageSchemaTarget in both handlers (L142/L194 vs L151/L197), and TestStorageSchemaRoutes_DenyScopedOperator asserts local.planReq/applyReq stay nil on denial. The gate holds at both layers.

Verified correct

  • storageSchemaReportResponse is nil-safe on both hops; a missing report becomes an explicit 500, not a nil dereference.
  • decodeOptionalStorageSchemaBody tolerates an empty body only via io.EOF; truncated JSON yields io.ErrUnexpectedEOF and still 400s.
  • DisallowUnknownFields on the apply body rejects a reused plan body carrying schema_files, so an operator can't silently converge to the embedded schema.
  • resolveStorageSchemaTarget checks TernDeployments.Endpoint before TernClient, preventing the defaultTernClient fallback from answering for a named deployment.
  • Bodies are size-limited by s.limitRequestBody, and writeBodyDecodeError maps *http.MaxBytesError to 413, so the unbounded schema_files map can't exhaust memory.
  • The 500 branch emits only the fixed summary; the cause is logged at the call site and the table test pins that db.example never reaches the body.

This review was generated by Claude Code (claude-opus-5).

@Kiran01bm Kiran01bm left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Approved on Kiran's (@kmuddukrishna) behalf by the scheduled review agent — no blocking findings at 4c9f197. See the review comment above; non-blocking findings and suggestions, if any, are not merge gates.

@aparajon
aparajon force-pushed the armand/storage-schema-api branch 2 times, most recently from f1ad009 to b722280 Compare September 14, 2026 19:05
@aparajon

Copy link
Copy Markdown
Collaborator Author

🤖 All five addressed in b7222801, and the branch is now rebased onto 7cb9fccc.

The two new operation names are missing from the metrics allowlist, so every decision is labelled operation="unknown". This was the sharpest of the five — a route whose denials cannot be alerted on is a gate nobody can watch. Both names are in knownDirectWriteAuthOperations now, and the reason nothing caught it is itself fixed: a registry test pins allowlist/handler parity, in the shape the recovered-panic registry already uses, so the next route that forgets fails CI instead of review.

writeStorageSchemaFailure collapses codes.Unimplemented into a generic 500. Split, using the UnsupportedCapability arm in log_handlers.go as the precedent. Unimplemented now says the deployment needs upgrading and Unavailable says it needs looking at; both are remedies without reading anyone's logs.

A gRPC-client resolution failure is returned verbatim with HTTP 400, leaking the configured endpoint address. Fixed. It answers with the fact that a client could not be built for the configured endpoint, and the cause — which names this deployment's own infrastructure — stays in the server's logs. You were right that this contradicted the redaction policy the failure writer enforces 150 lines below; the table test now covers this path too.

docs/auth.md's admin-only enumeration was not updated. Updated, so the prose and the Write row agree.

The doc comment says the route "is a GET" when both routes are POSTs. Reworded. A comment that invites someone to "restore" the route to a GET and drop it to the read tier is worse than no comment.


Replied by Claude Code (claude-opus-5) on Armand's behalf.

@Kiran01bm

Copy link
Copy Markdown
Collaborator

🤖 Review findings - created by Kiran's code review agent - for schemabot/pull/1394, b722280.
Verdict: 4 findings — 3 non-blocking (status taxonomy, dead test guard, misleading 501 text), 1 suggestion (routes ship inert).

Non-blocking

Server-side resolveStorageSchemaTarget failures are answered 400, contradicting this file's own taxonomy. storage_schema_handlers.go:161 writes every resolve error as StatusBadRequest, including line 117 ("its endpoint is configured but the client could not be built") and "this server does not expose its own storage schema" — both server conditions the caller cannot correct. A portless tern_deployments.west.production: "west.example" passes config validation, fails in NewGRPCClient, and returns 400; the identical unreachable-data-plane condition at line 279 returns 503 engine_unavailable. A CLI branching on 4xx ("your input is wrong, don't retry") versus 5xx takes the wrong branch, and the local no-service case answers 400 where the remote one answers 501.

The new registry guard names a gate function that does not exist, so three live call sites go unscanned. direct_write_auth_registry_test.go:39 keys on authorizeDirectWriteForPlan (the real gate is authorizeDirectWriteForStoredPlan) and omits authorizeDirectDatabaseWrite entirely. Re-running the test's own AST scan shows "apply" (plan_handlers.go:950), "lock_acquire" (lock_handlers.go:81) and "lock_release" (:159) are never scanned, while the other gates keep require.NotEmpty green. Add s.authorizeDirectDatabaseWrite(w, r, "lock_steal", …) without registering it and every decision records operation="unknown" — exactly what the doc comment says this test prevents.

The 501 body asserts one cause for Unimplemented and drops the other. writeStorageSchemaFailure discards errStorageSchemaUnsupported's message — which names both causes — for the hardcoded "it is running a release that predates them, so upgrade that deployment and retry" at storage_schema_handlers.go:277. A data plane on today's release whose embedder never called tern.WithStorageSchemaService hits the same path, so the operator upgrades an already-current deployment and gets the identical 501. Forwarding the underlying message (or naming the wiring cause) makes the response actionable.

General suggestions

Both routes ship registered but inert in any binary built from this head. No non-test code calls SetStorageSchemaService or tern.WithStorageSchemaService — serve.go:674 registers the data plane without it — yet service.go:855 registers POST /api/storage/schema/plan. An operator on this commit gets 400 locally and 501 "upgrade that deployment" remotely, neither of which names the missing wiring. Worth landing the embedder wiring alongside, or gating route registration on the service being set.

The one thing that could have broken, verified

Making storage schema reachable over HTTP could have handed a scoped operator a write path into the storage database. Both handlers run authorizeStorageSchemaOperation → authorizeDirectAdminWrite before touching target.service, both routes are non-GET and absent from auth.readPaths so TierForRequest returns TierWrite, and the new fixtures in route_authorization_sweep_test.go make the scoped-operator 403 an enumerated property of the route table rather than a per-handler habit. TestStorageSchemaRoutes_DenyScopedOperator asserts local.planReq/applyReq stay nil, proving a denied request neither reads nor converges storage.

Verified correct

  • resolveStorageSchemaTarget checks TernDeployments.Endpoint before TernClient and rejects a non-tern.StorageSchemaService client, so a named data plane can never be answered by control-plane storage.
  • storageSchemaService is written under Lock and read under RLock via localStorageSchemaService; nothing else in the package touches the field.
  • resp.GetReport()/GetPlanned()/GetRemaining() are nil-safe generated accessors and StorageSchemaReportFromProto(nil).APIType() returns nil, so the nil-report 500 branches are reachable and cannot panic.
  • Request bodies are bounded by the pre-existing 32 MiB limitRequestBody middleware, and writeBodyDecodeError maps *http.MaxBytesError to 413, so decoding before authorization cannot amplify memory.
  • decodeOptionalStorageSchemaBody treats io.EOF as an empty body and keeps DisallowUnknownFields, so {"deploymnet":"west"} is a 400 rather than a silent convergence of the wrong storage.
  • The new tiers_test cases pin both routes to TierWrite, and the middleware admits via the same function under test.

This review was generated by Claude Code (claude-opus-5).

@Kiran01bm Kiran01bm left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Approved on Kiran's (@kmuddukrishna) behalf by the scheduled review agent — no blocking findings at b722280. See the review comment above; non-blocking findings and suggestions, if any, are not merge gates.

@Kiran01bm

Copy link
Copy Markdown
Collaborator

🤖 Review findings - created by Kiran's code review agent - for schemabot/pull/1394, 6f2ed42.

Verdict: 5 findings — 2 blocking (routes inert, 30 s write timeout), 2 non-blocking, 1 suggestion.

Blocking

Nothing ever supplies a tern.StorageSchemaService, so both new routes are inert in every server built from this repo. pkg/serve/serve.go#L674 is tern.NewServer(client, s.logger).Register(gs) with the variadic opts ...ServerOption empty, and RegisterGRPC offers no option to pass one; SetStorageSchemaService / WithStorageSchemaService have no callers outside _test.go. The local route therefore always answers 400 "this server does not expose its own storage schema", and the remote route always returns Unimplemented → 501 "upgrade that deployment" against a deployment already on the newest release. No e2e or integration test covers the HTTP→adapter→database path — every new test injects a fake.

The new routes never lift the server's 30 s WriteTimeout, but the convergence they invoke runs on a 5-minute budget. pkg/api/storage_schema_handlers.go#L218 calls StorageSchemaApply with no SetWriteDeadline, while EnsureSchema builds its own EnsureSchemaTimeout = 5 min context and serve.go sets WriteTimeout: 30 * time.Second; every other long-running admin handler (handleChecksScan, handleChecksSynthesize, handleWebhookRedrive) calls extendWebhookOpsDeadline precisely for this. At t=30 s the write deadline fires: the convergence runs to completion server-side but the JSON body is never delivered, so the operator cannot tell a finished convergence from one that left statements behind — which the handler's own comment calls "the only thing this answer is for". The plan route sits on the same boundary, since StorageSchemaPlanTimeout is exactly 30 s.

Non-blocking

The registry-coverage test names a gate function that does not exist and omits one that does, silently skipping three operation call sites. pkg/metrics/direct_write_auth_registry_test.go#L39 lists "authorizeDirectWriteForPlan", whose only occurrence in the repo is that line — the real gate is authorizeDirectWriteForStoredPlan — and authorizeDirectDatabaseWrite is absent entirely. So "apply" (plan_handlers.go:950), "lock_acquire" (lock_handlers.go:81) and "lock_release" (:159) are never scanned; rename any of those literals and the operation starts recording as unknown on the dashboard while this test stays green.

Every resolveStorageSchemaTarget failure returns 400, including the two the code's own comments call the server's own fault. pkg/api/storage_schema_handlers.go#L161 answers StatusBadRequest for a TernClient construction failure ("the configured endpoint … a dial failure") and for a control plane built without a storage schema service — neither fixable by changing the request. Ten lines below, writeStorageSchemaFailure maps the same class of cause to 500/503/501, and the PR's own test asserts StatusServiceUnavailable for an unreachable data plane, so a CLI that splits 4xx from 5xx reports a broken server config as the operator's malformed request.

General suggestions

Both handlers decode the body before authorizing, unlike their three sibling admin handlers. pkg/api/storage_schema_handlers.go#L143 runs decodeStorageSchemaPlanRequest (with DisallowUnknownFields) before authorizeStorageSchemaOperation, so a scoped operator posting {"deploymnet":"west"} gets 400 json: unknown field instead of the 403 TestStorageSchemaRoutes_DenyScopedOperator asserts for the same principal with {}. handleChecksSynthesize, handleWebhookRedrive, handleChecksScan and handleChecksRepos all authorize as their first statement. No capability leaks — the denial just is not uniform, and the non-empty-body case is untested.

The one thing that could have broken, verified

Exposing two new mutating routes without the scoped-operator denial the repo's auth doctrine requires. Verified sound: readPaths is unchanged and contains only /api/pull, so both POSTs fall to default: return TierWrite; TestMutatingRoutesDenyScopedOperatorByDefault walks svc.apiRoutes() and requires a fixture for every write-tier pattern; and because authorizeStorageSchemaOperation runs before resolveStorageSchemaTarget, the observed 403 is a real authorization denial (the fake's planReq/applyReq stay nil) rather than an incidental 400 from an unconfigured target.

Verified correct

  • Tier classification: the only tiers.go change is a doc comment; new table cases assert the same TierWrite.
  • Nil-report handling: StorageSchemaReportFromProto returns a typed nil and APIType guards r == nil, so the 500 paths cannot panic.
  • In-process-client refusal is real: only *GRPCClient asserts StorageSchemaService, so a locally-configured deployment name hits the type-assertion refusal.
  • Unimplemented → 501 survives fmt.Errorf wrapping: grpc-go's status.Code routes through FromError→errors.As.
  • Caller attribution upholds AZ-3: resolveCaller returns the verified subject, so req.Caller cannot override the audited identity.
  • Request bodies are bounded: limitRequestBody wraps every ConfigureRoutes route at 32 MB, and overflow maps to 413.
  • Empty-body handling is strict and correct: io.EOF yields the zero value while DisallowUnknownFields stays on, pinned by TestHandleStorageSchemaApply_RefusesUnknownField.

This review was generated by Claude Code (claude-opus-5).

@aparajon
aparajon force-pushed the armand/storage-schema-api branch from 6f2ed42 to 3839be0 Compare September 15, 2026 22:50
@Kiran01bm

Copy link
Copy Markdown
Collaborator

🤖 Review findings - created by Kiran's code review agent - for schemabot/pull/1394, 3839be0.
Verdict: 5 findings — 3 non-blocking (guard under-covers, 4xx for server faults, no deadline extension), 2 suggestions.

Non-blocking

The new registry guard names a gate function that does not exist, so three operations are never checked. direct_write_auth_registry_test.go:39 lists authorizeDirectWriteForPlan (the real name is authorizeDirectWriteForStoredPlan) and omits authorizeDirectDatabaseWrite, so apply, lock_acquire and lock_release never reach the scan. Replicating the scan verbatim collects only 10 operations and require.NotEmpty keeps it green. Add s.authorizeDirectDatabaseWrite(w, r, "lock_steal", …) without a knownDirectWriteAuthOperations entry and every decision records as operation="unknown" (metrics.go:586) — exactly the regression this guard exists to catch.

Server-fault failures are answered 400, so a control-plane misconfiguration reads as a malformed request. resolveStorageSchemaTarget returns "this server does not expose its own storage schema" (:97) and a client-build failure (:117); both land on writeError(w, http.StatusBadRequest, …) at :161 and :207. A pre-deploy job that treats 4xx as "stop, your request is wrong" aborts on a bad configured endpoint instead of retrying or paging. It is also inconsistent within this PR: the identical no-adapter condition arriving remotely maps to 501 via writeStorageSchemaFailure, and an unreachable data plane to 503.

The apply handler runs a synchronous convergence on r.Context() without extending the 30 s write deadline. storage_schema_handlers.go:218 drives a bootstrap whose own budget is EnsureSchemaTimeout = 5 * time.Minute on a server built with WriteTimeout: 30 * time.Second (serve.go:224), while every sibling admin route first calls extendWebhookOpsDeadline (webhook_ops_handlers.go:123). A Spirit-backed ALTER TABLE that takes minutes still completes (it runs on its own context.Background() timeout), but the connection is torn down first and the planned/remaining pair never reaches the operator.

General suggestions

Nothing in the shipped binary calls SetStorageSchemaService, so both new routes always answer 400. The only production construction is serve.go:446 api.New(store, cfg, nil, logger), and serve.go:674 registers the tern server with no WithStorageSchemaService; every non-declaration match is in _test.go. Presumably the adapter lands later, but as merged the feature is inert and the remote error text ("upgrade that deployment and retry") points operators at the wrong remedy.

The guard resolves constant-named operations from a map it is still filling during the same walk. direct_write_auth_registry_test.go:109 calls collectStringConstants per file and scans that file's call sites immediately, so only constants from already-walked files resolve. A call site whose constant is declared in a later-sorting file makes stringArgument return named == false, the if !named { … return true } branch skips it, and the missing registry entry goes undetected. Today it passes only because every constant-named call site shares a file with its constant — collect all constants in a first pass, then scan.

The one thing that could have broken, verified

Routing a storage-schema request to the wrong plane. TernClient prefers a local client whenever s.config.Database(deployment) has a local DSN (service.go:451), so a deployment name colliding with a locally configured database could have had a data-plane convergence answered by control-plane storage. Checking TernDeployments.Endpoint before building the client is what prevents that, and the client.(tern.StorageSchemaService) assertion is a real second net: only *GRPCClient satisfies the interface (storage_schema.go:55) — LocalClient has no StorageSchemaPlan/Apply — so a collision is refused rather than silently misrouted.

Verified correct

  • Both new paths are POST with no readPaths entry, so TierForRequest lands them on TierWrite by the default rule, matching the handler doc and the new tiers_test.go cases.
  • Scoped-operator denial holds: both routes are in the sweep fixture and authorizeStorageSchemaOperation → AuthorizeDirectAdminWrite returns not_admin (403) for a write-capable non-admin.
  • status.Code unwraps through GRPCClient's fmt.Errorf("…: %w", err) (grpc-go v1.80 routes Code via FromError/errors.As), so Unimplemented still maps to 501, not the 500 default.
  • Storage schema RPCs are absent from retryServiceConfig's method lists, so a convergence is never automatically re-sent on UNAVAILABLE.
  • Nil-report handling is total: StorageSchemaReportFromProto(nil) returns nil and APIType() is nil-receiver safe, so a missing report yields 500 rather than a panic or an invented "converged".
  • Empty-body tolerance is right: Decode on http.NoBody returns io.EOF → zero request, while DisallowUnknownFields still rejects a misspelled field before anything is forwarded.

This review was generated by Claude Code (claude-opus-5).

@Kiran01bm Kiran01bm left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Approved on Kiran's (@kmuddukrishna) behalf by the scheduled review agent — no blocking findings at 3839be0. See the review comment above; non-blocking findings and suggestions, if any, are not merge gates.

@morgo morgo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Automated review on Morgan's behalf, at 29026a0.

Reviewing the delta since the approval at 3839be0c: two files moved, both in storage_schema_handlers.go and its test. Everything else in this PR is byte-identical to the approved state; the rest of the tree drift is the base branch moving underneath.

The change splits target-resolution failures by whose fault they are instead of answering every one with 400. I checked the partition itself, since that is the whole content of the change and a miscategorised path would be the bug:

Caller faults, left unwrapped and answered 400 — environment without a deployment, deployment without an environment, and a deployment/environment naming no configured data plane. All three are things the caller can fix from the response alone. ✓

Server faults, wrapped with an explicit status — no local storage schema service (501), a configured endpoint whose client will not build (503), a remote deployment resolving to an in-process client (500). None of these are fixable by the caller. ✓

The errors.As default is the safe direction: an unwrapped error stays 400, so adding a new caller-fault path needs no change here, and only a deliberate wrap can escalate a status.

The reasoning for why this matters is right and worth keeping: answering a server fault with 400 tells a pre-deploy gate its request was malformed, so it stops instead of retrying — which is the opposite of what a 503 should produce. Matching the local spellings to the 501/503 that writeStorageSchemaFailure already returns on the remote paths means one caller cannot see two different statuses for the same condition depending on which side resolved it.

Two nits, neither worth a round trip:

  • The doc comment says a server fault's "response deliberately does not carry the cause". That is exactly true of the 503 path, where the underlying error is logged and the response says only where to look. It is not true of the 500 path, which formats the in-process client's Go type through %T into the response body. Low consequence — this is a write-tier authenticated route, so the reader is already privileged — but the comment claims a property the code only has on one of the two paths.
  • 501 is >= http.StatusInternalServerError, so a server intentionally built without a storage schema service logs at Error on every such request. That is a build-time choice rather than a fault, and anything polling the route would produce a steady Error-level drip.

CI 41/41 SUCCESS.

Base automatically changed from armand/storage-schema-rpc to main September 16, 2026 17:09
aparajon and others added 4 commits September 16, 2026 13:09
…ailure has

Both storage schema operations were missing from the direct-write
authorization metric's allowlist, so every decision either route recorded
landed under operation="unknown" — the one label that makes a route's
denials invisible to the dashboard watching for them. A registry test now
pins the names that are written down at their gate, in the shape the
recovered-panic registry already uses.

The failure writer collapsed every data plane fault into a 500 pointing at
someone else's logs. An Unimplemented says the deployment needs upgrading
and an Unavailable says it needs looking at; both are remedies an operator
can act on without reading anything. A client that cannot be built for a
configured endpoint now answers with the fact and keeps the cause — which
names this deployment's own infrastructure — in the server's logs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…equest

Three of the ways a storage schema target fails to resolve are this server's
fault, and all three were answered 400. A build with no storage schema service,
a configured data plane whose client will not build, and a routing config that
sends a remote deployment to an in-process client are none of them things a
caller can fix by editing their request — and a pre-deploy gate that treats 4xx
as "malformed, stop" gives up on them instead of retrying or escalating.

They now answer 501, 503, and 500. The first two are the spellings the remote
paths already use for the same two conditions, so the same fault reads the same
whether it was local or came back over gRPC.

The log level follows the same split. A caller's mistake stays a warning: the
response says what to fix and nobody has to read this server's logs. A server
fault is an error, because the response deliberately withholds the cause and
these logs are where the operator goes next.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…t claims

The guard watched a gate named authorizeDirectWriteForPlan, which is not a
function in this repository — the real one is authorizeDirectWriteForStoredPlan
— and it did not list authorizeDirectDatabaseWrite at all. A gate named wrong
matches no call sites, the scan stays green on the ones it does find, and every
operation behind the misspelled or missing gate goes unchecked. Adding a
lock_steal call site with no registry entry left the guard passing, which is
exactly the regression it exists to catch; it now fails.

Two further holes, both of which let a missing entry pass:

Constants were collected and call sites scanned in the same directory walk, so
only constants from already-walked files resolved. A gate named through a
constant declared in a later-sorting file read as "named by a parameter" and
was skipped. Collecting every constant first, then scanning, removes the
dependence on filename order.

Nothing asserted the gates exist. Every name in the list is now checked against
the functions actually declared in the tree, so the next rename fails here
rather than silently narrowing what is watched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@aparajon
aparajon force-pushed the armand/storage-schema-api branch from 29026a0 to feef130 Compare September 16, 2026 17:09
@aparajon

Copy link
Copy Markdown
Collaborator Author

🤖 Five findings on 3839be0c. Three are fixed in d38e9f40 and feef130b; one is fixed two PRs up; one is the stack's shape. SHAs are post-rebase — the branch was rebased when #1393 landed.

The registry guard named a gate that does not exist. Correct, and worse than under-covering: authorizeDirectWriteForPlan matched nothing, so apply, lock_acquire and lock_release never reached the scan and require.NotEmpty kept it green on 10 operations. Fixed to authorizeDirectWriteForStoredPlan, added authorizeDirectDatabaseWrite, and added an assertion that every gate the test claims to watch is a function actually declared in the tree — a misspelling is now a failure rather than a silent hole. I ran your exact regression: a lock_steal call site with no registry entry now fails with operation "lock_steal" gated at pkg/api/lock_handlers.go:59 is missing from knownDirectWriteAuthOperations. finishDirectWriteDecision is in the gate set too, which is what covers the one gate that hardcodes its own operation name rather than taking it as a parameter.

The guard resolved constants from a map it was still filling. Also right, and the reason the name fix alone was not enough. Split into two passes — parse and collect every constant, then scan call sites over slices.Sorted(maps.Keys(parsed)). Verified with the case you described: a gated call whose operation constant is declared in a later-sorting file of the same package is now caught, where the single walk skipped it through the !named branch.

Server faults answered 400. Fixed: no local storage service → 501, a configured endpoint whose client cannot be built → 503, an endpoint that resolves to the in-process client → 500. They travel as one typed error through one refusal helper, so plan and apply cannot disagree, and a table test asserts both routes answer identically across the five cases. It also logs Error at 5xx and Warn below, which is the split a pre-deploy job's paging story depends on. 400 is now only what it should be: a deployment this server has no endpoint for, or a deployment/environment pair given half.

One gap worth naming, which your finding did not: the 503 branch has no test of its own. The table covers 501/500/400, and inducing 503 needs a config whose endpoint resolves but whose client build then fails. Happy to add it, but not as a commit that restarts this PR's CI — say the word and it goes in the next one.

The apply handler's write deadline. Real, and fixed in #1405 as 34ddce01: both routes lift the deadline to the budget their work is already bounded by, through extendWebhookOpsDeadline, the same helper the webhook operator routes use. Not in this PR because nothing can call the route until the adapter (#1404) and the CLI (#1395) land, and both merge ahead of it.

Nothing calls SetStorageSchemaService. That is #1404, the next PR up, which wires it in pkg/serve/serve.go. Until then both routes answer the no-adapter refusal, which is the honest answer for a server with no storage-schema service registered.


Replied by Claude Code (claude-opus-5) on Armand's behalf.

@aparajon
aparajon merged commit 5c2f1f0 into main Sep 16, 2026
41 checks passed
@aparajon
aparajon deleted the armand/storage-schema-api branch September 16, 2026 17:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants