the record is being maintained after the reboot
-
verity sdk is working, serializes models and push it to the ingestion pipeline on server. The ingestion pipeline hands over the model to orchestrator to build artifacts. Nothing automated
-
Serializes and uploads model at staging and gets artifacts and stores model withOUT any extension. straight what is returned from the llm returns. and the orchestration is set up till this
-
Nat (agent 2, evaluation) is in. The Hawkeye -> Nat handoff is automatic now: one
/ingestcall identifies the model AND evaluates it, andmodel_version.statusmoves offpendingon its own ->stagingon a pass,staging_failedotherwise. Ships with the SDK, soassemble(model, X_test=..., y_test=...)is the whole loop.- The eval mechanism is chosen from the fixture's
kind, not hardcoded. Onlylabeled_holdoutis registered; images/RAG/rollout are a module + a registry line later, not a rewrite (server/agents/brain2/nat/registry.py). predict()runs in a subprocess with every credential withheld (server/execution/sandbox.py). There is a test that proves the child cannot read SUPABASE_KEY.- Systemic metrics (latency percentiles, throughput, memory, CPU, GPU) are measured by that harness on every run and gate exactly like quality metrics. Verified: a model scoring accuracy 1.0 still fails on a latency threshold.
- Caveat, deliberately: those numbers are single-process, single-client, cold sandbox. Feasibility, not production p99. Real load is Falcon's.
- The eval mechanism is chosen from the fixture's
-
shifted the workflow to aws service. There were several challenges with seaweedfs, first i had to get a virtual machine, for hosting seaweedfs distributed storage or else there was no possible way. got the free aws account with $100 free credits, so spun up a s3 bucket, we could have also done a ec2 instance, but utilizing better infra since available.
-
Fury (agent 3, registry) is in. The Hawkeye -> Nat -> Fury handoff is automatic now — one
/ingestcall identifies, evaluates, AND registers. This changes what entry 3 said: a passing verdict no longer stops atstaging— it goes straight toproduction, since Fury is what actually decides promotion now, not the eval step itself.stagingis effectively retired as a resting state;staging_failedandpendingare unaffected.nameis now required on every upload (assemble(model, name=..., ...)), and it's the only thing that decides "is this a new model or a new version of one I already know." Uploads grouped under the same name share onemodelrow and one lineage.- Exact byte-for-byte re-uploads under the same name are a true no-op — checked before
anything else runs, so a repeat costs nothing (no Hawkeye call, no Nat call, no
sandbox execution). Verified live: re-uploading an identical model returned the
existing record with
deduplicated: true, no new row. - Promotion is fully automatic on a passing verdict, with no comparison against the
incumbent — a new version that clears its own thresholds replaces whatever's
currently
production, and the one it replaces moves toarchived. Both keep their ownpromoted_frompointer to the eval_run that justified them, so the archived version's evidence isn't lost. Verified live end-to-end against real S3/Supabase/Groq: first version of a new model -> production; a second version -> production, first -> archived; the same second version re-uploaded -> deduped, no new row. - Accepted, undecided risk, not solved here: a version that barely clears its own bar can replace an incumbent that was scoring far better, since nothing compares them. Mitigated only by visibility (both eval_runs stay on record), not prevented.
- Fixed along the way, not code: a real repo-hygiene bug where
verity/testswas gitignored, so every prior commit touching SDK tests had silently no-op'd across this whole project — the suite existed on disk but was never actually in git history until this was caught and fixed. Still not automated past here: api-fication (nothing actually serves aproductionmodel yet), Falcon, theagent_runaudit trail, comparative promotion gating.
-
client/is no longer an untouched scaffold — added an MVP intake form (single page,client/src/app/page.tsx) that calls/ingestdirectly from the browser and renders the real response: manifest, per-threshold pass/fail, a verdict stamp. Ships a bundled demo model+fixture (client/public/demo/) so the whole Hawkeye->Nat->Fury loop is one click, no Python needed to try it.- The server had no CORS config at all — added it, scoped to localhost:3000, since a browser can't call a cross-origin API without it.
- Client-side SHA-256 (Web Crypto) computes the same digest the server now verifies (the fix-wave's finding #3) — the frontend can't get away with a stale/wrong hash any more than the CLI can.
- Fixed a second gitignore-too-broad bug, same category as
verity/tests: the root.gitignore's baredemo/pattern was also silently matchingclient/public/demo/— scoped it to/demo/(root only). - Verified against the real backend, not just a build check: a Node script replicating the exact browser flow confirmed all three response shapes the UI renders — pass to production, an archived-replacement, and a deduplicated repeat. Still frontend-only for identify+evaluate+register; there's no view of a model once registered (no GET endpoint exists server-side yet), so this only ever shows the result of the upload just made, not a history/dashboard.
-
Falcon (agent 4, observability) is in — the V1 loop's four agents are all built. A promotion to
productionnow switches monitoring on by itself, and the SDK reports live traffic back from wherever the customer serves the model.- Falcon needs no LLM. Everything it does is deterministic: when Fury promotes a version,
promoted_fromalready points at the eval_run that justified it, and that row already holds measured latency percentiles, throughput, memory and quality scores. The baseline is lifted from evidence that already exists rather than guessed. - The honesty constraint that shaped the whole design: those eval numbers come from a
single-process, single-client, cold sandbox. Production latency under real concurrency
will be materially higher. So the config records them tagged
"basis": "sandbox_feasibility"for context and for V7's rule engine — and V1 compares nothing and alerts on nothing. A baseline that's wrong by construction would just train people to ignore it. Newmonitoring_configtable (Schemas.md never specced one, only the README diagram named it). verity.monitor(model, model_version_id=...)wraps the customer's model;.predict()behaves identically and telemetry ships from a background thread. The governing rule is that telemetry can never be why inference fails or slows: enqueue is non-blocking and drops on a full queue, the HTTP call never touches the predict path, and if the model raises, the caller gets that exact exception back — recorded, then re-raised unchanged.GET /models/{id}/telemetryand a panel in the intake UI make it visible, so this isn't write-only.inputs/predictioncolumns exist but stay null at V1 — they're for V7 drift detection, and leaving them empty removes the whole sampling question.- Found and fixed during review, not shipped: a background security review caught that
statuswas an unconstrained string (anything != "ok" counts as an error, so a client sending "OK" would silently inflate the error rate) and that the batch size was unbounded; and the SDK reporter leaked a thread + atexit reference per instance, could stall shutdown for minutes against a dead endpoint, and had a racy drop counter. - Declined deliberately:
/telemetryhas no auth. Real, but systemic — the whole app has none by design until V1.5, and/ingestis equally open while accepting arbitrary pickles. Securing one endpoint while that stands would be theatre. Still not automated past here: api-fication (nothing serves aproductionmodel yet — Verity will, once built; until then the customer serves it and the SDK reports back), alerting and drift (V7), theagent_runaudit trail, comparative promotion gating.
- Falcon needs no LLM. Everything it does is deterministic: when Fury promotes a version,
-
Housekeeping, one real outage fix, and the api-fication decision. Nothing new shipped in the pipeline — this entry exists so the next reader isn't confused by three things that changed underneath them.
- Groq retired the entire Llama line.
llama-3.3-70b-versatilestarted returning 404model_not_found, which broke both LLM agents at once.GET /openai/v1/modelson the live key showed what was actually reachable;server/agents/provider.pynow defaults toopenai/gpt-oss-120b. Two tests had pinned the old model as a bare string literal and failed for the wrong reason — they now assert againstDEFAULT_MODELimported fromagents.provider, so they pin "the provider default is used" rather than which model that happens to be this month. The bug was in the test, not the code: pinning a vendor's model name is pinning something you don't control. - The repo still contained an entirely different, abandoned project. The root
pyproject.tomldeclaredverity-eval("local-first RAG evaluation") with torch, transformers, sentence-transformers and pinecone as dependencies, plus a matchingrequirements.txt,uv.lock, a 1.1 GB.venv, and averity_eval.egg-info/. Its[tool.setuptools.packages.find] include = ["verity*"]claimed the SDK's package name from the wrong directory. All deleted. The real packages areverity/pyproject.toml(SDK) andserver/pyproject.toml(server).Architechture.mdandagent.mdwent with it — pre-reboot documents whose content the README now covers, and whose content had started to actively contradict the current design. - Docs moved into
docs/, leaving onlyREADME.mdat the root:architecture.md(wasverity-architecture-and-workflow.md),Schemas.md,Metrics.md,progression.md, andreference/mlflow.md.Schemas.mdandMetrics.mddeliberately kept their filenames — roughly fifteen code comments and both committed design specs refer to them by bare name in prose, and moving a file leaves that prose true where renaming it would make fifteen comments lie. - Two documentation lies found while doing it, both worth more than the tidying.
Schemas.mddocumented sixmanifestcolumns —io_schema,serving_pattern,platform,confidence,review_required,declared_overrides— that no migration ever created, and omitted four that exist. The table now marks which is which. Andarchitecture.mdhad no section on Falcon at all: it was written before Falcon existed and still claimed three agents, one route, and 102/21 tests. Both are the same failure mode — a document that described intent and was never re-read against reality. - api-fication is now designed rather than open. Settled: one container image per promoted
version, built from the training environment the SDK captures at
assemble()time (so a pickle is never loaded against a different sklearn than it was written with); request schema generated from the model's own introspected surface; deploy fires automatically on promotion, non-fatal like Falcon; Verity proxies/models/{name}/predictto the live container. Local Docker first, behind a runtime interface so a cloud runner is a new class rather than a rewrite. Scope is narrowed on purpose to tabular/classical ML — sklearn, XGBoost, LightGBM — finished end to end before a second model class starts. - Correction on the record, because an earlier entry implied otherwise: the plan was for some time that Verity would never serve inference, and that the customer would always host the model with the SDK reporting back. That was my inference from an unanswered question, not a decision anyone made, and it had already been written into the Falcon spec as settled fact. Verity serves; self-hosting plus SDK telemetry stays as the secondary path.
- Groq retired the entire Llama line.
-
api-fication is in — Verity serves the models it promotes. A version that reaches
productionis now built into its own container image, started, health-checked, and put behindPOST /users/{user_id}/models/{name}/predict. This is the first entry where "the deployment side ends" is literally true rather than aspirational.- One image per version, built from the training environment. The SDK captures its own
scikit-learn/numpy/scipy/cloudpickle/Python versions atassemble()time and ships them; the image is built from those pins. A pickle is only reliably loadable against the versions that wrote it, so a single shared serving image was never an option — and the capture has to happen client-side, because introspecting on the server would faithfully report the wrong machine. A free consequence: the artifact is copied into the image, so the serving container needs zero credentials, the same principle the eval sandbox already had, reached from a different direction. - The request schema is measured, not guessed.
execution.sandbox.introspect()readsn_features_in_,feature_names_in_,classes_andpredict_probaoff the fitted estimator inside the same scrubbed subprocesspredict()already ran in, at ingest, beside Hawkeye. Structure is a fact about the object; only semantics need an LLM.feature_namesbeing None is a real answer — it makes the served API positional rather than named, instead of inventing column names. manifest.io_schemais populated for the first time. It had been specced in Schemas.md since the first draft and never created.telemetry_event.inputsand.predictionare non-null for the first time in this project's history. They were created with Falcon and left empty because a customer-hosted model has no reason to ship payloads back. Verity serving is what makes them writable, and therefore what unblocks drift detection — that is why serving came before the drift metrics rather than after.- Deploy is non-fatal, like Falcon. It runs after the version is already
production, so a build failure writes afaileddeployment row and returnsdeployment: nullrather than 500ing a promotion that genuinely succeeded. Verified live by pointing the server at an unreachable Docker daemon:DockerRuntimeraised,deploy()recorded the reason in the real database, and the orchestrator swallowed it — all three layers behaved. - Verified live end-to-end: a real 585 MB image built and started, the proxy returned real
predictions, a second upload under the same name promoted and tore down the container it
replaced (old one
Exited (0), proxy re-routed with no intervention), and telemetry rows landed carrying inputs and predictions. - Two bugs found by running it for real, not by tests. The SDK's 60s HTTP timeout could
not survive a cold container build (~3 min), so the client reported
ReadTimeoutfor a request the server went on to complete successfully — raised to 600s and made configurable. Anduv syncwithout--extra devsilently pruned pytest from the server venv, souv run pytestfell back to the global interpreter; the suite kept passing against the wrong packages, and a pandas test that should have skipped "passed" by accident. Fixed withuv sync --extra dev, after which the full suite runs on the venv it claims to. - One honest loose end:
manifest.serving_patternwas created and is never written. Stamping it would mean updating an append-only evidence row after the fact, and thedeploymentrow already records how a version is served. It stays empty rather than being written dishonestly; a later migration drops it if nothing claims it. Still not automated past here: drift and data-quality metrics (now unblocked), alerting (V7), theagent_runaudit trail, comparative promotion gating, and anything beyond a single local replica — deployments live and die with the machine running Verity.
- One image per version, built from the training environment. The SDK captures its own
-
Falcon detects and notifies now — what entry 9 still listed as "alerting (V7)" moved forward into V1, brought closer by api-fication itself: once the proxy measures real production latency (entry 9) and sees real production inputs, comparing live traffic against itself rather than against a cold-sandbox figure stopped requiring anything V3's analytics store was supposed to unlock first.
- Two checks, zero LLM calls, reusing machinery that already existed rather than
reimplementing it.
server/agents/brain4/falcon/detect.py'sdetect_systemic_anomalycompares a version's trailing 15 minutes of traffic (error_rate, thenlatency_p95_ms) against the 15 minutes before that — never againsteval_reference, which is a cold-sandbox estimate that always looks better than real traffic and would just manufacture false alarms.detect_quality_anomalyre-runs Nat's ownscore()/apply_thresholds()against accumulated delayed labels, against the exactmetric_setand thresholds frozen at promotion time — not a second scoring implementation, the same one. BelowWINDOW_MIN_EVENTS=20events in either window, orMIN_LABELS=30accumulated labels, both are a deliberate no-op: a handful of events producing a scary-looking number is noise, not signal. - No scheduler, no poller — both checks run inline, triggered by the events that already
move data.
check_systemicfires fromTelemetrySink.flush()(once per distinctmodel_version_idin the batch, after the write succeeds) and fromPOST /telemetry, covering both a Verity-served model and a customer-hosted one reporting through the SDK.check_qualityfires from the newPOST /predictions/{prediction_id}/outcomesonce its batch of labels lands. Every call site wraps the check in its owntry/excepton top of the check's own internal one — belt and suspenders around the one rule that must never break: a detection bug can never turn a successful write into a failed response. prediction_id, minted before the row that would hold it exists. The proxy'spredict()generatespred_<uuid4>before calling the container, not read back from a database insert —TelemetrySinkonly queues the write, so the biginttelemetry_event.iddoesn't exist yet when the response goes out, and a customer reporting a delayed outcome needs something to correlate against immediately. Newlabel_eventtable, upserted on(telemetry_event_id, instance_index)so a corrected label overwrites rather than double-counts.- The row is the truth; email is best-effort on top of it.
record_and_notifywrites a newalert_eventrow unconditionally, first, then attempts an email via AWS SES (same account api-fication's S3 usage already lives in) only if the model has a registeredalert_email— a raising or failing send is swallowed and never un-writes the alert.emailed_atstaying null is the only record delivery didn't happen; nothing retries it. alert_emailthreads all the way fromverity.assemble(model, ..., alert_email=...)throughupload()→/ingest→orchestrator.build_artifact()→server/agents/brain3/fury/registry.py'sregister()→create_model(). Deliberately one-way: the existing-model branch ofregister()never touches a model'salert_emailon a later version's upload, so re-uploading an already-registered model can't silently change who gets notified.- Verified live, not just against the unit suite — with one constraint worked around
rather than fought.
check_systemiccompares two windows exactly 15 minutes apart byoccurred_at, which is set server-side and isn't backdateable through any API; genuinely exercising it with two real, temporally-separated windows would cost ~30 real minutes of waiting for one verification step, which wasn't worth burning. What was run for real: started the server, usedverity --demoto get a promoted, Docker-deployed model, sent 35 clean predictions in one burst (one of 36 attempts hit the unrelated flake noted below) —GET /models/{id}/alertsstayed empty, confirmingcheck_systemicruns without error and correctly stays silent with no aged baseline window yet (not a false negative; there was nothing 15-30 minutes old to compare against). Then, against the same model, reported 32 deliberately wrong outcomes (actual = 1 - predicted) through the new outcomes route — three realqualityalerts fired onceMIN_LABELSwas crossed, detail{"metric": "accuracy", "op": ">=", "value": 0.75, "actual": 0.0}, thresholds read back verbatim from the promotingeval_run. A second version was assembled withalert_email="ops@example.com"set and put through the same wrong-outcome burst specifically to exercise the SES-failure path (noSES_SENDER/SES-authorized credentials configured in this environment — only S3's, which don't authorizeses:SendEmail): two more alerts fired,emailed_atstayed null on every alert across both runs, and every triggering request —/predict, and all 63/predictions/.../outcomescalls — returned 200. The systemic check's own comparison arithmetic is whattest_falcon_detect.pyandtest_falcon_monitor.pyprove instead, with synthetic timestamps exactly 15 and 30 minutes apart — that is the actual proof of correctness for that half, not a claim of a live soak test that wasn't run. - Found live, not a regression in this feature: two of 171 live requests logged during
this verification hit a
transient
httpx/SupabaseRemoteProtocolError: Server disconnectedinsidefind_model/find_production_version— pre-existing code this task didn't touch, both simple retries, named here for the record rather than silently absorbed. - Accepted risks, named in the design spec, not solved here: one global
RELATIVE_INCREASE_THRESHOLDfor every model regardless of how bursty its traffic naturally is; a model whose customer never reports a single outcome is invisible tocheck_qualityforever, indistinguishable from "the model is fine"; each window only ever compares against its immediate predecessor, so a slow decline spread across many windows never trips any single comparison (needs a longer historical baseline — the analytics store V3 already earmarks for telemetry volume, not invented here); best-effort email only, no retry queue for a failed SES send. - Security finding surfaced during implementation, deliberately not fixed:
POST /predictions/{prediction_id}/outcomeshas no ownership check. Crucially, this does not depend on leaking anyone'sprediction_id:/predictis equally unauthenticated, so an attacker can mint their ownprediction_idfor free by calling/predictagainst the target model, then report whatever outcome they like against it. Same class of gap/ingest,/telemetry, and/predictalready carry by the standing, already-reasoned decision that no route has auth until V1.5 — securing this one route alone would be close to worthless, since it'd only confirm a caller owns something trivially self-mintable. Called out specifically here because this route feeds directly into the alerting signal, and suppression is the sharper risk than manufacture: fabricating CORRECT outcomes to hide a genuinely degraded model's accuracy is silent and indistinguishable from healthy — worse than a false alarm, which is merely noisy and self-correcting. The real fix is V1.5 auth on/predict(closing where aprediction_idcan be minted), not a check here. This flips from "acceptable" to "must fix before V7 ships" once Falcon acts autonomously on its own signal — today a human always reviews the alert first, bounding the blast radius to wasted attention rather than an automated action. Still not automated past here: input drift detection (a separate mechanism — needs a distributional distance metric that doesn't exist yet), autonomous retraining (a human decides, always), theagent_runaudit trail, comparative promotion gating, and anything beyond a single local replica.
- Two checks, zero LLM calls, reusing machinery that already existed rather than
reimplementing it.
-
Container serving can now run on AWS ECS Fargate instead of only local Docker — done specifically to get CloudWatch Logs for free and a serving process that survives independently of wherever the local server happens to be running.
serving/runtime.pywas already built with aContainerRuntimeseam anticipating exactly this (build/run/stop,DockerRuntimethe only implementation);FargateRuntimeis the second one, selected viaVERITY_CONTAINER_RUNTIME=docker|fargate(deploy.py), defaulting todockerso nothing about existing local dev changes unless explicitly opted in.- The seam had a real gap once a second implementation actually needed it:
deploy.pyhardcodedf"http://localhost:{host_port}"rather than trusting the runtime to say where it's reachable — invisible while Docker was the only implementation, sincelocalhostwas always correct by construction.run()now returnsendpoint_urldirectly;host_portis optional in the dict and nullable in thedeploymentrow, meaningful for Docker's ephemeral port, absent for Fargate. build()delegates the localdocker buildto an internalDockerRuntime, then pushes to ECR.run()registers a new revision of one task definition family, launches withassignPublicIp=ENABLED(the server calling this runs locally, not inside the VPC), polls until the task reachesRUNNING, then resolves its ENI to a public IP.stop()is a directecs.stop_task. Task lifecycle matches Docker's exactly: a version's task runs until a new promotion under the same name replaces it — no autoscaling, no idle-shutdown, the same "single replica" risk api-fication already accepted, not a new one.- A real bug found live, not by review:
docker-py'simages.push()does not raise on a failed push by default — it returns the streamed log as an opaque object, and anyerrorentry in it is silently ignored unless the caller checks. An early version of_push_to_ecrdidn't check, so a real push against the real ECR repo silently produced zero images whilebuild()reported success — only surfaced because the live Fargate test then failed downstream with a confusingCannotPullContainerError. Fixed by requesting the decoded stream (stream=True, decode=True) and raisingContainerRuntimeErroron anyerrorentry, with a test that proves the raise actually fires (and that the message isn't a redundant double-wrap of itself — a second real bug the fix's own review caught before it shipped). - Live-verified twice, at two different layers. First, the isolated runtime test
(
test_a_real_model_deploys_to_fargate_and_answers_health, gated behind@pytest.mark.awsandVERITY_RUN_FARGATE_LIVE_TEST=1— never inferred from the presence of AWS credentials, matching the existing@pytest.mark.dockerprecedent but opt-in instead of auto-detected): a real image built, pushed, launched as a real Fargate task, reachedRUNNING, answered/health, stopped. Second, the full app stack end to end:VERITY_CONTAINER_RUNTIME=fargate, a realverity --demopromotion, a real deployment row withendpoint_urla genuine public IP (notlocalhost) andhost_portnull,/predictanswered correctly both through Verity's own proxy route and by hitting that public IP directly from outside the VPC. The task was stopped and confirmed gone afterward. - Two transient failures along the way, neither a code defect: IAM policy propagation
delay right after permissions were granted (resolved itself within minutes, confirmed
by a standalone reproduction outside pytest), and one demo run that legitimately failed
its own eval gate on an LLM-proposed
resource.cpu_time_sthreshold too tight for a trivial demo model — ordinary Nat threshold-proposal noise, not related to Fargate. - Accepted risks, named in the design spec, not solved here: no autoscaling or idle-shutdown (a continuously-billed task for as long as it's the live version); a real cold-start on every replacement, on top of whatever the image build/push took; a fresh public IP per task with no load balancer or stable DNS name in front of it.
- The seam had a real gap once a second implementation actually needed it:
-
The frontend gets a Models tab — a real browsable registry, not just an upload form that forgets what you gave it. Four new read-only routes (
GET /users/{user_id}/models,GET /models/{model_id}/versions,GET /model_versions/{model_version_id},GET /model_versions/{model_version_id}/download-urls) back a newserver/registry.pymodule, kept deliberately separate fromorchestrator.py— one assembles a response from values just computed mid-pipeline, the other re-queries already-stored rows for a version created any time in the past.- The two list routes return an empty list for an unknown user/model rather than 404 (no user table to check existence against); the two detail routes 404, since those are keyed to one specific thing just clicked on. Deliberately asymmetric, not an inconsistency.
- The detail route's response is shaped to match the frontend's existing
IngestResulttype exactly, which is what letsclient/src/components/version-detail.tsxreuseEvidenceReportcompletely unmodified — the same component now renders both a just-uploaded result and a version pulled back up later. - Downloads are presigned S3 URLs (15-minute expiry), generated on request — no file bytes pass through the Verity server, matching how MLflow itself handles an S3-backed artifact store.
- Live-verified via the real
veritySDK, not the browser's demo form: two versions of one model uploaded back to back (one failed its eval gate, one promoted toproductionwith a real local Docker deployment), then confirmed through the new routes end to end — both versions listed newest-first under their shared model, the detail route rendering correctly for both the passing and failing version, and a presigned artifact URL that produced a real200download of the actual pickled bytes. - Found and fixed along the way, not a code bug: a fresh worktree's venv had never
had
uv sync --extra devrun, souv run pytestwas silently resolving to some other environment lackingsupabase— the exact standing gotcha this project's own task briefs already warn about. One test's report was rejected by review for claiming a test count without pasting real command output, which is what caught it; fixed by running the sync once, no code changes needed. - Out of scope, named in the design spec: run comparison, metric-over-time charts, search/filtering, tagging, and promotion/staging buttons — browse, detail, and download only, not a full MLflow reimplementation.
- Accepted risk, not solved here: the download route makes a model's artifact
bytes reachable by anyone who can reach the server and guess or already knows a
user_id— the list routes enumerate model/version ids from there, so nothing after that first guess needs guessing. The 15-minute presigned-URL expiry bounds how long a leaked URL keeps working; it does not gate who can mint one. Same standing "no auth until V1.5" decision every other route already carries, named here because this is the first route that hands back the underlying file.