Skip to content
Closed
Show file tree
Hide file tree
Changes from 3 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion docs/auth.md
Original file line number Diff line number Diff line change
Expand Up @@ -638,7 +638,7 @@ permission levels, called **tiers** in configuration and logs:
| Access | Operations |
|---|---|
| Read | List databases, pull live schemas, view stored plans, status, progress, logs, history, and locks |
| Write | Create plans, apply changes, stop or resume work, cut over, cancel, revert, skip revert, roll back, acquire or release locks, change settings, and run check or webhook maintenance |
| Write | Create plans, apply changes, stop or resume work, cut over, cancel, revert, skip revert, roll back, acquire or release locks, change settings, run check or webhook maintenance, and inspect or converge SchemaBot's own storage schema |

Creating a plan requires write access because it stages a change. Reading a
plan that already exists requires only read access.
Expand All @@ -650,6 +650,12 @@ to writes under `forward_auth`.
In the route rules, `GET` and `HEAD` requests are reads, as is `POST /api/pull`.
Other requests require write access by default.

`POST /api/storage/schema/diff` reads without changing anything, and still
requires write access under that default. It reports the internal shape of
SchemaBot's own bookkeeping database, and its sibling route converges that
database, so both belong to the people who operate the server rather than to
everyone who can see the schema changes it runs.

<a id="per-database-operator-scoping"></a>

### Grant a team access to its database
Expand Down
241 changes: 241 additions & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -1042,6 +1042,247 @@ storage:
Leave the flag false during normal operation and revert it after the removal
converges.

### Ask what storage DDL is outstanding

A deploy that did not converge leaves one question open: which storage DDL is
still outstanding. Two commands answer it, and both read the live storage
database. A release tag never stands in for that read: what a release would
converge to and what the storage actually converged to differ exactly when a
deploy has failed.

`storage diff` is read-only. It takes no lock and holds no transaction, so it
is safe at any time, including against production during an incident.

```console
$ schemabot storage diff
schemabot on db-1.example (mysql) needs 3 statements: 3 outstanding, against the schema embedded in v1.2.3.

Outstanding, and run automatically on the next boot or apply (3):

ALTER TABLE `applies` ADD COLUMN `driver_note` varchar(255) NOT NULL DEFAULT '' AFTER `lease_owner`;
ALTER TABLE `checks` ADD COLUMN `blocked_reason` varchar(64) NOT NULL DEFAULT '' AFTER `state`;
CREATE TABLE `check_gate_audit` (
`id` BIGINT UNSIGNED AUTO_INCREMENT,
`check_id` BIGINT UNSIGNED NOT NULL,
PRIMARY KEY (`id`)
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COLLATE=utf8mb4_0900_ai_ci;

Converge it with: schemabot storage apply
```

The statements are printed bare and one per line so a whole section can be
pasted into a client as it stands. The exit status is the machine-readable half
of the answer: `0` when the storage needs nothing, `2` when statements are
outstanding, and `1` when the read itself failed. A pre-deploy gate needs those
three apart, since "converged" and "unreachable" call for opposite decisions.

The headline answers the two questions that decide what the rest of the output
means. `schemabot on db-1.example (mysql)` is the database that was read —
reported by whoever read it, so it is not re-derived from a DSN, a config file,
or a deployment name. `against the schema embedded in v1.2.3` is the schema it
was compared against; [Which schema you are asking
about](#which-schema-you-are-asking-about) is how to change that.

`storage apply` converges the database by running the same bootstrap the next
boot would run: the same differ, the same refusal of destructive statements,
and the same advisory lock, so two operators running it at once serialize the
way two booting pods do. It previews the statements and prompts before running
them; `--auto-approve` (`-y`) skips the prompt for scripted maintenance.

```console
$ schemabot storage apply
schemabot on db-1.example (mysql) needs 1 statement: 1 outstanding, against the schema embedded in v1.2.3.

Outstanding, and run automatically on the next boot or apply (1):

ALTER TABLE `applies` ADD COLUMN `driver_note` varchar(255) NOT NULL DEFAULT '' AFTER `lease_owner`;

Run these statements against schemabot on db-1.example (mysql)? Only 'yes' will be accepted: yes
Ran 1 statement against schemabot on db-1.example (mysql).
schemabot on db-1.example (mysql) is converged.
```

Destructive statements are refused here exactly as they are at startup, and for
the same reason: a binary older than the storage sees the newer schema's tables
as surplus. A refusal is reported rather than silently dropped, and
`--allow-destructive` opts in per invocation, widening the deployment's standing
`allow_destructive_schema_changes` policy without ever narrowing it.

```console
$ schemabot storage diff
schemabot on db-1.example (mysql) needs 1 statement: 1 destructive, against the schema embedded in v1.2.3.

Destructive, and refused; surplus state stays in place (1):

-- check_gate_audit: DROP TABLE destroys data
DROP TABLE `check_gate_audit`;

Converge it with: schemabot storage apply
```

### Reach the right storage database

Both commands take a target, and which path applies is stated rather than
discovered. Nothing falls back from one to the other: a deployment that cannot
be reached through the API is an error naming the deployment, never a report
about a different database that happened to be reachable.

| Target | Reads |
|---|---|
| no flags | the storage of the server the CLI is pointed at |
| `--deployment <name> -e <environment>` | that data plane's own storage, over the gRPC connection that already exists between the two |
| `--dsn <dsn>` or `--config <file>` | the storage database this workstation opens itself |

A data plane owns its storage database and generally sits where a workstation
cannot dial it, so `--deployment` routes through the control plane: the control
plane asks the data plane, and the data plane reads its own storage with its own
embedded schema files. That is also what makes the answer trustworthy, since the
binary that reports the diff is the binary whose next boot would run it.

The direct path exists for when the server is down, including when it is down
because its own schema bootstrap is failing. `--dialect` states the storage
family when a DSN's form does not say; it applies only to a direct connection.

Both routes are admin-only and both sit at the write tier, the read-only diff
included, because the diff exposes the internal shape of SchemaBot's bookkeeping
database. Both are `POST` requests, so they take the write tier by the default
rule rather than by an exception. See [Authentication and
authorization](auth.md#what-read-and-write-access-include).

### Which schema you are asking about

The live side of the diff is always a read of the database. The desired side is
schema *files*, and three things can supply them:

| Desired schema | Where the files come from |
|---|---|
| no flag | the embedded files of the binary that answers the request |
| `--schema-dir <path>` | that directory's `.sql` files, read by the CLI |
| `--release <tag>` | that tag's `pkg/schema/<dialect>/` files, fetched by the CLI |

With no flag, the answer describes the release that is **currently running**:
the files are compiled in (`go:embed` over `pkg/schema/mysql/` and
`pkg/schema/postgres/`), so through the API it is the server or data plane that
answered, and on the direct path it is the CLI binary you are running. That is
the right default — it is what the next boot would converge — and it is the
wrong question before a roll, when the running release reports convergence while
the release about to deploy still has work to do.

`--release` asks that question without a binary of that release:

```console
$ schemabot storage diff --deployment west -e production --release v1.4.0
schemabot on db-1.example (mysql), deployment west in production needs 1 statement: 1 outstanding, against the schema files of release v1.4.0 in block/schemabot.

Outstanding, and run automatically on the next boot or apply (1):

ALTER TABLE `applies` ADD COLUMN `driver_note` varchar(255) NOT NULL DEFAULT '' AFTER `lease_owner`;

These are what schemabot on db-1.example (mysql), deployment west in production needs in order to match the schema files of release v1.4.0 in block/schemabot, not what its own next boot would run. To converge them, run that release's binary against this database — its container image is that release — or let the release's first boot converge them.
```

Because the storage schema is declarative, one diff against the release you are
rolling to covers however many releases lie between; there is nothing to step
through.

The report names the schema it used, always, and never relabels a schema you
supplied as the answering binary's own. That line is the difference between two
correct reports about the same database, so read it before acting on the
statements.

Details of the two selectors:

- **`--schema-dir <path>`** reads `*.sql` directly from a checkout or an
extracted image layer, one file per storage table. Point it at the dialect
directory (`pkg/schema/mysql`), not at its parent. A directory with no `.sql`
files is an error naming the path — a diff against an empty schema would
report every existing table as surplus.
- **`--release <tag>`** fetches the files over the repository's contents API at
that tag. It reads the schema directory for the dialect the *live storage*
runs, which it learns by first asking the target — one extra read-only diff,
paid only by this flag. `--release-repo` points at a fork or mirror
(`block/schemabot` by default), `GITHUB_API_URL` at a different API host, and
`GITHUB_TOKEN` or `GH_TOKEN` authorizes the fetch. A repository the CLI cannot
read is an error naming the token to set and `--schema-dir` as the offline
alternative. Naming both selectors is refused rather than resolved by
precedence.

`storage apply` has neither flag. A convergence runs the schema embedded in the
binary running it, so that it does exactly what that binary's next boot would
do — the property that makes it usable as a pre-deploy step at all, and the one
that keeps an older binary from being handed newer schema to destroy. Passing
either flag to `apply` is refused with the two real ways to converge a release:
run that release's binary, or let its first boot do it.

### Deploying a release that changes the storage schema

Every startup converges the storage schema on its own, so the routine case
needs none of this. Reach for the commands when the release notes name a
storage schema change, when the tables involved carry a long history, or when a
pod is not starting.

1. **Before the roll, ask the new release what it will run.** Name the release
being deployed, from whatever CLI you have to hand:

```bash
schemabot storage diff --deployment west -e production --release v1.4.0
```

Exit status 0 means that release's boot has nothing to do and the rest of
this does not apply. A binary of the new release answers the same question
with no flag, which is what to use where the tag cannot be fetched:

```bash
schemabot storage diff --dsn "$STORAGE_DSN" # run from the new release's binary
```

2. **Decide whether the boot should do it.** Additive DDL inside the
five-minute startup budget is fine when the tables are small. It is not fine
when they are not: on MySQL an index added to an existing storage table runs
as Spirit online DDL, a table copy whose cost grows with row count, and every
pod in the roll pays it. Converge once, ahead of the roll, instead:

```bash
schemabot storage apply --dsn "$STORAGE_DSN" # from the new release's binary
```

The convergence has to come from a binary of the new release: `apply` runs the
schema embedded in whatever binary runs it, and there is no flag that points
it at a release's files. Use that release's container image as a one-shot job
if there is no binary to hand.

3. **If you converged ahead of the roll, re-check right before it.** A table or
a column you created early survives a boot of the current release, because
dropping one is destructive and is refused. **An index does not.** Dropping
an index destroys no data, so it falls outside that refusal, and any boot of
the still-running older release converges the new index away without
comment — a pod restart, a scale-up, a health-check replacement. Re-run the
step 1 diff immediately before rolling, and treat a long gap between
pre-creating an index and deploying as a gap the index probably did not
survive.

4. **After the roll, confirm through the API, per deployment.**

```bash
schemabot storage diff # this server's storage
schemabot storage diff --deployment west -e production # a data plane's storage
```

Now the running binary is the new release, so exit status 0 is the
confirmation that its storage converged.

5. **If a pod is crashlooping, ask directly.** A failed storage bootstrap keeps
the server from accepting traffic at all, so the API cannot answer for it.
The direct path can, with the same binary the pod runs, and the statements it
prints are the ones the pod is failing on. `storage apply` from there clears
it under the same advisory lock the pods are contending for.

During a rollback window the diff reports the newer release's tables and columns
as refused destructive statements and exits 2. That is the expected steady state
rather than drift: the surplus state is deliberate, and it is what lets the
release be rolled forward again. A pre-deploy gate keyed on exit status 0 will
flag it, which is the correct signal to pause on.

## Support Channel

SchemaBot can add an opt-in support link to GitHub PR comments so authors know
Expand Down
28 changes: 15 additions & 13 deletions docs/invariants.md
Original file line number Diff line number Diff line change
Expand Up @@ -312,17 +312,18 @@ clamp (`pkg/webhook/plan_drift.go`); the request body limit (`pkg/webhook/handle

### AV-9: SchemaBot never destroys its own storage to start

The startup schema bootstrap converges SchemaBot's own storage additively, and decides before it
writes. On MySQL a destructive statement (a `DROP TABLE`, or an `ALTER TABLE` carrying a `DROP
COLUMN`) is refused unless destructive storage changes are explicitly allowed, and a statement
whose destructive clauses cannot be partitioned out is refused *whole*. Refusing the whole
statement runs strictly less than any split of it, so the fallback can never widen what the
bootstrap executes, and startup continues on the safe remainder. On PostgreSQL the convergence is
additive-only and gates on the entire drift set before touching anything, so a change needing
manual remediation aborts the pass rather than leaving storage half-converged. *Breaks if
violated:* the first instance of a rolling deploy drops state the rest of the fleet is still
reading. *Enforced:* the per-dialect bootstrappers (`pkg/api/ensure_schema.go`,
`pkg/api/ensure_schema_postgres.go`).
Every convergence of SchemaBot's own storage — at startup, or on an operator's command — is
additive, and decides before it writes. On MySQL a destructive statement (a `DROP TABLE`, or an
`ALTER TABLE` carrying a `DROP COLUMN`) is refused unless destructive storage changes are
explicitly allowed, and a statement whose destructive clauses cannot be partitioned out is refused
*whole*. Refusing the whole statement runs strictly less than any split of it, so the fallback can
never widen what the bootstrap executes, and startup continues on the safe remainder. On
PostgreSQL the convergence is additive-only and gates on the entire drift set before touching
anything, so a change needing manual remediation aborts the pass rather than leaving storage
half-converged. *Breaks if violated:* the first instance of a rolling deploy drops state the rest
of the fleet is still reading. *Enforced:* the per-dialect bootstrappers
(`pkg/api/ensure_schema.go`, `pkg/api/ensure_schema_postgres.go`), which the operator-facing
storage schema surface calls rather than reimplements (`pkg/api/storage_schema.go`).

### AV-10: Anything the PR can do, the CLI can do

Expand Down Expand Up @@ -1326,8 +1327,9 @@ to another. *Enforced:* server-side routing (`pkg/tern/target_router.go`) and so
An API route not classified as a read is treated as a write. Planning counts as a write, since it
stages a change. A target that cannot be resolved never authorizes. A configured grant that could
never match any request is a startup error rather than silent dead config. And a new mutating
endpoint cannot ship without a test proving it denies unauthorized callers. *Enforced:* route
classification with a structural sweep test over the route table (`pkg/api/service.go`).
endpoint cannot ship without a test proving it denies unauthorized callers. *Enforced:* the tier
classification (`pkg/auth/tiers.go`), with a structural sweep test over the route table
(`pkg/api/service.go`).

### AZ-3: Identity comes from a verified lane

Expand Down
14 changes: 14 additions & 0 deletions docs/release.md
Original file line number Diff line number Diff line change
Expand Up @@ -271,6 +271,20 @@ statement (see [configuration.md](./configuration.md)). A destructive
change is a coordinated operation and belongs in the release notes with
instructions, not in a routine patch release.

The diff does not have to be read out of the files by hand. Name the release
being tested and the command answers it against the live database:

```bash
schemabot storage diff --deployment west -e production --release v1.4.0
```

Name the release, or run the command from a binary of it — with neither, the
answer describes whatever release is currently serving, which before a roll is
the old one. Operators pre-creating an index ahead of the roll should also read
[Deploying a release that changes the storage schema](./configuration.md#deploying-a-release-that-changes-the-storage-schema)
— a pre-created index is removed again by any boot of the still-running earlier
release, because dropping an index is not destructive and so is not refused.

### 3. The public Go API

SchemaBot is importable as a Go library, not only runnable as a binary, so its
Expand Down
Loading
Loading