Skip to content

Backend pods still migrate when migrationJob is enabled, causing lock races on upgrade #134

Description

@sdolidze

Bug

When migrationJob.enabled=true, the backend Deployment pods still run database migrations on startup, in addition to the dedicated pre-upgrade migration Job. During a helm upgrade this means two or more processes (the Job plus every rolling backend pod) attempt to migrate the same database concurrently. They contend on the Knex migration lock, and if a migrator is killed mid-migration the lock is left held, which then blocks every subsequent startup.

Steps to reproduce

  1. Deploy the chart with migrationJob.enabled=true (and no image.command override).
  2. Run helm upgrade to a release that applies a non-trivial migration, against a database large enough that the migration takes more than a few seconds.
  3. The pre-upgrade *-migrate Job starts migrating; meanwhile the rolling backend pods start and their entrypoint also runs knex migrate:latest.
  4. Observe: backend pods crash-loop, and the logs show the lock is held.

Expected behavior

When the migration Job is enabled, it should be the only process that runs migrations. Backend pods should start the server without attempting to migrate, so there is exactly one migrator per upgrade.

Actual behavior

Multiple migrators race for the lock. The losers fail and (because the entrypoint uses set -e) the pod exits non-zero and crash-loops. If a migrator is killed mid-run the lock is orphaned, after which every migrator — the Job included — fails with:

Can't take lock to run migrations: Migration table is already locked
If you are sure migrations are not running you can release the lock manually by running 'knex migrate:unlock'
Migration table is already locked
MigrationLocked: Migration table is already locked

The instance then stays down until the lock is cleared manually and a single migrator is allowed to finish.

Root cause

The chart never tells the backend pods to skip migrating when the Job is enabled.

templates/migrationJob.yaml — a pre-install,pre-upgrade hook Job that runs migrations once:

annotations:
  helm.sh/hook: pre-install,pre-upgrade
...
command:
  - pnpm
  - -F
  - backend
  - migrate-production

templates/backendDeployment.yaml:49 — the backend container command, with no awareness of migrationJob.enabled:

command: {{ .Values.image.command }}

image.command is unset by default, so the container falls through to the app image entrypoint, which migrates unconditionally before starting the server (docker/prod-entrypoint.sh in lightdash/lightdash):

# Migrate db
pnpm -F backend migrate-production
# Run prod
exec "$@"

So enabling migrationJob adds a migrator without removing the per-pod one. On a rolling upgrade you get 1 Job + N backend replicas all running knex migrate:latest against the same database:

helm upgrade
   |
   +-- pre-upgrade Job ----> migrate ----\
   +-- backend pod 1 -------> migrate -----+--> contend on knex_migrations_lock
   +-- backend pod 2 -------> migrate -----/      (killed mid-run => lock orphaned => all startups blocked)
   +-- ...

The worker template already does the right thing — it sets its own command (templates/_worker-deployment.tpl, node dist/scheduler.js) and never migrates — which is why only the backend pods are affected.

Affected versions

Every chart version since the pre-upgrade migration Job was introduced (#121). The race is independent of the Lightdash app version; it is triggered by the upgrade process whenever migrationJob.enabled=true.

Fix

When migrationJob.enabled, give the backend container an explicit command that starts the server without migrating, so the Job is the sole migrator. Minimal chart-only change in templates/backendDeployment.yaml:

command: {{ if .Values.migrationJob.enabled }}["dumb-init", "--", "node", "dist/index.js"]{{ else }}{{ .Values.image.command }}{{ end }}

(dumb-init is preserved as PID 1; node dist/index.js is the image's default CMD, so this is the normal boot minus the migrate step.)

A cleaner alternative is to add a SKIP_DB_MIGRATIONS-style env toggle to prod-entrypoint.sh in lightdash/lightdash and have the chart set it when migrationJob.enabled, so the chart sets an env var rather than overriding the whole command. Either approach closes the race; maintainers' preference on where the switch lives.

Workaround

On affected versions, override the backend command in values to skip the entrypoint migration (keep migrationJob.enabled: true):

image:
  command: '["dumb-init", "--", "node", "dist/index.js"]'

It must be the quoted JSON-string form — backendDeployment.yaml renders command bare (no toJson), so a plain YAML list renders as a single mashed argument. image.command is consumed only by the backend Deployment, so this does not affect workers. Do not override image.args, which is shared with the worker template.

Out of scope

  • The separate /api/v1/health connection-pool pressure during upgrades (migration-status query per probe) — tracked upstream in lightdash/lightdash; mitigate by pointing k8s probes at /api/v1/livez (no DB access).
  • Long, non-transactional migrations that can leave a partially-applied schema when interrupted — a different failure mode from the race itself.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions