Bug
When migrationJob.enabled=true, the backend Deployment pods still run database migrations on startup, in addition to the dedicated pre-upgrade migration Job. During a helm upgrade this means two or more processes (the Job plus every rolling backend pod) attempt to migrate the same database concurrently. They contend on the Knex migration lock, and if a migrator is killed mid-migration the lock is left held, which then blocks every subsequent startup.
Steps to reproduce
- Deploy the chart with
migrationJob.enabled=true (and no image.command override).
- Run
helm upgrade to a release that applies a non-trivial migration, against a database large enough that the migration takes more than a few seconds.
- The pre-upgrade
*-migrate Job starts migrating; meanwhile the rolling backend pods start and their entrypoint also runs knex migrate:latest.
- Observe: backend pods crash-loop, and the logs show the lock is held.
Expected behavior
When the migration Job is enabled, it should be the only process that runs migrations. Backend pods should start the server without attempting to migrate, so there is exactly one migrator per upgrade.
Actual behavior
Multiple migrators race for the lock. The losers fail and (because the entrypoint uses set -e) the pod exits non-zero and crash-loops. If a migrator is killed mid-run the lock is orphaned, after which every migrator — the Job included — fails with:
Can't take lock to run migrations: Migration table is already locked
If you are sure migrations are not running you can release the lock manually by running 'knex migrate:unlock'
Migration table is already locked
MigrationLocked: Migration table is already locked
The instance then stays down until the lock is cleared manually and a single migrator is allowed to finish.
Root cause
The chart never tells the backend pods to skip migrating when the Job is enabled.
templates/migrationJob.yaml — a pre-install,pre-upgrade hook Job that runs migrations once:
annotations:
helm.sh/hook: pre-install,pre-upgrade
...
command:
- pnpm
- -F
- backend
- migrate-production
templates/backendDeployment.yaml:49 — the backend container command, with no awareness of migrationJob.enabled:
command: {{ .Values.image.command }}
image.command is unset by default, so the container falls through to the app image entrypoint, which migrates unconditionally before starting the server (docker/prod-entrypoint.sh in lightdash/lightdash):
# Migrate db
pnpm -F backend migrate-production
# Run prod
exec "$@"
So enabling migrationJob adds a migrator without removing the per-pod one. On a rolling upgrade you get 1 Job + N backend replicas all running knex migrate:latest against the same database:
helm upgrade
|
+-- pre-upgrade Job ----> migrate ----\
+-- backend pod 1 -------> migrate -----+--> contend on knex_migrations_lock
+-- backend pod 2 -------> migrate -----/ (killed mid-run => lock orphaned => all startups blocked)
+-- ...
The worker template already does the right thing — it sets its own command (templates/_worker-deployment.tpl, node dist/scheduler.js) and never migrates — which is why only the backend pods are affected.
Affected versions
Every chart version since the pre-upgrade migration Job was introduced (#121). The race is independent of the Lightdash app version; it is triggered by the upgrade process whenever migrationJob.enabled=true.
Fix
When migrationJob.enabled, give the backend container an explicit command that starts the server without migrating, so the Job is the sole migrator. Minimal chart-only change in templates/backendDeployment.yaml:
command: {{ if .Values.migrationJob.enabled }}["dumb-init", "--", "node", "dist/index.js"]{{ else }}{{ .Values.image.command }}{{ end }}
(dumb-init is preserved as PID 1; node dist/index.js is the image's default CMD, so this is the normal boot minus the migrate step.)
A cleaner alternative is to add a SKIP_DB_MIGRATIONS-style env toggle to prod-entrypoint.sh in lightdash/lightdash and have the chart set it when migrationJob.enabled, so the chart sets an env var rather than overriding the whole command. Either approach closes the race; maintainers' preference on where the switch lives.
Workaround
On affected versions, override the backend command in values to skip the entrypoint migration (keep migrationJob.enabled: true):
image:
command: '["dumb-init", "--", "node", "dist/index.js"]'
It must be the quoted JSON-string form — backendDeployment.yaml renders command bare (no toJson), so a plain YAML list renders as a single mashed argument. image.command is consumed only by the backend Deployment, so this does not affect workers. Do not override image.args, which is shared with the worker template.
Out of scope
- The separate
/api/v1/health connection-pool pressure during upgrades (migration-status query per probe) — tracked upstream in lightdash/lightdash; mitigate by pointing k8s probes at /api/v1/livez (no DB access).
- Long, non-transactional migrations that can leave a partially-applied schema when interrupted — a different failure mode from the race itself.
Bug
When
migrationJob.enabled=true, the backend Deployment pods still run database migrations on startup, in addition to the dedicated pre-upgrade migration Job. During ahelm upgradethis means two or more processes (the Job plus every rolling backend pod) attempt to migrate the same database concurrently. They contend on the Knex migration lock, and if a migrator is killed mid-migration the lock is left held, which then blocks every subsequent startup.Steps to reproduce
migrationJob.enabled=true(and noimage.commandoverride).helm upgradeto a release that applies a non-trivial migration, against a database large enough that the migration takes more than a few seconds.*-migrateJob starts migrating; meanwhile the rolling backend pods start and their entrypoint also runsknex migrate:latest.Expected behavior
When the migration Job is enabled, it should be the only process that runs migrations. Backend pods should start the server without attempting to migrate, so there is exactly one migrator per upgrade.
Actual behavior
Multiple migrators race for the lock. The losers fail and (because the entrypoint uses
set -e) the pod exits non-zero and crash-loops. If a migrator is killed mid-run the lock is orphaned, after which every migrator — the Job included — fails with:The instance then stays down until the lock is cleared manually and a single migrator is allowed to finish.
Root cause
The chart never tells the backend pods to skip migrating when the Job is enabled.
templates/migrationJob.yaml— apre-install,pre-upgradehook Job that runs migrations once:templates/backendDeployment.yaml:49— the backend container command, with no awareness ofmigrationJob.enabled:image.commandis unset by default, so the container falls through to the app image entrypoint, which migrates unconditionally before starting the server (docker/prod-entrypoint.shinlightdash/lightdash):So enabling
migrationJobadds a migrator without removing the per-pod one. On a rolling upgrade you get1 Job + N backend replicasall runningknex migrate:latestagainst the same database:The worker template already does the right thing — it sets its own command (
templates/_worker-deployment.tpl,node dist/scheduler.js) and never migrates — which is why only the backend pods are affected.Affected versions
Every chart version since the pre-upgrade migration Job was introduced (#121). The race is independent of the Lightdash app version; it is triggered by the upgrade process whenever
migrationJob.enabled=true.Fix
When
migrationJob.enabled, give the backend container an explicit command that starts the server without migrating, so the Job is the sole migrator. Minimal chart-only change intemplates/backendDeployment.yaml:(
dumb-initis preserved as PID 1;node dist/index.jsis the image's defaultCMD, so this is the normal boot minus the migrate step.)A cleaner alternative is to add a
SKIP_DB_MIGRATIONS-style env toggle toprod-entrypoint.shinlightdash/lightdashand have the chart set it whenmigrationJob.enabled, so the chart sets an env var rather than overriding the whole command. Either approach closes the race; maintainers' preference on where the switch lives.Workaround
On affected versions, override the backend command in values to skip the entrypoint migration (keep
migrationJob.enabled: true):It must be the quoted JSON-string form —
backendDeployment.yamlrenderscommandbare (notoJson), so a plain YAML list renders as a single mashed argument.image.commandis consumed only by the backend Deployment, so this does not affect workers. Do not overrideimage.args, which is shared with the worker template.Out of scope
/api/v1/healthconnection-pool pressure during upgrades (migration-status query per probe) — tracked upstream inlightdash/lightdash; mitigate by pointing k8s probes at/api/v1/livez(no DB access).