Skip to content

fix(service)!: share worker lifecycle for consensus shutdown - #1513

Open
lklimek wants to merge 5 commits into
v1.8-devfrom
fix/consensus-wait-shutdown
Open

lklimek wants to merge 5 commits into
v1.8-devfrom
fix/consensus-wait-shutdown

Conversation

@lklimek

@lklimek lklimek commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Consensus shutdown now waits for its receive loop and file workers before node teardown releases their resources. This is the focused lifecycle-core extraction from #1515.

Issue being fixed or feature implemented

Node shutdown could return while consensus still used its data directory, causing teardown failures such as directory not empty. A shared worker lifecycle replaces the consensus-only cancellation and WaitGroup workaround.

What was done?

  • Added attempt-scoped contexts, synchronized worker admission, cancellation and drain to BaseService. Optional OnDrain runs after registered work finishes.
  • Migrated consensus State, timeout ticker, WAL and autofile group; removed their redundant lifecycle fields and Wait overrides. Explicitly joined AutoFile timer/signal work.
  • Joined consensus handoffs so they cannot start State after reactor shutdown. A caller may cancel its wait after admission; the reactor continues to own and join startup. Node waits for reactors before closing event sinks and stores.
  • Preserved parent lifetimes for blocksync application during handover, connection error callbacks, and consensus processing already in flight during direct State.Stop. Parent cancellation still aborts this work.
  • Removed the node test's extra leaktest shutdown barrier. Added upgrade guidance, an in-tree cancellation-hook audit, explicit autofile Close ownership, and documented best-effort router shutdown notifications.

This PR is independent of #1515 and carries only its core. RPC/WebSocket ownership, mempool recheck drain, broad P2P/ABCI/reactor migrations and the broader durable-finalization/dependency-lifetime changes remain outside this PR. Plain goroutines are not automatically tracked; this does not promise that Node.Wait joins every repository helper.

How Has This Been Tested?

  • Confirmed red then green regressions for shared manual-stop cancellation, late consensus handoff and callback context preservation.
  • Existing blocksync handover regressions exposed the context compatibility issue and passed after its focused correction.
  • Affected packages passed go test -race -tags=deadlock -p 1: service, autofile, consensus, node, blocksync and P2P connections.
  • Both original node shutdown regressions passed 100 repetitions with GOMAXPROCS=2 and race/deadlock enabled. Late handoff passed 10 repetitions; existing blocksync handover cases passed 5 repetitions.
  • Goimports formatting and scoped golangci-lint passed with 0 issues.
  • Blocked FinalizeBlock and handoff cancellation regressions failed before the fixes and passed afterward. They verify durable state after direct State.Stop, parent cancellation, message metadata, and reactor ownership after a caller stops waiting.
  • Follow-up race/deadlock tests passed for consensus, node, blocksync, autofile, service, and the router's canceled-context peer cleanup. The lifecycle regression subset also passed 10 repetitions. Scoped lint passed with 0 issues.
  • CI and both e2e networks (dashcore, rotate) passed on 8e431fc. Follow-up CI results are pending; this remains a draft.

Breaking Changes

Manual Stop cancels the context passed to OnStart. Stop invokes OnStop synchronously but registered work is joined by Wait; callers must Stop then Wait before releasing worker resources. Hooks run outside the lifecycle lock and must not wait for their own registered work. TimeoutTicker implementations now require Wait. Failed startup may retry after worker cleanup; successfully stopped services cannot restart.

Checklist:

  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated relevant unit/integration/functional/e2e tests
  • I have made corresponding changes to the documentation

For repository code-owners and collaborators only

  • I have assigned this pull request to a milestone

🤖 Co-authored by Claudius the Magnificent AI Agent

PR Hygiene · b8a6beb

  • Bots — thepastaclaw ✓
  • Self-review — post /self-reviewed
  • Build failed
  • Approvals — you own every area touched; none needed

When every merge requirement is met, the PR Hygiene check passes. Reviewer limits do not block merging; other required GitHub checks and protections still apply.

Join the receive routine before State.Wait returns and cancel its context
on explicit Stop. This prevents node cleanup racing in-flight consensus
file writes. Add a regression test for the shutdown contract.

Co-Authored-By: OpenAI Codex GPT-6 <noreply@openai.com>

<sub>🤖 Co-authored by [Claudius the Magnificent](https://github.com/lklimek/claudius) AI Agent</sub>
@lklimek lklimek added the claudius-review Trigger Claudius AI code review on this PR label Sep 24, 2026
@coderabbitai

coderabbitai Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: dashpay/tenderdash/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: d8695057-091d-434c-a242-87ae6286c0ba

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@thepastaclaw

thepastaclaw commented Sep 24, 2026 •

Copy link
Copy Markdown

✅ Final review complete — no blockers (commit b8a6beb) · triage: normal

lklimek and others added 2 commits September 25, 2026 09:15
Port the focused lifecycle core from #1515, replacing the local State wait workaround while retaining other service ownership mechanisms.

Co-Authored-By: OpenAI GPT-6 <noreply@openai.com>
@lklimek lklimek changed the title fix(consensus): wait for receive routine during shutdown fix(service)!: share worker lifecycle for consensus shutdown Sep 25, 2026

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claudius review — verdict: Request changes. 15 findings across 3 reviewers (0 CRITICAL, 0 HIGH, 5 MEDIUM, 8 LOW, 2 INFO). One LOW-severity finding is classified blocking because it plausibly trips the G-GROWTH gate (unbounded resource growth). The fourth planned reviewer (project-reviewer-adams) was skipped per this pipeline's early-stop rule once the security reviewer's own gate returned a blocking candidate.

Blocking — please confirm before merge:

  • SEC-008 — internal/libs/autofile/group.go:82-86: OpenGroup's head AutoFile is now opened with context.WithoutCancel, so its goroutine, file descriptor, and SIGHUP registration are released only by an explicit Close(). In-tree WAL usage looks covered (traced every OnDrain/Wait call site, found no dropped Group), but if any rotation or reopen path ever skips Close(), this leaks per-occurrence and grows unboundedly over a long-running validator's life. Please confirm Close() is unconditionally reached on every group rotation before merging — this is exactly the kind of thing static reading alone can't fully settle.

Worth a maintainer's close look (non-blocking, but not to be waved away):

  • SEC-001 — the refactor now cancels a service's context before calling OnStop() (previously the reverse). This defeats the pre-PR safeguard that let an in-progress consensus ApplyCommit finish before the timeout ticker stopped, with no replacement guard added. Caveat, in fairness: the guard may already have been partly defeated by pre-existing Reactor.OnStop ordering, so this could narrow an existing gap rather than open a brand-new one — but it's exactly the kind of shutdown-correctness regression a PR titled "fix consensus shutdown" shouldn't be introducing.
  • CALL-001 — that same cancel-before-OnStop reordering is a contract change hitting every service.Implementation in the tree (~30 embedders), of which this PR audited and bridged only 2 (blocksync synchronizer, p2p connection). SEC-006/SEC-007 found a third unmigrated call site (p2p/router.go's routePeer) with a real, if low-impact-in-production, behavioral delta.
  • DOC-001 — this is a breaking change (fix(service)!:) to a public package's runtime contract with no UPGRADING.md entry.
  • SEC-002 — an unchecked type assertion in blocksync/synchronizer.go:255 that would panic if OnStart is ever invoked outside its Start() wrapper; the same PR already uses the safe comma-ok idiom for the equivalent case in connection.go, so this is inconsistency more than a proven live bug.
  • SEC-003 — SwitchToConsensus now blocks uninterruptibly because the caller's context is discarded after admission; a wedged WAL replay can no longer be cancelled by its caller.

Also flagged (LOW/INFO, no action required to merge): two stale godocs left behind by the refactor (ticker.go, wal.go), an architecture-doc gap on the Stop-during-startup case, an unsynchronized field write in State.OnStart's rollback path, a nil-channel-block edge case in the new Stopping(), and — credit where due — the new lifecycle_test.go suite is genuinely solid engineering, not padding.

Full detail in the report.json/report.html CI artifacts.
📊 View full HTML review report

Comment thread libs/service/service.go
Comment thread libs/service/service.go
Comment thread internal/consensus/state.go
Comment thread internal/blocksync/synchronizer.go Outdated
Comment thread internal/consensus/reactor.go
@github-actions github-actions Bot removed the claudius-review Trigger Claudius AI code review on this PR label Sep 25, 2026
@lklimek

lklimek commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator Author

Checked all five open threads and the full 15-finding report against 8e431fc. No code changes in this verification pass. The five threads remain open; individual assessments are in replies.

The reported rotation leak is not substantiated: Group.rotateFile retains the same AutoFile, closes its current descriptor under lock, and reopens lazily. It does not allocate a new worker or SIGHUP registration per rotation. WAL OnDrain calls Group.Close; failed WAL startup and repair also close/join their old group. Explicit Close ownership should be documented, and logjack should defer Close, but its process-exit cleanup gap does not establish unbounded validator growth or a merge blocker.

Remaining small items: stale ticker/WAL hook comments and the architecture document's missing startup-stop exception are valid. The synchronizer comment should say parent lifetime: the captured parent can be the blocksync reactor, so manual reactor shutdown also cancels application.

The rollback data-race claim is not demonstrated by a concurrent reader: startup failure precedes receiveRoutine admission, and BaseService defers OnStop while startup is active. A custom successfully stopped ticker must be recreated for retry, as the architecture document already requires for composite children. A lock around assignments alone would not establish reader synchronization. The Stopping nil-channel behavior is documented; a blocked ScheduleTimeout on an unstarted ticker is an API edge case requiring a concrete supported call path before treating it as a regression.

The substantive remaining concerns are the global cancellation-contract audit (including router shutdown notifications), in-flight consensus finalization policy, handoff caller-cancellation policy, and upgrade guidance. Static verification only; no new tests were run.

🤖 Co-authored by Claudius the Magnificent AI Agent

Keep processing on the State parent's lifetime while canceling intake, and
let handoff callers cancel their wait without abandoning reactor-owned work.
Document the global cancellation contract, cleanup ownership and migration
boundaries; harden context lookup and close logjack's group on return.

Co-Authored-By: OpenAI GPT-6 <noreply@openai.com>
@lklimek

lklimek commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the actionable feedback in ec8d306:

  • Direct State.Stop cancels intake while already-running consensus processing retains the State parent's context. The regression blocks actual FinalizeBlock, verifies Wait stays blocked, then verifies the committed height is persisted. Parent cancellation still aborts finalization; reactor/node shutdown may require WAL recovery. This deliberately does not pull in the detached dependency graph from fix(service)!: join background work before shutdown completes #1515.
  • Canceling a handoff caller releases its wait; admitted startup remains reactor-owned. The regression verifies both that startup remains live after caller cancellation and that reactor shutdown joins it.
  • Added UPGRADING.md migration guidance and an inventory of all in-tree stop-hook groups. Router bookkeeping cleanup still runs, while shutdown broadcasts are explicitly best effort; existing canceled-context peer-manager tests pass.
  • Made the synchronizer context assertion safe, corrected hook/startup documentation, documented explicit autofile Close ownership, and added logjack's deferred Close.

The two main regressions were red before the fix, then green. Affected package tests pass with race/deadlock enabled; the lifecycle regression subset passed 10 repetitions. Final consensus/node tests and scoped lint pass (0 lint issues). New CI is pending.

The speculative per-rotation resource-growth finding is not reproduced: rotation reuses one AutoFile and closes its descriptor; final Close joins its worker. No speculative rollback locks or Stopping contract changes were introduced.

🤖 Co-authored by Claudius the Magnificent AI Agent

@lklimek
lklimek marked this pull request as ready for review September 30, 2026 12:06
@github-actions github-actions Bot added the waiting-bots Waiting for the review bots to report on this head label Sep 30, 2026

@thepastaclaw thepastaclaw left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Final validation — Phase 1 + Phase 2

The lifecycle changes correctly coordinate worker admission, cancellation, draining, consensus handoff ownership, and file cleanup within the PR's stated partial-migration scope. Source inspection and the shutdown regressions support the supplied clean findings; no actionable in-scope defects were confirmed. All seven affected packages passed a fresh race/deadlock-enabled test run using a compatible BLS build, and the worktree remains unchanged.

Review provenance

Source: reviewer 1: muse-spark-1.3-contributor (agent: phase1-reviewer, role: general); reviewer 2: muse-spark-1.3-contributor (agent: phase1-reviewer, role: tenderdash-consensus-security); reviewer 3: gpt-6.1-sol (agent: phase2-reviewer, role: general); reviewer 4: gpt-6.1-sol (agent: phase2-reviewer, role: tenderdash-consensus-security); final verifier: gpt-6.1-sol (agent: sol-verifier, role: final-verifier)

  • Triage: normal by gpt-6.1-sol (effort low) — The diff introduces intricate, cross-cutting worker lifecycle and shutdown synchronization changes in libs/service/service.go and consensus integration, but does not change consensus rules, cryptography, funds handling, peer-facing deserialization, or storage migrations.
  • Phase 1 reviewers: muse-spark-1.3-contributor — general (completed, effort xhigh); agent phase1-reviewer, muse-spark-1.3-contributor — tenderdash-consensus-security (completed, effort xhigh); agent phase1-reviewer
  • Phase 1 model: muse-spark-1.3-contributor — not quota-gated; passed over gemini-3.8-flash-high (antigravity below 15% reserve: weekly 13% left, 5h 100% left), glm-5.3-flash (not used above high effort; tier asks max)
  • Fresh verifier: gpt-6.1-sol — final-verifier; agent sol-verifier
  • Phase 2 reviewers: gpt-6.1-sol — general (completed, effort high); agent phase2-reviewer, gpt-6.1-sol — tenderdash-consensus-security (completed, effort high); agent phase2-reviewer

@github-actions

Copy link
Copy Markdown

Bots are done — your move: post /self-reviewed.
Full checklist in the description.

@github-actions github-actions Bot added waiting-self-review Waiting for the author to post /self-reviewed and removed waiting-bots Waiting for the review bots to report on this head labels Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-self-review Waiting for the author to post /self-reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants