fix(daemon): allow supervised restarts to become ready - #23
Conversation
Intent: - Stop reporting a failed daemon restart when launchd has accepted the exit and the replacement needs more than 15 seconds to answer. - Preserve headless recovery without sending operators toward unnecessary break-glass UI. Implementation: - Name the readiness deadline as daemonRestartTimeout. - Extend the post-acknowledgement readiness window from 15 to 30 seconds. - Derive the failure text from the same constant. Outcomes: - The CLI covers the 22-second replacement observed on Flagg instead of emitting a false failure. Verification: - go test ./cmd/secrets ./internal/daemon passed. - A live RPC restart replaced PID 349 with PID 98339; status then reported 258 secrets and uptime 0m. Follow-ups: - Host bootstrap still needs the separate joelclaw sudoers repair for genuinely unresponsive RPC.
|
Important Review skippedBot user detected. To trigger a single review, invoke the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Intent: - Stop healthy launchd processes from becoming intermittently unresponsive during the once-per-minute expired-lease sweep. - Keep status, list, lease, and restart RPCs responsive when audit persistence is slow. Implementation: - Remove and copy expired leases while holding the lease mutex. - Release the mutex before writing one durable audit entry per expiration and persisting the compacted lease set. - Add an injectable audit function and a concurrency regression that blocks audit logging while proving List remains responsive. Outcomes: - Audit fsync latency no longer serializes every RPC that reads the lease table. Verification: - go test -race ./internal/lease ./internal/daemon ./cmd/secrets passed. - Regression proves List returns while expiration audit logging is deliberately blocked. Follow-ups: - Deploy the rebuilt daemon through the root-owned service installer after merge; the running service binary is intentionally not replaced through an unprivileged path.
|
Added the underlying contention fix in The daemon was not crashing. The fix now removes/copies expired leases under the mutex, releases it, then performs audit and persistence work. A race-enabled regression blocks audit logging and proves Verification: |
Why
A live supervised restart on Flagg took about 22 seconds to become healthy. The CLI stopped waiting at 15 seconds and reported failure even though launchd had accepted the restart and brought up a healthy replacement.
That false failure sends agents toward break-glass recovery when routine headless restart already worked.
What changed
Verification
go test ./cmd/secrets ./internal/daemonuptime: 0mThe separate joelclaw host-bootstrap PR installs exact passwordless launchd recovery for a genuinely wedged RPC loop.