Problem
A running sync all can still starve the entire dbrain process. In the investigation for #159, a zero-work /favicon.ico request measured approximately 17.2 seconds while sync was active. The delay is not explained by the request's database work and can affect HTTP and MCP responsiveness.
This is the remaining Channel B — process starvation problem. It is distinct from the deferred link-capture, SQLite-pool, and synchronous-feed-fetch mitigations delivered in #162.
Goal
Keep HTTP and MCP request latency bounded and responsive while sync all is running by identifying and removing or bounding the process-wide CPU, GC, scheduler, or tsnet starvation mechanism.
Investigation
- Reproduce the stall with a production-shaped
sync all and concurrent no-work HTTP, API, and MCP requests.
- Capture request latency distributions and correlate them with CPU saturation, Go scheduler/runtime metrics, GC pauses, goroutine counts, memory pressure, network activity, and tsnet behavior.
- Attribute the delay to the actual sync stage and mechanism before choosing a mitigation; do not assume SQLite contention or the semantic lease is the cause.
- Test the fix under sustained sync load, including the stages and local model/network configuration that reproduce the original stall.
Acceptance criteria
- The root cause is documented with reproducible evidence and a clear distinction between database/lease wait and process-wide starvation.
- HTTP and MCP no-work/read-only probes remain responsive throughout
sync all, with the prior multi-second starvation outlier eliminated or materially bounded and any remaining limit explicitly documented.
- The verification includes before/after latency evidence plus runtime/GC/CPU/tsnet observations from the same workload.
- The mitigation does not bypass authoritative write coordination or silently weaken sync correctness.
- Add regression or diagnostic coverage that prevents the starvation mechanism from returning unnoticed.
Related work
Problem
A running
sync allcan still starve the entire dbrain process. In the investigation for #159, a zero-work/favicon.icorequest measured approximately 17.2 seconds while sync was active. The delay is not explained by the request's database work and can affect HTTP and MCP responsiveness.This is the remaining Channel B — process starvation problem. It is distinct from the deferred link-capture, SQLite-pool, and synchronous-feed-fetch mitigations delivered in #162.
Goal
Keep HTTP and MCP request latency bounded and responsive while
sync allis running by identifying and removing or bounding the process-wide CPU, GC, scheduler, or tsnet starvation mechanism.Investigation
sync alland concurrent no-work HTTP, API, and MCP requests.Acceptance criteria
sync all, with the prior multi-second starvation outlier eliminated or materially bounded and any remaining limit explicitly documented.Related work