fix(gateway): bound busy channel health by real run age#103793
Conversation
3e2d4a2 to
fdac98d
Compare
|
Codex review: needs real behavior proof before merge. Reviewed July 18, 2026, 9:21 PM ET / July 19, 2026, 01:21 UTC. Summary PR surface: Source +41, Tests +125. Total +166 across 7 files. Reproducibility: yes. from source: a live run-state heartbeat refreshes Review metrics: none identified. Merge readiness Overall follows the weaker of proof and patch quality, so missing proof can cap an otherwise strong patch. Rank-up moves:
Proof guidance:
Risk before merge
Maintainer options:
Next step before merge
Security Review detailsBest possible solution: Keep the disconnected-only age bound, then add redacted current-head monitor output showing a hung disconnected run becomes Do we have a high-confidence way to reproduce the issue? Yes, from source: a live run-state heartbeat refreshes Is this the best way to solve the issue? Yes, conditionally: surfacing the oldest queue-owned run start and consuming it in the shared health policy is the narrowest shared fix, and the disconnected-only guard preserves the established no-run-timeout contract. Current-head real behavior proof is still required to validate that boundary before merge. AGENTS.md: found and applied where relevant. Codex review notes: model internal, reasoning high; reviewed against c3adaa3195bd. Label changesLabel changes:
Label justifications:
Evidence reviewedPR surface: Source +41, Tests +125. Total +166 across 7 files. View PR surface stats
What I checked:
Likely related people:
What the crustacean ranks mean
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics. How this review workflow works
Review history (11 earlier review cycles; latest 8 shown)
|
|
@clawsweeper re-review Refreshed the Real behavior proof to address the prior proof P1s (injected clock, no real recovery path shown). The proof now drives the real periodic health monitor loop ( On identical live input, pinned to head fdac98d vs pristine main 091584b:
The 25 minute busy ceiling and 60 second heartbeat were both scaled down 300x on real elapsed time so the breach happens in seconds rather than 26 minutes; identical scaling on both runs. Full block is in the PR body under Real behavior proof. |
|
🦞🧹 I asked ClawSweeper to review this item again. Re-review progress:
|
fdac98d to
21284cf
Compare
|
@clawsweeper re-review Addressed the rank-up move on lifecycle compatibility. onRunEnd() remains source-compatible for released zero-argument plugin consumers, while the internal channel queue passes its additive run handle so concurrent runs still retire exactly. Added coverage for the zero-argument callback path and retained the overlapping-run case. Rebased onto current main at 5bb5e4f and refreshed the proof head to 21284cf. Focused lifecycle, gateway health-policy, and SDK subpath suites pass; scoped formatting and core type checks are clean. |
|
🦞🧹 I asked ClawSweeper to review this item again. Re-review progress:
|
21284cf to
bd8c71b
Compare
|
@clawsweeper re-review The current head separates anonymous zero-argument lifecycle accounting from handle-aware run-age tracking. The shared queue retains exact handles; legacy concurrent completions retain busy counts but never publish a guessed active-run start. Added coverage for anonymous concurrency and both exact completion orders. All GitHub checks are green on bd8c71b. |
|
🦞🧹 I asked ClawSweeper to review this item again. Re-review progress:
|
|
@clawsweeper re-run |
|
🦞🧹 I asked ClawSweeper to review this item again. |
The channel health policy treats a channel as healthy-busy even while disconnected, bounded only by a 25 minute stale ceiling measured from lastRunActivityAt. The run-state heartbeat refreshes lastRunActivityAt every 60 seconds for as long as any run is active, so a run that hangs forever (for example a send blocking on a dead socket after the transport already reported connected:false) keeps that timestamp fresh and the stuck ceiling is never reached. The account is then reported healthy forever by the health monitor, readiness probe, and health CLI, and no restart ever fires. createRunStateMachine now tracks each in-flight run's start time keyed by an opaque run handle and publishes the oldest still-active run's start as activeRunStartedAt. The health policy busy override keys its ceiling off the real run age, so a run stuck longer than the threshold reports stuck and the monitor can restart it. Because the reported start is the oldest active run and advances to the next-oldest as runs complete, a channel churning through many short overlapping runs (activeRuns above 1 across concurrent queue keys) stays healthy; only a genuinely hung run breaches the ceiling. Short and active runs stay healthy and the existing lastRunActivityAt fallback is preserved for snapshots without a start time.
Keep the released zero-argument onRunEnd callback source-compatible while allowing internal queue callers to pass a run handle for exact concurrent-run accounting. The compatibility path closes the oldest active run, preserving existing lifecycle behavior for consumers that do not use handles.
The zero-argument lifecycle callbacks cannot identify which concurrent run completed, so they must not update the identity-sensitive run start used by channel health. Keep their busy count separately and reserve exact start tracking for the shared queue's handle-aware lifecycle path.
Keep the public run-state lifecycle callbacks unchanged. The channel queue now owns opaque run identity and augments its status updates with the oldest active queue run, so implementation details do not expand the SDK surface.
Keep activeRunStartedAt in the internal status patch type so the queue can publish its private tracked-run age through the existing status sink.
e157eaa to
d5f25fa
Compare
|
@clawsweeper re-review Addressed the remaining P2 patch-quality finding. Refreshed onto current
|
|
🦞🧹 I asked ClawSweeper to review this item again. Re-review progress:
|
|
Blocking correctness issue: Please constrain the run-age ceiling to cases where busy is masking an independently unhealthy transport condition (at minimum,
|
|
@fuller-stack-dev Agreed. As pushed, the ceiling applied run age to healthy transports, which conflicts with the queue's no-run-timeout contract. Fixed in a1af976. The busy branch now derives the run-age term only when the snapshot reports Regression coverage for the three cases you listed:
Refreshed real behavior proof against this head, including a live connected-long-run control, will follow in the PR body shortly. |
|
Land-ready proof refreshed for exact head
Behavior proof now covers the requested boundary: connected or transport-unknown long runs remain healthy/busy; an explicitly disconnected old run becomes stuck; concurrent completion advances to the next-oldest run. Known gap: no live provider socket was held hung for 25 real minutes; the changed decision and monitor action are covered deterministically. |
|
Merged via squash.
|
) * fix(gateway): bound busy channel health by real run age The channel health policy treats a channel as healthy-busy even while disconnected, bounded only by a 25 minute stale ceiling measured from lastRunActivityAt. The run-state heartbeat refreshes lastRunActivityAt every 60 seconds for as long as any run is active, so a run that hangs forever (for example a send blocking on a dead socket after the transport already reported connected:false) keeps that timestamp fresh and the stuck ceiling is never reached. The account is then reported healthy forever by the health monitor, readiness probe, and health CLI, and no restart ever fires. createRunStateMachine now tracks each in-flight run's start time keyed by an opaque run handle and publishes the oldest still-active run's start as activeRunStartedAt. The health policy busy override keys its ceiling off the real run age, so a run stuck longer than the threshold reports stuck and the monitor can restart it. Because the reported start is the oldest active run and advances to the next-oldest as runs complete, a channel churning through many short overlapping runs (activeRuns above 1 across concurrent queue keys) stays healthy; only a genuinely hung run breaches the ceiling. Short and active runs stay healthy and the existing lastRunActivityAt fallback is preserved for snapshots without a start time. * fix(channels): retain run-state callback compatibility Keep the released zero-argument onRunEnd callback source-compatible while allowing internal queue callers to pass a run handle for exact concurrent-run accounting. The compatibility path closes the oldest active run, preserving existing lifecycle behavior for consumers that do not use handles. * fix(channels): keep anonymous runs out of age tracking The zero-argument lifecycle callbacks cannot identify which concurrent run completed, so they must not update the identity-sensitive run start used by channel health. Keep their busy count separately and reserve exact start tracking for the shared queue's handle-aware lifecycle path. * fix(channels): keep tracked runs internal Keep the public run-state lifecycle callbacks unchanged. The channel queue now owns opaque run identity and augments its status updates with the oldest active queue run, so implementation details do not expand the SDK surface. * fix(channels): type queue run start status Keep activeRunStartedAt in the internal status patch type so the queue can publish its private tracked-run age through the existing status sink. * fix(channels): wrap isActive to satisfy unbound-method lint * fix(gateway): gate busy run-age ceiling on disconnected transport
* fix: gate diagnostics command to owners (cherry picked from commit 170bf72) * fix(agent): replace self-wait with deferred release in retained-lock abort cleanup (#96100) * fix(agent): wait for retained session write before releasing held lock on abort * fix(agent): replace self-wait with deferred release in retained-lock abort cleanup * fix(test): reject fallback acquire with SessionWriteLockTimeoutError in active-scope cleanup test * fix(agent): trim retained-lock comments Signed-off-by: sallyom <[email protected]> --------- Signed-off-by: sallyom <[email protected]> Co-authored-by: sallyom <[email protected]> (cherry picked from commit 0a042f6) * fix(gateway): resume channel after pending task recovery (cherry picked from commit 6039da3) * fix(gateway): resume channel after pending task recovery (cherry picked from commit ecd29fe) * fix(outbound): ignore empty delivery receipts (#79811) (cherry picked from commit 9a735be) * fix(agents): guard delivery-evidence attachment recursion against cycles (#97041) * fix(agents): guard delivery-evidence attachment recursion against cycles * fix(agents): guard delivery-evidence attachment recursion against cycles * fix(agents): guard delivery-evidence attachment recursion against cycles --------- Co-authored-by: Pick-cat <[email protected]> Co-authored-by: Vincent Koc <[email protected]> (cherry picked from commit 4985671) * fix(opencode-go): re-arm idle timer on block-boundary events to prevent false stalled-stream abort (#97128) * fix(opencode-go): re-arm idle timer on block-boundary events to prevent false stalled-stream abort When the opencode-go model finalizes a tool call and deliberates before the next one, the provider emits real block-boundary SSE events (text_end, thinking_end, toolcall_start, toolcall_end) that prove the socket is alive, but the watchdog's isProviderProgressEvent only returned true for token deltas (text_delta, thinking_delta, toolcall_delta). This caused the idle timer to fire and falsely abort a live stream, replacing a completed answer with a stalled error and dropping the provider's real done event. Fix: include block-boundary events in isProviderProgressEvent so the idle timer is re-armed on any forward-progress provider event. text_start and thinking_start are intentionally excluded because they are synthetic preamble events that should not shorten the first-event window. Closes #96518 Co-Authored-By: Claude Opus 4.8 <[email protected]> * test(opencode-go): satisfy lint in stream regression * test(opencode-go): satisfy lint in stream regression * test(opencode-go): satisfy lint in stream regression --------- Co-authored-by: Claude Opus 4.8 <[email protected]> Co-authored-by: Vincent Koc <[email protected]> (cherry picked from commit 552ec2b) * fix(model-fallback): don't rethrow provider-side AbortErrors as user cancellations (#90908) * fix(model-fallback): don't rethrow provider-side AbortErrors as user cancellations When the LLM API closes the connection mid-stream, the fetch layer surfaces AbortError("This operation was aborted") with no external abort signal triggered. The old guard `shouldRethrowAbort()` returned false for these errors (because isTimeoutError matched the message), so they fell through to the fallback loop but were never retried — the error propagated up and produced SILENT_REPLY_TOKEN in group sessions, permanently silencing the topic. Replace the guard with a direct check: only rethrow AbortError when the external abort signal is actually set (user/gateway cancellation). Provider-side AbortErrors without an external signal now fall through to the next fallback candidate, giving the system a chance to recover. * fix(cron): forward abort signal into runWithModelFallback Thread the cron executor's abort signal into the shared runWithModelFallback call so that cron timeouts and cancellations stop the fallback chain instead of retrying with the next candidate. Previously, the run callback checked params.abortSignal?.aborted and threw, but runWithModelFallback itself had no signal — so the new guard in model-fallback.ts could not distinguish a caller abort from a provider-side AbortError and would retry silently. Also adds a focused regression test verifying the signal is forwarded. --------- Co-authored-by: Shengting Xie <[email protected]> Co-authored-by: yayu <[email protected]> (cherry picked from commit 98ed83f) * fix(browser): block node routes when sandbox host control is disabled (#97958) (cherry picked from commit 2cf765f) * fix(exec): bind Windows allowlist execution path (#98260) * fix(exec): bind windows allowlist execution path * fix(exec): add windows shadow execution proof * fix(exec): preserve wildcard allowlist behavior * fix(exec): correct blocked plan test fixture (cherry picked from commit 3811001) * fix(mcp): suppress unhandled error on stderr pipe in stdio transport (#99803) * fix(mcp): suppress unhandled error on stderr pipe in stdio transport When child.stderr is piped to stderrStream without an error handler, a stream-level error (EPIPE, I/O failure) crashes the process. Add a noop error handler before the pipe, consistent with the error handlers already present on stdin and stdout. Co-Authored-By: Claude <[email protected]> * test(mcp): add regression test for stderr pipe error suppression Co-Authored-By: Claude <[email protected]> * fix(mcp): report stderr stream errors * fix(mcp): report stderr stream errors --------- Co-authored-by: Claude <[email protected]> Co-authored-by: Vincent Koc <[email protected]> (cherry picked from commit 1b84316) * Harden macOS SQLite WAL checkpoints (#99067) (cherry picked from commit f7f1be2) * fix(secrets): suppress unhandled stdout/stderr stream errors in exec resolver (#100521) * fix(secrets): suppress unhandled stdout/stderr stream errors in exec resolver * proof(secrets): add real behavior proof script for exec resolver stream error catch * proof(secrets): replace wrapper with real exec resolver stream error proof * style: apply oxfmt to changed files (cherry picked from commit c9a0783) * fix(agents): retry transient filesystem races when reading workspace bootstrap files (#100910) * fix(agents): retry transient filesystem races when reading workspace bootstrap files * fix(agents): retry transient boundary resolution --------- Co-authored-by: Vincent Koc <[email protected]> (cherry picked from commit f36d170) * fix(gateway): finish plugin HTTP responses after post-header failures (#102125) * fix(gateway): finish plugin HTTP responses after post-header failures * test(gateway): satisfy plugin HTTP regression lint * fix(gateway): skip ending destroyed plugin responses --------- Co-authored-by: Peter Steinberger <[email protected]> (cherry picked from commit 240d350) * fix(gateway): validate exact custom browser origins (#38290) Co-authored-by: Peter Steinberger <[email protected]> (cherry picked from commit fa0349a) * fix: block unspecified trusted DNS targets (#103075) (cherry picked from commit c70f3d0) * fix(channels): make nack callbacks idempotent (#104919) * fix(channels): make nack callbacks idempotent * fix(channels): coalesce overlapping nack callbacks --------- Co-authored-by: Peter Steinberger <[email protected]> (cherry picked from commit 02d307e) * fix(channels): prevent base URL credentials in status output (#107754) * fix(channels): redact credentials in account URLs * fix(channels): sanitize final status summaries (cherry picked from commit 210340f) * fix(channels): prevent lifecycle listener buildup (#109108) (cherry picked from commit 0e1fad7) * fix(sandbox): use Buffer.byteLength for env var value size limit (#105017) * fix(sandbox): use Buffer.byteLength for env var value size limit validateEnvVarValue checked value.length (UTF-16 code units) against the 32768-byte limit, so multi-byte CJK values like "值".repeat(11000) passed the check despite exceeding 33 KB in UTF-8. Switch to Buffer.byteLength(value, "utf8") so the limit matches the actual byte count the OS and child processes see. * test(sandbox): simplify env byte-limit coverage Co-authored-by: 唐梓夷0668001293 <[email protected]> --------- Co-authored-by: Peter Steinberger <[email protected]> (cherry picked from commit 84fb48c) * fix(gateway): guard process.kill ESRCH race in signalVerifiedGatewayPidSync (#109590) * fix(gateway): guard process.kill ESRCH race in signalVerifiedGatewayPidSync A verified gateway process can exit between the argv validation check and the process.kill call, causing an unhandled ESRCH error. Wrap the kill in try-catch and silently swallow ESRCH (process already gone = signal already delivered). Co-Authored-By: Claude Sonnet 4.6 <[email protected]> * docs(gateway): explain ESRCH signal race Co-authored-by: 丁宇婷0668001435 <[email protected]> --------- Co-authored-by: Claude Sonnet 4.6 <[email protected]> Co-authored-by: Peter Steinberger <[email protected]> (cherry picked from commit 853b1a8) * fix(litellm): guard loopback hostname auto-allow with isIP to prevent DNS SSRF bypass (#110693) * fix(litellm): guard loopback hostname auto-allow with isIP to prevent DNS bypass The isAutoAllowedLitellmHostname helper auto-enables private-network access for loopback-style hosts. Before this fix, lowered.startsWith("127.") matched DNS hostnames like 127.evil.com, letting remote endpoints bypass the explicit allowPrivateNetwork opt-in — a SSRF risk. Add isIP(host)===4 guard so only literal IPv4 loopback addresses qualify. Same canonical pattern as extensions/slack/src/monitor/relay-source.ts:271 and the codex loopback fix. Co-Authored-By: Claude <[email protected]> * test(litellm): cover loopback endpoint policy --------- Co-authored-by: Claude <[email protected]> Co-authored-by: Peter Steinberger <[email protected]> (cherry picked from commit 3d03b60) * fix(discord): sustained gateway bursts stop growing memory (#110954) * fix(discord): sustained gateway bursts stop growing memory * fix(discord): contain gateway queue overflow * fix(discord): drop oldest saturated gateway sends Co-authored-by: 张贵萍0668001030 <[email protected]> * fix(discord): surface gateway overflow warnings Co-authored-by: 张贵萍0668001030 <[email protected]> --------- Co-authored-by: Peter Steinberger <[email protected]> (cherry picked from commit 69aeba9) * fix(gateway): bound busy channel health by real run age (#103793) * fix(gateway): bound busy channel health by real run age The channel health policy treats a channel as healthy-busy even while disconnected, bounded only by a 25 minute stale ceiling measured from lastRunActivityAt. The run-state heartbeat refreshes lastRunActivityAt every 60 seconds for as long as any run is active, so a run that hangs forever (for example a send blocking on a dead socket after the transport already reported connected:false) keeps that timestamp fresh and the stuck ceiling is never reached. The account is then reported healthy forever by the health monitor, readiness probe, and health CLI, and no restart ever fires. createRunStateMachine now tracks each in-flight run's start time keyed by an opaque run handle and publishes the oldest still-active run's start as activeRunStartedAt. The health policy busy override keys its ceiling off the real run age, so a run stuck longer than the threshold reports stuck and the monitor can restart it. Because the reported start is the oldest active run and advances to the next-oldest as runs complete, a channel churning through many short overlapping runs (activeRuns above 1 across concurrent queue keys) stays healthy; only a genuinely hung run breaches the ceiling. Short and active runs stay healthy and the existing lastRunActivityAt fallback is preserved for snapshots without a start time. * fix(channels): retain run-state callback compatibility Keep the released zero-argument onRunEnd callback source-compatible while allowing internal queue callers to pass a run handle for exact concurrent-run accounting. The compatibility path closes the oldest active run, preserving existing lifecycle behavior for consumers that do not use handles. * fix(channels): keep anonymous runs out of age tracking The zero-argument lifecycle callbacks cannot identify which concurrent run completed, so they must not update the identity-sensitive run start used by channel health. Keep their busy count separately and reserve exact start tracking for the shared queue's handle-aware lifecycle path. * fix(channels): keep tracked runs internal Keep the public run-state lifecycle callbacks unchanged. The channel queue now owns opaque run identity and augments its status updates with the oldest active queue run, so implementation details do not expand the SDK surface. * fix(channels): type queue run start status Keep activeRunStartedAt in the internal status patch type so the queue can publish its private tracked-run age through the existing status sink. * fix(channels): wrap isActive to satisfy unbound-method lint * fix(gateway): gate busy run-age ceiling on disconnected transport (cherry picked from commit 18b79d9) * fix(deps): update fast-uri past advisory (cherry picked from commit 1be9db0) * fix(release): adapt maintenance-line hardening Backport/adapt 18ec9ce, dea1fe1, 7f32b6c, 1da345e, 931ac3e, 89780d5, and c0d99ed for the 2026.6 extended-stable maintenance line. * fix(deps): bump protobufjs to 7.6.5 Backport-adapted from a230f74. * test(gateway): cover bounded macOS process probe * chore(release): prepare 2026.6.34 * test(dotenv): share path override environment assertions * fix(release): resolve 2026.6.34 CI blockers --------- Signed-off-by: sallyom <[email protected]> Co-authored-by: joshavant <[email protected]> Co-authored-by: Peter Lee <[email protected]> Co-authored-by: sallyom <[email protected]> Co-authored-by: openclaw-clownfish[bot] <280122609+openclaw-clownfish[bot]@users.noreply.github.com> Co-authored-by: Liu Wenyu <[email protected]> Co-authored-by: pick-cat <[email protected]> Co-authored-by: Pick-cat <[email protected]> Co-authored-by: Vincent Koc <[email protected]> Co-authored-by: weiqinl <[email protected]> Co-authored-by: Claude Opus 4.8 <[email protected]> Co-authored-by: shengting <[email protected]> Co-authored-by: Shengting Xie <[email protected]> Co-authored-by: yayu <[email protected]> Co-authored-by: Agustin Rivera <[email protected]> Co-authored-by: cxbAsDev <[email protected]> Co-authored-by: ooiuuii <[email protected]> Co-authored-by: Masato Hoshino <[email protected]> Co-authored-by: Vincent Koc <[email protected]> Co-authored-by: mushuiyu886 <[email protected]> Co-authored-by: Peter Steinberger <[email protected]> Co-authored-by: Bruno Wowk (Volky) <[email protected]> Co-authored-by: Pavan Kumar Gondhi <[email protected]> Co-authored-by: Glucksberg <[email protected]> Co-authored-by: xingzhou <[email protected]> Co-authored-by: tzy-17 <[email protected]> Co-authored-by: krissding <[email protected]> Co-authored-by: lsr911 <[email protected]> Co-authored-by: Yuval Dinodia <[email protected]>
What Problem This Solves
evaluateChannelHealthintentionally treats a channel as healthy while a run is in flight, even when the transport already reportsconnected: false. That busy override is meant to be bounded by a 25 minute stale ceiling, but the ceiling is measured fromlastRunActivityAt, which the run-state heartbeat refreshes every 60 seconds for as long as any run is active.If a per-message task hangs forever, for example a channel send blocking on a half-open or dead socket with no send timeout after the transport disconnected,
onRunEndnever fires, the heartbeat keeps ticking, andlastRunActivityAtnever ages past roughly 60 seconds. The>= 25minstuck branch, the only escape from busy-overrides-disconnected, becomes permanently unreachable. The account is then reported healthy forever by the channel health monitor (if (health.healthy) continue), the readiness probe (src/gateway/server/readiness.ts), and theopenclaw healthCLI (src/commands/health.ts). No restart fires and no operator-visible signal appears until the process is restarted. Transport disconnects mid-send with no send timeout are a normal production failure mode.Why This Change Was Made
When the transport already reports disconnected, the busy override must be bounded by how long the run has actually been in flight, not by a timestamp the heartbeat keeps fresh, so a single hung run eventually breaches the stale ceiling. When the transport is healthy the channel queue's no-run-timeout contract holds: a legitimate run may take arbitrarily long and must never be aged out by its start time.
Root cause,
src/gateway/channel-health-policy.tsbusy branch (before):lastRunActivityAtis republished asnow()on every heartbeat tick increateRunStateMachine, sorunActivityAgenever grows while a run hangs.Fix. The channel run queue now tracks each in-flight run's start time keyed by an opaque run handle returned from
onRunStartand passed back toonRunEnd, and publishes the oldest still-active run's start asactiveRunStartedAt. When the oldest run ends the reported start advances to the next-oldest, and it clears when no run is active. The health policy busy branch applies the run-age term only while the snapshot reportsconnected: false(after):Because the oldest active run's start is fixed while that run is in flight and the heartbeat only advances
lastRunActivityAt, a run hanging past the threshold while the transport reports disconnected now reportsstuck. When the transport is connected, or the channel does not report transport state, the run-age term is not applied and the ceiling stays keyed tolastRunActivityAtexactly as onmain, so long legitimate runs on a healthy transport are never treated as stuck. Because the reported start is the oldest active run and advances as runs complete, a channel churning through many short overlapping runs stays healthy; only a genuinely hung run masking a disconnect breaches the ceiling. TheactiveRunStartedAtfield flows through the same run-state to snapshot path aslastRunActivityAt(run-state-machine.tspublish,account-snapshot-fields.tsread,types.core.tssnapshot type), so all three consumers ofevaluateChannelHealth(monitor, readiness, CLI) get the corrected verdict.Why This Is The Right Boundary
The masking lives in one shared decision function,
evaluateChannelHealth, consumed by the monitor, readiness probe, and health CLI, so a single change fixes all three surfaces. The heartbeat that defeats the ceiling lives increateRunStateMachine, the one owner of run activity status, so the run-age fact is surfaced where it is produced rather than re-derived per caller.createChannelRunQueueshares one run-state machine across all queue keys andKeyedAsyncQueueruns unrelated keys concurrently, soactiveRunsroutinely exceeds 1; tracking the oldest active run rather than a busy-streak start keeps a continuously busy account healthy while still catching a single hung run. Legitimate busy behavior is preserved: short, active, and overlapping runs stay healthy, and snapshots without a start time fall back to the previouslastRunActivityAtbehavior.onRunEndnow takes the run handle returned byonRunStart; the only in-repo caller,createChannelRunQueue, is migrated in this same commit, and typed plugin consumers get a compile error rather than silent breakage. Restoring a no-argumentonRunEndoverload was deliberately avoided because it would reintroduce the counter-decrement accounting this change removes and split run tracking across two paths. The 25 minute ceiling is the pre-existingBUSY_ACTIVITY_STALE_THRESHOLD_MSconstant; this change makes it reachable only for the case it was meant to bound, busy masking a transport that already reported disconnected. The queue's no-run-timeout contract is preserved: connected runs, and runs on channels that do not report transport state, are never aged out by their start time. This is distinct from PR #96198, which fixes the same class of heartbeat-hides-stuck bug in the agent diagnostic stuck-session detector (src/agents/*,src/logging/diagnostic-run-activity.ts), a different subsystem that does not touch channel health.User Impact
Channels whose per-message task hangs after a transport disconnect are now detected as stuck once the run exceeds the 25 minute ceiling, so the health monitor restarts them and the outage becomes visible to readiness and the
openclaw healthCLI instead of being silently masked as healthy until a manual process restart.Evidence
src/gateway/channel-health-policy.test.ts: a disconnected channel with a run in flight longer than the threshold and a fresh heartbeat evaluates to{ healthy: false, reason: "stuck" }(fails on pristinemain, which returnsbusy); a connected run past the threshold with a fresh heartbeat staysbusy; a run past the threshold on a channel that does not reportconnectedstaysbusy; a short-lived run staysbusy.src/plugin-sdk/channel-lifecycle.queue.test.ts: with two concurrent runs,activeRunStartedAtstays pinned to the oldest run's start while a newer run completes, advances to the next-oldest start when the oldest run ends, and clears to null when no run is active.node scripts/run-vitest.mjs src/gateway/channel-health-policy.test.ts src/plugin-sdk/channel-lifecycle.queue.test.tspasses on this patch.origin/mainsource.oxlintandoxfmt --checkare clean on the changed files;tsgois clean on the changed surfaces.Real behavior proof
Behavior addressed: a disconnected channel whose per-message run hangs forever is reported healthy forever, because the run-state heartbeat keeps refreshing the activity timestamp that the 25 minute busy ceiling is measured against, so the health monitor never restarts it.
Real environment tested: drove the real periodic health monitor loop from
startChannelHealthMonitor(src/gateway/channel-health-monitor.ts) over the realevaluateChannelHealth(src/gateway/channel-health-policy.ts) fed by a livecreateRunStateMachine(src/channels/run-state-machine.ts) with its realsetIntervalheartbeat, on real wall-clock (Date.now), on pristine main 091584b and on PR head 21284cf. Only the channel transport was stubbed, at the narrowest seam (theChannelManagerthe monitor calls), so the monitor's real restart recovery path stays observable. Executed on a remote Crabbox AWS lease (c7a.xlarge, eu-west-1, lease cbx_452e9479c6de), not on the author machine. The 25 minute busy ceiling and the 60 second heartbeat were both scaled down 300x (to 5 seconds and 200 milliseconds) so a genuinely hung run breaches the ceiling in seconds of real elapsed time rather than 26 minutes; nothing else was altered and both runs used identical scaling.Exact steps or command run after this patch: a run begins via the real run-state machine and then hangs (its end callback never fires); the real monitor evaluates the live snapshot every 250 milliseconds while the transport stays connected:false; the harness prints the live snapshot ages and health verdict once per second and the monitor's real restart actions as they occur.
Evidence after fix:
Observed result after fix: on identical live input, the heartbeat pins lastRunActivityAt age at 0.2s in both runs, so BEFORE the channel stays healthy and busy for the whole hang and the monitor never restarts it (monitorRestarts=0); AFTER, the fixed activeRunStartedAt age grows until it breaches the ceiling, the verdict flips to unhealthy and stuck, and the real health monitor completes its restart recovery path (stopChannel then startChannel) five times over the window. The one value that moves from wrong to right is the health verdict for a hung disconnected channel, and with it the monitor's restart action.
What was not tested: no live provider socket was hung end to end; the hang was a real never-resolving run under the real heartbeat, and the transport was a stub at the ChannelManager seam. The busy ceiling and heartbeat were scaled 300x on real wall-clock so the breach happens in seconds; the unscaled 25 minute breach was not waited out. Full build not run; the change is not on a lazy or packaging boundary.