← Corpus / augment-it / issue

Crawl replies can be lost — the eternal 'crawling…' spinner (no client timeout, reconnect-dropped invokes, tight dispatch ceiling)

The Curry Foundation crawl finished in 87s and the tab spun forever: no client deadline, reconnect-dropped invokes, a 300s ceiling.

Path
issues/Crawl-Replies-Can-Be-Lost-Eternal-Spinner-No-Client-Timeout.md
Authors
Michael Staton
Augmented with
Claude Code on Claude Fable 5
Tags
Issue · Augment-It · Didi-Crawl · Workspace · Transport · Reliability

Crawl replies can be lost

The incident (2026-07-24 evening, logs-confirmed)

Operator fires a links crawl on The Beth & Ravenel Curry Foundation → UI stuck on “crawling…”. prompt-runner logs show the truth: crawl completed … ms: 87542. The work succeeded; the reply never made it back. Earlier team crawls (Quell 211s, Truist 147s) completed too — 211s uncomfortably close to the 300s dispatch ceiling.

The three stacked gaps

  1. No client-side deadline. workspace.invoke waits forever; a lost reply is an eternal spinner with a disabled button.
  2. WS reconnects drop pending invokes. The transport reconnects with backoff, but in-flight invokes don’t survive the new socket — and workspace-service container rebuilds (three today) sever every session. Any crawl pending across one is orphaned.
  3. 300s dispatch ceiling vs multi-minute crawls. The NATS request timeout was set before live team-crawl timings existed.

Shipped mitigations

  • 660s withDeadline race on crawlSearch and crawlTeam — the spinner always resolves; the error names that the run may still have finished server-side and to re-fire.
  • organization.crawl dispatch ceiling 300s → 600s (pack.fan_out’s precedent).

Remaining (folded elsewhere)

  • Transport-level invoke re-correlation across reconnects (or server-held results claimable by request id) → the liveness sweep (#21).
  • Crawl progress frames (#35) would surface a dead reply path within seconds instead of minutes.