← Corpus / augment-it / plan
Augment from DB · Phase 5 — stream-scan mode: scan a pulse stream, badge the already-known, one-click the new into the corpus
Stream scan is a mode of search-and-add, not a third remote: the stream URL is the authoritative index, deduped against `content_items`.
- Path
- plans/Augment-From-DB-Phase-5-Stream-Scan-Mode.md
- Authors
- Michael Staton
- Augmented with
- Claude Code on Claude Fable 5
- Tags
- Plan · Augment-It · Augment-From-DB · Phase-5 · Stream-Scan · Entity-Pulse · Pulse-Streams
Augment from DB · Phase 5 — stream-scan mode
Spec reference
Implements Phase 5 (v1.1) of [[../specs/Augment-From-DB-Flow]], per decision D3 (mode of search-and-add, not a third remote). Branch: rebuild/turbo-rsbuild.
Two design facts found during authoring that make this cheap:
runOfficialBlogPackalready has the exact seam:curated_index_urls— when non-empty, discovery (SerpApi + homepage + path-guess) is SKIPPED and the pack harvests straight from the given index pages. A stream scan IS “curated index = the stream URL.” No entity-pulse changes at all.- social-search has no DB access, so corpus dedup is a cross-service NATS call (the domains.ts → content-ingest precedent): a new
content.urls.checkverb on record-surrealdb-resolver ({urls[]} → {existing[]}againstcontent_items’ unique url index). Service-to-service only — no workspace map entry.
Steps
content.urls.check—checkContentUrlsinresolver.ts(SELECT url FROM content_items WHERE url IN $urls); handler inhandlers.ts(content.urls.check.requested). Shared-ledger read: no client filter, content_items is keyed by unique URL.stream-scan.ts(social-search) —scanStream(nc, {org_slug, stream_url, stream_kind, client, max_items?}):runOfficialBlogPack({row_id: 'org:'+org_slug, row_url: stream_url, curated_index_urls: [stream_url], max_posts_total})→content.urls.checkover the item URLs → items +already_in_corpusflags. Blog/RSS/newsroom kinds are the dependable path; social-wall kinds pass through the same call flagged experimental (non-goal: no reliability commitment).server.ts— subscribeorganization.stream.scan.requested; ok:false on failure (search.fire’s contract, same reason).capabilities.ts—organization.stream.scan→ subject, 60s timeout (multi-stage, like pack.entity_pulse).- org-workbench —
AdditiveListgains an optional per-entry action (entryaction={{label, fn}}); OrgCard wires it on the Pulse-streams list only: “scan” per stream →requestSearchwith an envelope extended bystream: {url, kind}(targetcorpus— scan adds land inorg_corpus). - search-and-add — when the envelope carries
stream, the App enters scan mode: TermBar hidden, stream URL + “Re-scan” shown, results viaorganization.stream.scan;ResultRowgains aknownbadge (“in corpus”) and disables ➕ for known items; ➕ routes through the existing corpus add +entity-updatedbroadcast. - Verify — rebuild both service containers; live NATS proof against Aspen’s seeded blog stream (
https://www.aspeninstitute.org/blog/, added in Phase 2’s proof): scan → items with flags →organization.corpus.addone item → re-scan → its flag flips to true. svelte-check + builds on both remotes + shell. One real Firecrawl-credit scan, accepted. - Changelog; spec status → Shipped (all five phases done; operator browser walk-throughs noted as pending); commit + push as
attempt(augment-from-db, stream-scan, step5):.