Work / Chitragupt / Wiki / Synthesis
document-pipeline-boundaries
Synthesiscanonicalverified 2026-07-19
SYNTHESIS.DOCUMENT-PIPELINE-BOUNDARIESDocument pipeline — five boundary schemas
Summary
Every uploaded document travels through five stages: ingest → parse → identity → classify → ledger. Each stage owns a narrow, typed contract for its input and output, projected off the physical DocumentRecord. The schemas live in packages/shared/src/schemas/document-boundaries.ts alongside toIngestedBlob/toParserOutput/toClassifiedDocument helpers. A single per-upload trace_id (UUID v4, minted at ingest) is stamped onto the doc row, every ledger entry, and every structured log line, so a doc's full journey is one grep.
Built from
- inbox-document-status — the state machine each stage flips
- inbox-document-types — the DOCUMENT_TYPES registry each stage looks up
- ledger-entry-types — LedgerEntry's
entry_typeenum - upload-only — the concept that makes the pipeline load-bearing
- read-only-review — why LedgerEntry is server-only-write
The five stages
| Stage | Boundary schema | Input | Output | Owner |
|---|---|---|---|---|
| Ingest | IngestedBlobSchema |
GCS finalize event + tenant + category | Validated blob + trace_id + doc row seeded at parse_pending |
apps/functions/src/triggers/storage/handlers/ingest-vault-document.ts |
| Parse | ParserOutputSchema |
IngestedBlob |
form_type, temporal, fields, ay, identity_extracted, content_hash, instance_key |
apps/functions/src/parser/run-and-persist.ts (delegates to packages/parser/src/dispatcher.ts) |
| Identity | IdentityBindingSchema |
ParserOutput |
identity_id (matched entity or self) OR pending_identity_confirmation handoff with transfer_to_member_uid / identity_extracted evidence |
apps/functions/src/parser/identity-matcher.ts |
| Classify | ClassifiedDocumentSchema |
IdentityBinding + dedup overlay |
status (final), bucket, duplicate_of / supersedes |
apps/functions/src/parser/run-and-persist.ts (via applyDedupOverlay) + confirmation callables |
| Ledger emit | LedgerEntrySchema |
ClassifiedDocument + confirm gate passed |
Per-form LedgerEntry rows with deterministic idempotency_key |
apps/functions/src/triggers/firestore/ledger-mappers/* + _lib.ts:makeEntry |
Trace id contract
- Minted once at ingest via
randomUUID()(ingest-vault-document.ts). - Stamped onto the doc row's
trace_idfield alongside the ingest-seed status write. - Read back by every downstream stage (parser, identity, dedup, confirm, recompute, ledger mapper) and included in every structured log line.
- Copied onto every emitted
LedgerEntrybymakeEntry(viaMapperContext.traceId). nullon legacy rows / test fixtures / bulk backfill entries — every consumer treats it as an optional string.
RecomputeSignal
Not a Firestore doc — an in-process value that flows from onDocumentConfirmed → recomputeAfterConfirm and from categorizeLedgerEntry → its dirty-flag writes. Shape:
{ uid, ay, pillar: "tax" | "expense" | "portfolio" | "capital_gains", trace_id, entry_ids? }
entry_ids is populated by the categorise path so the engine can skip rows it already saw; the doc-confirm path recomputes over the full AY snapshot and omits it.
Related
- inbox-document-status — status transitions each boundary respects
- inbox-document-types — DOCUMENT_TYPES registry the parser + identity + classify stages consume
- ledger-entry-types — LedgerEntry.entry_type contract
- pillar-inbox — where the ClassifiedDocument surfaces
- pillar-tax · pillar-expense · pillar-portfolio — RecomputeSignal targets
Last refresh
2026-07-19 — created alongside the boundary-schemas module + trace_id pipeline.
Every project of mine is written down like this.
Read the résumé