Skip to content
Ritesh FirodiyaGet in touch

Work / Chitragupt / Wiki / Synthesis

document-pipeline-boundaries

Synthesiscanonicalverified 2026-07-19

SYNTHESIS.DOCUMENT-PIPELINE-BOUNDARIES

Document pipeline — five boundary schemas

Summary

Every uploaded document travels through five stages: ingest → parse → identity → classify → ledger. Each stage owns a narrow, typed contract for its input and output, projected off the physical DocumentRecord. The schemas live in packages/shared/src/schemas/document-boundaries.ts alongside toIngestedBlob/toParserOutput/toClassifiedDocument helpers. A single per-upload trace_id (UUID v4, minted at ingest) is stamped onto the doc row, every ledger entry, and every structured log line, so a doc's full journey is one grep.

Built from

The five stages

Stage Boundary schema Input Output Owner
Ingest IngestedBlobSchema GCS finalize event + tenant + category Validated blob + trace_id + doc row seeded at parse_pending apps/functions/src/triggers/storage/handlers/ingest-vault-document.ts
Parse ParserOutputSchema IngestedBlob form_type, temporal, fields, ay, identity_extracted, content_hash, instance_key apps/functions/src/parser/run-and-persist.ts (delegates to packages/parser/src/dispatcher.ts)
Identity IdentityBindingSchema ParserOutput identity_id (matched entity or self) OR pending_identity_confirmation handoff with transfer_to_member_uid / identity_extracted evidence apps/functions/src/parser/identity-matcher.ts
Classify ClassifiedDocumentSchema IdentityBinding + dedup overlay status (final), bucket, duplicate_of / supersedes apps/functions/src/parser/run-and-persist.ts (via applyDedupOverlay) + confirmation callables
Ledger emit LedgerEntrySchema ClassifiedDocument + confirm gate passed Per-form LedgerEntry rows with deterministic idempotency_key apps/functions/src/triggers/firestore/ledger-mappers/* + _lib.ts:makeEntry

Trace id contract

  • Minted once at ingest via randomUUID() (ingest-vault-document.ts).
  • Stamped onto the doc row's trace_id field alongside the ingest-seed status write.
  • Read back by every downstream stage (parser, identity, dedup, confirm, recompute, ledger mapper) and included in every structured log line.
  • Copied onto every emitted LedgerEntry by makeEntry (via MapperContext.traceId).
  • null on legacy rows / test fixtures / bulk backfill entries — every consumer treats it as an optional string.

RecomputeSignal

Not a Firestore doc — an in-process value that flows from onDocumentConfirmed → recomputeAfterConfirm and from categorizeLedgerEntry → its dirty-flag writes. Shape:

{ uid, ay, pillar: "tax" | "expense" | "portfolio" | "capital_gains", trace_id, entry_ids? }

entry_ids is populated by the categorise path so the engine can skip rows it already saw; the doc-confirm path recomputes over the full AY snapshot and omits it.

Related

Last refresh

2026-07-19 — created alongside the boundary-schemas module + trace_id pipeline.

Every project of mine is written down like this.

Read the résumé