Skip to content
Ritesh FirodiyaGet in touch

Work / Chitragupt / Wiki / Decisions

2026-08-03-dev-self-validation-round-2

Decisioncanonicalverified 2026-08-03

DECISION.2026-08-03.DEV-SELF-VALIDATION-ROUND-2

yarn db:self round 2 — real bugs found across all four pillars, most fixed

Decision

Full-corpus yarn db:self (425 real docs, self.ritesh, FY2018-19 → FY2026-27) was re-run as an end-to-end audit of Inbox, Tax, Expense, and Portfolio. This follows on from 2026-07-25-dev-self-validation-findings — that session's four fixes held up; this session found a new, more serious set of bugs, mostly in Expense and Portfolio, plus one external blocker. Six parallel research passes (one per pillar + a hardcoded-value sweep + a live browser cross-check against wireframes) diagnosed root causes before any fix landed, then three parallel fix passes implemented them.

Real bugs found and fixed this session

  1. Expense: SIP/MF debits miscounted as consumption spend, Invested KPI permanently ₹0. bank-statement.ts never tagged investments_savings-category debits with entry_class: "investment" — they fell through to "expense". Fixed: debits categorised as investments_savings now get entry_class: "investment". apps/functions/src/triggers/firestore/ledger-mappers/bank-statement.ts.
  2. Expense: non-expense documents (GST return, TIS, ITR payment challan, broker ledger) leaked into the /expense/transactions list as huge, sometimes-duplicated outflows. These mappers deliberately leave entry_class unset ("audit-only" rows), but the website's row filter (entryClass !== null && entryClass !== "expense" && entryClass !== "transfer") let null fall through as visible. Fixed the filter to require an explicit "expense"/"transfer" match. apps/website/src/store/ledger.ts.
  3. Expense: category-donut taxonomy drift. The wiki mandates exactly 8 donut buckets (Housing · Food & groceries · Loan EMI · Other · Transport · Tax · Subscriptions · Health); the schema carries 28 fine-grained categories (base + tax-linked chips) with no collapse layer, and no loan_emi category existed at all (home-loan EMIs were bucketed as housing). Added loan_emi to EXPENSE_CATEGORIES, rerouted the narration-based categoriser's loan-EMI patterns to it, and added a bucket-collapse map in BurnDonut.tsx so the chart always renders exactly the 8 canonical buckets regardless of how many raw categories exist underneath.
  4. Portfolio: statement history collapsed to one row per instrument, not one row per year. portfolio-position-mapper.ts keyed positions by ${form_type}__${uidPrefix} with no period component — 19 confirmed EPF passbooks (one per year) collapsed onto a single Firestore row; "latest upload wins" silently discarded 18 years of history. Rekeyed the position id to ${holdingId}__${ay}; portfolio-review-engine.ts now reads every position but aggregates only the LATEST-AY row per holding (summing across years would have multiplied cumulative EPF/PPF balances).
  5. Portfolio: nonsensical "your return" figures (+69,335% observed live). EPF/PPF positions always carry cost_basis_paise: 0 (the passbook reports a balance, never a contribution total). Pooling them into a portfolio-wide (nav − invested) / invested let a near-zero denominator (from the few holdings that DO have a cost basis, e.g. a ₹20K ELSS buy) swallow a huge numerator (the EPF/PPF balance). Fixed by computing total_return_pct only over positions with a real cost basis (new nav_with_cost_basis_paise / invested_with_cost_basis_paise fields on PortfolioReview); the dashboard's Portfolio-vs-Nifty trajectory chart independently recomputed the same broken ratio from raw fields and needed the same fix threaded through PortfolioHistoryPoint and history.ts.
  6. Portfolio: capital gains double-counted for AY2022-23. Two real, differently-named source files (fo.xlsx, equity.xlsx) turned out to be byte-different but content-identical exports of the same Zerodha Tax P&L report. The dedup engine correctly detected this and stamped superseded_by on the older doc — but the seed script's autoConfirm() re-read the doc and force-wrote status: "confirmed" without checking whether the fresh read still showed "unreviewed", racing against the dedup write and re-confirming an already-superseded duplicate. The production confirmDocument callable does NOT have this bug — it calls assertConfirmable(status) against a fresh read. Fixed the seed script to mirror that guard. This was a validation-harness bug, not a product bug, but it was corrupting every yarn db:self run's capital-gains numbers.
  7. Tax: waterfall showed "Clear ₹0" for the current AY despite ₹14.6L computed tax vs ₹60,858 TDS. position.kind was correctly "incomplete" (no Form-16/salary-slip confirmed yet for that AY — itself downstream of bug #8 below), but the waterfall's label ternary silently fell through any non-refund/non-owe kind to "Clear". Fixed to render "Incomplete — salary not confirmed yet" / —, matching the pattern PositionCard.tsx already used correctly. apps/website/src/app/(app)/tax/components/TaxPage.tsx.
  8. Tax: tax_review_runs stuck dirty: true with no recompute and no UI signal. on-self-identity-written.ts marks every AY dirty on any identities/self write; only the 5-minutely scheduled drain clears it, which doesn't fire in the emulator, and the website never reads dirty anywhere. Added a synchronous recompute of the current open AY right after the dirty-stamp write; past AYs still rely on the drain.
  9. Tax: deduction caps hardcoded twice in the UI, not sourced from the AY-versioned rule tables. TaxPage.tsx re-declared §80C/§80CCD(1B)/§80D caps as bare literals (with the §80CCD(1B) cap duplicated a second time in the same file) instead of reading getOldRegimeCapPaise(ay, key) from packages/shared/src/rules/lib/deduction-caps.ts — the same source the compute engine uses. Re-exported the helper from the shared rules barrel and wired the UI to call it per-AY.
  10. Tax: 15-question checklist actually has 19 questions, AND was silently truncated to 15 rendered rows. Header hardcoded "15 QUESTIONS"; separately, entries.slice(0, 15) cut the rendered list short even though the progress counter used the true length. Both replaced with the real INTERROGATOR_CATALOG_V1.length (19).
  11. Fake per-document confidence badge. DocumentsFedIn.tsx rendered a two-way literal (73%/99%) keyed only on a status boolean. Added the real classification_confidence field to DocumentRecord and wired it through; hides the badge when null instead of fabricating a value.
  12. Inbox: 13th folder bucket mislabeled and wrong semantics. Was "Misc / Unclassified" (folder-routing based); wiki specifies "⚠ Needs your attention", status-based (pending_identity_confirmation / needs_ocr / needs_classification / pending_password). Relabeled and rewired the membership predicate in FolderTree.tsx + DocumentList.tsx.
  13. Inbox: missing storage meter, missing "Needs your action" shelf, text-pill status instead of a glyph. All three added/fixed — storage meter extracted into a shared StorageMeter.tsx (also de-duplicating a hardcoded free-tier-only cap on /settings/storage), new NeedsActionShelf.tsx mounted between the dropzone and the document list, new StatusGlyph.tsx replacing the text pill on doc rows.
  14. Stale copy referencing a removed feature. /sign-in still advertised "CA-ready handoff packet — 1-click PDF export", removed per 2026-07-28-remove-ca-handoff-surface. Replaced with copy describing the surviving Invite-CA/Hire-CA paths.
  15. Portfolio review page: missing "vs balanced ideal" table, wrong donut legend taxonomy, missing observation panels (stubbed with an unrelated "Sector exposure" card), missing SEBI disclaimer strip, block-order drift. All fixed — allocation donut now collapses PortfolioAssetClass onto the 4 canonical buckets with a Now/Ideal/Δ table against the 60/25/10/5 balanced-ideal split; the three narrative panels (What lost you money / What made you money / Gaps in your safety net) now derive from real top_holdings + concentrated_holding_count fields, rendering an honest "Not enough data yet" when ungrounded rather than inventing figures; one PortfolioSebiFooterStrip replaces two duplicated inline mentions.
  16. Inbox: rejection_reason only ever set for one LLM-fallback failure mode (input_too_large). Every other terminal failure (insufficient_credits, schema_invalid, no_tool_use) left it null, so the Review-page banner gave no signal that automatic classification never ran at all. Added reason strings for all terminal modes except not_configured (dev/emulator-only).

Found, NOT fixed — needs a human decision

  • fy2025-26-federal-bank-housing-loan-provisional-certificate.pdf (form_type home-loan-interest-cert, confirmed doc b4c96889469658e0) carries borrower_name: "RAKESH SHANTILAL FIRODIYA" — not Ritesh — yet its §24(b)/§80C amounts feed directly into self.ritesh's own tax review with no identity/co-borrower gate. Shared middle name suggests a real joint loan with a brother, in which case this is legitimate and just needs a co-borrower UX (e.g. "confirm this is your share"); alternatively it's a misfiled document. The system currently has no borrower-name-vs-account-owner cross-check for loan-interest-cert (or similarly self-attributed) document types — bank docs key by IFSC+account, but loan docs key only by loan-account-number-last-4. Needs a product call before building the gate.
  • PAN mismatch, no reconciliation. Form-16 reports employee_pan: "ACCPF6873C" (0.95 confidence) vs identities/self.pan: "ACCPF6973C" — a one-digit transposition. Doesn't affect any tax math (PAN isn't a compute input) and is purely informational, but nothing in the pipeline cross-checks PAN consistency across confirmed documents. Worth a lightweight reconciliation flag if a reconcile-cross-source.ts-style pass is ever built out further.
  • Debt-MF three-regime capital-gains split is dead code. ledger-mappers/mf-cas.ts passes through four debt-MF fields (debt_stcg_realised, debt_ltcg_indexed_realised, debt_ltcg_flat_realised, debt_ltcg_slab_realised) that don't exist anywhere in the mf-cas parser or its field schema — capital_gains/{ay} debt-MF buckets will always read 0 regardless of real debt-MF activity. Fixing this requires the MF-CAS parser to actually split per-scheme by (acquisition date × transfer date), a real parsing feature, not a wiring bug. Logged here rather than attempted in this pass.
  • ELSS current_value_paise mirrors cost_basis_paise (no real gain/loss visible). Investigated and confirmed this is NOT a regex bug — the real source document (a Zerodha "Statement of investments in Tax Saving (ELSS) Funds") is a purchase confirmation only; it never prints a current NAV. The mapper's existing fallback (floor to cost basis, suppress ytd_return_pct to null rather than claim a false 0%) is already the most honest handling possible without a live market-data feed. Not actionable without a new data source.
  • Direct equity (demat) holdings have no path into Portfolio at all. None of the 34 registered parsers represent "current holdings as of date X" for a brokerage account — broker-cg is realised P&L, broker-trade-log/broker-dividend-log/broker-ledger-statement are transaction logs, none are a holdings/positions statement. A user with 17 Zerodha statements uploaded sees only 3 Portfolio positions (PPF/ELSS/EPF) because there is genuinely no parser type for a demat holdings/portfolio-value statement. This is a real V1 scope gap, not a routing bug — flagged for inbox-document-types / ROADMAP, not fixed here.

External blocker (not a code bug)

The Anthropic API key backing the parser's LLM-fallback classification (ANTHROPIC_API_KEY in apps/functions/.env.local) has an exhausted credit balance — confirmed via a direct API call returning "Your credit balance is too low to access the Anthropic API." This is coded as a non-retryable terminal failure by design (correctly, per 2026-07-25-dev-self-validation-findings bug #2's fix), so ~130 of 425 docs land in needs_classification / needs_ocr instead of being auto-classified. Of those, roughly 55-60 are plausible hr-letter / travel-receipt / generic-receipt / utility-bill candidates (offer letters, relieving letters, ESOP grant docs, ride/travel invoices) that would likely classify correctly once the LLM path works again; the rest (~22 identity-document photos, resumes, personal correspondence) have no matching parser regardless and are expected to sit unclassified. Needs a billing top-up on the Anthropic account, not a code change.

Impact

  • pillar-expense, pillar-portfolio, pillar-tax, pillar-inbox — all four pillar entity pages' canonical values remain accurate; no canonical value changed, only implementation bugs against those values.
  • New schema fields: nav_with_cost_basis_paise / invested_with_cost_basis_paise on PortfolioReview; invested_with_cost_basis_paise / value_with_cost_basis_paise on PortfolioHistoryPoint; loan_emi added to EXPENSE_CATEGORIES.
  • apps/functions/scripts/seed-self-user.mjs now guards against the dedup-race described in bug #6 — future yarn db:self runs should not reproduce double-counted capital gains from duplicate broker exports.
  • Companion decisions logged by the same session: 2026-08-03-fix-checklist-question-count-mislabel, 2026-08-03-correct-widget-order-entity-drift.

Status

Active.

Sources

  • This chat session, 2026-08-03. Six parallel research agents (one per pillar audit + hardcoded-value sweep + live browser cross-check) followed by three parallel fix agents (Tax UI, Inbox UI, Portfolio UI), plus direct fixes to the Expense ledger mapper, Portfolio position mapper/review engine/history callable, Tax waterfall label, and the seed script.
  • Supersedes nothing; extends 2026-07-25-dev-self-validation-findings.

Every project of mine is written down like this.

Read the résumé