Skip to content
Ritesh FirodiyaGet in touch

Work / AskCal / Wiki / Synthesis

accuracy-findings

Synthesiscanonicalverified 2026-09-12

SYNTHESIS.ACCURACY-FINDINGS

Accuracy: what the published evidence says

Researched 12 Sep 2026, before spending anything on scans. Two questions: which model, and how do we test the thesis without weighing 50 dishes first.

The benchmark worth trusting

Vision-Language Models for Image-Based Dietary Assessment (bioRxiv, Jul 2026; PMC13483877) — ten approaches over 3,229 images from Nutrition5k, which has scale-weighed ground truth.

Model Calorie MAE CCC Ingredient Jaccard Cost / 1K images
Gemini 3.0 Flash 80.7 kcal 0.767 N.A. N.A.
Gemini 3.1 Flash Lite N.A. 0.754 0.655 $0.59
Gemini 2.0 Flash N.A. 0.742 0.621 $0.10
Qwen2-VL-7B (open weights) N.A. N.A. N.A. self-hosted
GPT-4o / 4o-mini / 5 Mini, Claude Haiku 4.5, FatSecret N.A. N.A. N.A. N.A.

Per-model figures not printed in the abstract are N.A. rather than guessed.

No human or dietitian baseline is reported, so "is 80.7 kcal good?" is still unanswered by this paper. Four annotators reviewed 440 images but their own error was not disclosed.

What this changes

The plan costed the workload on Claude Haiku 4.5 at ~$2.30 per 1,000 scans. Gemini 3.1 Flash Lite benchmarks within noise of the best model at $0.59 per 1,000 — roughly 4× cheaper — and 2.0 Flash is $0.10 per 1,000. AI cost as a share of revenue was the number that could quietly kill this business; at $0.59/1K it stops being the binding constraint.

We currently default to gemini-3.6-flash, which postdates this benchmark and is therefore unmeasured. Worth running our own eval across 3.6-flash, 3.1-flash-lite and 2.0-flash before committing — that is exactly what packages/eval is for, and the sweep is cheap at these prices.

A benchmark NOT to trust

foodvision-bench reports PlateLens at ±1.1% MAPE on 231 USDA-weighed meals. Treat as unreliable:

  • ±1.1% MAPE on photo calorie estimation would beat a dietitian with a food scale. The peer-reviewed number above is an 80.7 kcal MAE, a different universe.
  • PlateLens tops both tiers, photo and manual entry.
  • Its "independently replicated" figure matches its own vendor claim exactly.
  • Commercial systems are scored via human transcription; only the open-source baselines (CLIP-ViT-L/14 ±10.4%, SigLIP-SO-14 ±11.5%) are deterministic.

Those open-source baselines are plausible and useful as a floor: a general image-text model with no food-specific training lands around ±10% MAPE.

Testing the thesis

Nutrition5k (Google Research, CC BY 4.0, commercial use permitted) — ~5,000 plates, every ingredient weighed on a rig as it was added, nutrition derived per-gram from USDA. packages/eval ingests it via pnpm --filter @askcal/eval ingest.

It settles the baseline, not the thesis. Nutrition5k is Google cafeteria food: plated, Western, largely unmixed, shot overhead under even light — the "photogenic American plate" this product positions away from. It answers "does clarifying help at all?" cheaply and immediately. It cannot answer "does clarifying help on a fried rice", which is the actual claim.

So the test set is two halves:

Half Source Answers Status
Baseline Nutrition5k, free Does asking beat guessing at all? ingestion built
Thesis Hand-weighed, mixed/sauced/shared, per dataset/README.md Does it help where the category fails? still to weigh

Reporting a good Nutrition5k score as proof of the thesis would be the most expensive mistake available here: it would validate the product on exactly the food the product is not for.

Sources

Every project of mine is written down like this.

Read the résumé