Work / AskCal / Wiki / Synthesis
accuracy-findings
Synthesiscanonicalverified 2026-09-12
SYNTHESIS.ACCURACY-FINDINGSAccuracy: what the published evidence says
Researched 12 Sep 2026, before spending anything on scans. Two questions: which model, and how do we test the thesis without weighing 50 dishes first.
The benchmark worth trusting
Vision-Language Models for Image-Based Dietary Assessment (bioRxiv, Jul 2026; PMC13483877) — ten approaches over 3,229 images from Nutrition5k, which has scale-weighed ground truth.
| Model | Calorie MAE | CCC | Ingredient Jaccard | Cost / 1K images |
|---|---|---|---|---|
| Gemini 3.0 Flash | 80.7 kcal | 0.767 | N.A. | N.A. |
| Gemini 3.1 Flash Lite | N.A. | 0.754 | 0.655 | $0.59 |
| Gemini 2.0 Flash | N.A. | 0.742 | 0.621 | $0.10 |
| Qwen2-VL-7B (open weights) | N.A. | N.A. | N.A. | self-hosted |
| GPT-4o / 4o-mini / 5 Mini, Claude Haiku 4.5, FatSecret | N.A. | N.A. | N.A. | N.A. |
Per-model figures not printed in the abstract are N.A. rather than guessed.
No human or dietitian baseline is reported, so "is 80.7 kcal good?" is still unanswered by this paper. Four annotators reviewed 440 images but their own error was not disclosed.
What this changes
The plan costed the workload on Claude Haiku 4.5 at ~$2.30 per 1,000 scans. Gemini 3.1 Flash Lite benchmarks within noise of the best model at $0.59 per 1,000 — roughly 4× cheaper — and 2.0 Flash is $0.10 per 1,000. AI cost as a share of revenue was the number that could quietly kill this business; at $0.59/1K it stops being the binding constraint.
We currently default to gemini-3.6-flash, which postdates this benchmark and
is therefore unmeasured. Worth running our own eval across 3.6-flash,
3.1-flash-lite and 2.0-flash before committing — that is exactly what
packages/eval is for, and the sweep is cheap at these prices.
A benchmark NOT to trust
foodvision-bench reports PlateLens at ±1.1% MAPE on 231 USDA-weighed meals. Treat as unreliable:
- ±1.1% MAPE on photo calorie estimation would beat a dietitian with a food scale. The peer-reviewed number above is an 80.7 kcal MAE, a different universe.
- PlateLens tops both tiers, photo and manual entry.
- Its "independently replicated" figure matches its own vendor claim exactly.
- Commercial systems are scored via human transcription; only the open-source baselines (CLIP-ViT-L/14 ±10.4%, SigLIP-SO-14 ±11.5%) are deterministic.
Those open-source baselines are plausible and useful as a floor: a general image-text model with no food-specific training lands around ±10% MAPE.
Testing the thesis
Nutrition5k (Google
Research, CC BY 4.0, commercial use permitted) — ~5,000 plates, every
ingredient weighed on a rig as it was added, nutrition derived per-gram from
USDA. packages/eval ingests it via pnpm --filter @askcal/eval ingest.
It settles the baseline, not the thesis. Nutrition5k is Google cafeteria food: plated, Western, largely unmixed, shot overhead under even light — the "photogenic American plate" this product positions away from. It answers "does clarifying help at all?" cheaply and immediately. It cannot answer "does clarifying help on a fried rice", which is the actual claim.
So the test set is two halves:
| Half | Source | Answers | Status |
|---|---|---|---|
| Baseline | Nutrition5k, free | Does asking beat guessing at all? | ingestion built |
| Thesis | Hand-weighed, mixed/sauced/shared, per dataset/README.md |
Does it help where the category fails? | still to weigh |
Reporting a good Nutrition5k score as proof of the thesis would be the most expensive mistake available here: it would validate the product on exactly the food the product is not for.
Sources
- VLMs for Image-Based Dietary Assessment (PMC) · bioRxiv
- Nutrition5k · paper
- foodvision-bench — cited as a negative example
- Benchmarking Foundation Model Dietary Estimates (PMC)
Every project of mine is written down like this.
Read the résumé