Precision nutrition hinges on one deceptively simple problem: knowing how much someone actually ate. Dietary self-report tools remain notoriously inaccurate, and the promise of AI-powered food recognition has generated significant excitement — but rigorous head-to-head comparisons against human judgment have been scarce, especially outside Western food environments.
This study enrolled 128 adults in Astana, Kazakhstan, to estimate portion sizes of 51 foods and 8 beverages from standardized photographs. Participants were randomized to either unassisted visual estimation or atlas-assisted estimation using a visual food guide calibrated to Central Asian cuisine; a multitask AI model trained on regional food images was evaluated in parallel against actual weighed portions. Atlas-assisted human estimation achieved the best accuracy, with a mean absolute error of roughly 81 grams and a mean absolute percentage error near 45%. The AI model landed in the middle — MAE approximately 97 grams and MAPE around 68% — meaningfully outperforming unaided human guesswork (MAE ~134 g, MAPE ~79%) but falling notably short of atlas-guided estimation. All three differences reached statistical significance.
These findings carry important nuance for the broader AI-in-nutrition field. Most AI portion-estimation models have been trained and validated predominantly on Western or East Asian food datasets; performance tends to degrade substantially when applied to cuisines with different plating conventions, communal serving styles, or irregular food geometries — all characteristic of Central Asian dietary contexts. The intermediate performance of the AI here is actually a meaningful result: it suggests region-specific training data can bring AI into a competitive range, though not yet to the accuracy of a simple visual reference tool. The roughly 45% mean percentage error even for atlas-assisted humans underscores that portion estimation at scale remains a formidable challenge regardless of method. For public health nutrition surveillance, atlas-assisted approaches remain the pragmatic gold standard until AI models can be trained on far larger, culturally diverse datasets. This study is incremental but practically valuable — it benchmarks a genuine gap and points toward where training data investment is most needed.