Publication date: 19 juni 2026
University: Wageningen University

Artificial intelligence and computer vision for quantitative assessment of eating behavior and urban food landscapes

Summary

Eating behavior and food environment accessibility are fundamental determinants of nutritional intake and health outcomes, yet traditional assessment methods—manual video annotation for eating behavior and observational audits for food environments—remain labor-intensive, time-consuming, and difficult to scale. This thesis investigated how artificial intelligence can automate assessment of eating behavior and urban food landscapes to support nutrition research at both individual and population levels.

At the population level, a data-driven methodology was developed to assess urban food landscapes using restaurant menu data from online delivery platforms in Boston, London, and Dubai (Chapter 2). Machine learning matched menu items to the U.S. FoodData Central database, enabling calculation of nutritional indices at neighborhood level. Database coverage varied substantially by city—Boston (71%), London (56%), and Dubai (42%)—reflecting limitations in region-specific nutritional data availability. Dietary fiber demonstrated significant inverse associations with obesity in both London (p=0.001) and Boston (p=0.004), while higher socioeconomic neighborhoods consistently showed better access to nutrient-rich foods. This methodology provides a scalable alternative to traditional food environment assessment, enabling policymakers to identify neighborhoods at risk for inadequate nutritional access.

At the individual level, a systematic review following PRISMA guidelines identified five methodological categories for automated eating behavior detection from video recordings: facial landmarks, deep learning, optical flow, active appearance model, and video fluoroscopy (Chapter 3). Facial landmarks emerged as the most promising approach for detecting both bites and chews. Building on these findings, a computationally efficient rule-based system utilizing 468 3D facial keypoints was developed for automated bite counting, achieving 79% accuracy with available annotation and 71.4% accuracy without annotation input across 164 videos from 15 participants, with consistent performance across varying food textures (Chapter 4).

To improve detection accuracy, transformer-based architectures were subsequently developed using 103 annotated videos from 36 participants consuming standardized meals (Chapter 5). The Vision Transformer model achieved 98.45% frame accuracy and 86.2% counting accuracy for bite detection, effectively capturing global spatial relationships through self-attention mechanisms. For chew detection, a CNN-LSTM architecture outperformed the Vision Transformer (85.56% versus 69.21% accuracy), as sequential models better captured the temporal dynamics inherent to chewing behaviors. Sip detection using pretrained VideoMAE achieved 61% accuracy, constrained by limited training data and annotation challenges.

These validated models were integrated into the Automated Meal Video Analysis (AMVA) Toolkit, an open-source cloud-native platform deployed on AWS serverless infrastructure with GDPR-compliant data handling and privacy-preserving facial masking capabilities (Chapter 6). The system reduced manual annotation time 40-fold—from six weeks to six hours for 118 videos—while maintaining processing scalability linear with video duration.

This thesis demonstrates that artificial intelligence can successfully automate assessment of eating behavior and urban food landscapes, transforming labor-intensive manual processes into scalable, objective measurement systems. Current limitations include geographic constraints in nutritional databases, cultural bias toward Western dietary patterns, and insufficient training data for drinking behaviors (Chapter 7). Future work should prioritize expanding cultural data representation, improving model generalizability across diverse eating contexts, and developing integrated multimodal AI systems for comprehensive behavioral assessment. By bridging food environments, eating behavior, and health outcomes, this thesis positions artificial intelligence as a foundational methodology for scalable, data-driven nutrition science, enabling analyses previously infeasible due to cost, time, and scalability constraints.

See also these dissertations

We print for the following universities