JIM 2026;
3 (3): e1211
DOI: 10.61012/JIM_202608_1211
How reliable are large language models in metabolic dietetics? a cross-sectional benchmarking evaluation
Topic: Dietetics
Category: Original article
![]()
Abstract
Objective: The aim of this study was to evaluate the performance of free-tier large language models (LLMs) in providing dietary recommendations for inherited metabolic disorders (IMDs), with a focus on scientific accuracy, completeness, and response consistency of generated responses across models under real-world conditions.
Materials and Methods: This cross-sectional benchmarking study assessed four LLMs – ChatGPT, GitHub Copilot, Gemini and Claude – using 30 standardized clinical questions on dietary management in IMDs. Models were tested under real-world free-tier conditions, without fine-tuning or prompt engineering. Each question was submitted three times to each model. Responses were independently evaluated by three metabolic dietitians using a 6-point ordinal scale for scientific accuracy and completeness, while response consistency was assessed across repeated outputs. Inter-rater agreement was assessed using the intraclass correlation coefficient (ICC).
Results: A total of 359 unique responses were collected for evaluation. Safety guardrails were activated in 18.4% of responses, with substantial variability across models and clinical domains. Of these, 80.3% resulted in a complete block and were excluded from the analysis. Consequently, the final sample consisted of 306 responses. Model performance was heterogeneous: ChatGPT achieved the highest and most consistent scores, with no severe inaccuracies or omissions. Copilot, Gemini and Claude showed intermediate performance, with relevant variability. Across clinical domains, responses related to diet planning and clinical monitoring were generally more stable, whereas dietary supplementation and nutritional requirements showed greater variability and clinically relevant errors. Poor response consistency was observed in 64.7% of responses, indicating limited consistency across repeated identical prompts. Inter-rater agreement is defined as poor for accuracy, completeness, and response consistency.
Conclusions: The performance of free-tier LLMs in IMD dietary management is heterogeneous, with substantial limitations in response consistency and in clinically complex scenarios. These findings indicate that current models are not yet suitable for autonomous clinical use and can be dangerous to use without a healthcare professional’s oversight.
![]()
Graphical Abstract
To cite this article
How reliable are large language models in metabolic dietetics? a cross-sectional benchmarking evaluation
JIM 2026;
3 (3): e1211
DOI: 10.61012/JIM_202608_1211
Publication History
Submission date: 04 Jun 2026
Revised on: 29 Jun 2026
Accepted on: 16 Jul 2026
Published online: 31 Aug 2026