He asked AI to count carbs 27000 times. It couldn't give the same answer twice
A study of 27,000 AI queries reveals that leading AI models (GPT, Claude, and Gemini) produce inconsistent carbohydrate estimates for the same food images, with variations large enough to cause dangerous insulin dosing errors in diabetes management applications. The research identifies two critical failure modes: systematic bias that consistently over/underestimates carbs, and unpredictable variability where a single query can produce catastrophic outliers—Claude performs best with 100% of estimates in safe ranges, while Gemini 2.5 Pro shows 12% of queries posing severe hypoglycemia risk. For IT organizations, this demonstrates that AI models cannot yet be safely deployed in high-stakes, health-critical applications without additional safeguards, and highlights the need for rigorous testing, transparency about model limitations, and human-in-the-loop verification systems before adopting AI in regulated healthcare environments.
