How good are frontier models at physics?

This study suggests frontier AI models are significantly better at physics than current benchmarks indicate, because many “failures” were actually caused by broken grading, flawed reference solutions, or ambiguous questions rather than model reasoning errors. For CIOs and technology leaders, the strategic takeaway is that benchmark scores may understate the real value of advanced models in technical and scientific workflows, but also that evaluation quality is becoming a critical governance issue as vendors and internal teams make adoption decisions based on unreliable tests. IT organizations should expect faster maturation of AI capabilities in specialized domains and shift toward more rigorous, expert-validated assessments before committing to enterprise use cases or procurement.

Hacker News3 min read
Read full article
How good are frontier models at physics?

Read the full story at Hacker News →