When LLM judges agree, should we believe them?
The article argues that CIOs and AI leaders should not treat agreement among multiple LLM judges as inherently trustworthy, because correlated models can produce false confidence when their outputs are driven by shared training, prompts, or model families. By using dependence-aware aggregation with Ising models, organizations can better distinguish independent evidence from repeated mistakes, improving evaluation accuracy by 9% to 14% in tests and making LLM-as-a-judge systems more reliable for production use. For IT organizations, the key implication is that AI quality assurance and model governance should incorporate statistical checks for judge diversity and correlation rather than relying on simple majority voting.
Hacker News3 min read
