Show HN: A new benchmark for testing LLMs for deterministic outputs
A new structured output benchmark (SOB) reveals that most LLM evaluation frameworks miss critical production risks—existing benchmarks validate JSON schema compliance but fail to catch hallucinated values that silently break downstream systems. The benchmark demonstrates that while leading models (GPT-5.4, GLM-4.7, Qwen3.5) achieve 85%+ overall scores, value accuracy—the metric that matters for production—ranges from 75-80%, meaning organizations relying on LLMs for data extraction from invoices, medical records, and PDFs should expect 20-25% of fields to require human review without additional safeguards. This represents a critical gap between perceived model reliability and actual operational fitness for deterministic structured output tasks that enterprises increasingly depend on.