Podcast· The AI Daily Brief
The Best Way to Test New AI Models
Oct 7, 2026 · 49 min · Clio 85
Clio listened
Nathaniel and Nufar Gaspar walk CIOs through building a personal AI benchmark to test new models against real work, not public leaderboards. They show how to choose tasks, blind-test models via OpenRouter, use a separate judge model, and weigh accuracy against cost, latency, and workflow fit.
Key points
- Public benchmarks are increasingly weak because model training data overlaps with them; real value comes from your own use-case benchmark.
- Choose 4-6 repeatable tasks that matter, include at least one “wish list” task, and score outputs blind.
- Run each request in a fresh chat, hide model names, and test multiple candidates plus your current baseline.
- Use side-by-side review and optionally an AI judge, but validate the judge because it may not match your taste.
- Decisions should be switch, split, or stay; cost, latency, company policy, and habit can outweigh small quality gains.
- The demo compared models like Opus, GPT, Grok, Kimi, and Gemini, and showed tradeoffs between quality, speed, and expense.