Podcast· The AI Daily Brief

The Best Way to Test New AI Models

Oct 7, 2026 · 49 min · Clio 85

Clio listened

Nathaniel and Nufar Gaspar walk CIOs through building a personal AI benchmark to test new models against real work, not public leaderboards. They show how to choose tasks, blind-test models via OpenRouter, use a separate judge model, and weigh accuracy against cost, latency, and workflow fit.

Key points

  • Public benchmarks are increasingly weak because model training data overlaps with them; real value comes from your own use-case benchmark.
  • Choose 4-6 repeatable tasks that matter, include at least one “wish list” task, and score outputs blind.
  • Run each request in a fresh chat, hide model names, and test multiple candidates plus your current baseline.
  • Use side-by-side review and optionally an AI judge, but validate the judge because it may not match your taste.
  • Decisions should be switch, split, or stay; cost, latency, company policy, and habit can outweigh small quality gains.
  • The demo compared models like Opus, GPT, Grok, Kimi, and Gemini, and showed tradeoffs between quality, speed, and expense.

Key moments

Episode page at the publisher
The Best Way to Test New AI Models - CIO Daily Brief