benchmarks.
5 writings found
Latest Archives
Claude Fable 5.1: When Reasoning Effort Actually Matters
Testing Anthropic's new Fable 5.1 across reasoning levels reveals a wild cost/quality tradeoff. The pelican benchmark shows reasoning might not be what we think it is.
Why Your ASR Model's Leaderboard Score Is Lying to You
High benchmark scores hide critical failures in speech recognition. The Monsoon dataset exposes why measuring accuracy matters less than measuring whose accuracy.
Why ASR Benchmarks Hide the Real Problem
Speech recognition leaderboards measure what's easy to measure, not what matters. A new dataset exposes how models fail differently across populations.
When Benchmarks Break: A Laptop Model Drew Better Pelicans Than Claude Opus
A quantized 21GB model running locally outperformed Anthropic's flagship on SVG generation. What this tells us about AI benchmarks and model comparison.
When Benchmark Performance Stops Meaning What We Think It Means
A quantized local model outdraws Claude Opus 4.7 at pelicans on bicycles. What does that tell us about AI benchmarks? Probably nothing good.