1 writing found
Testing Anthropic's new Fable 5.1 across reasoning levels reveals a wild cost/quality tradeoff. The pelican benchmark shows reasoning might not be what we think it is.