Back to feed
News Story
APriority76
Artificial Analysis (X)
2 sources

Grok 4.6 Shows Big Gains on AA-Briefcase Benchmark

Grok 4.6 achieved significant improvements on the AA-Briefcase benchmark, which tests long-horizon agentic knowledge work tasks. The model performs on par with Claude Fable 5 but at a much lower cost, at $4.42 per task compared to Claude Fable 5's $22.30. This release highlights xAI's competitive edge in both performance and cost efficiency.

Primary report

Artificial Analysis (X)

Primary source

Same-event coverage

Also covered by 1 sources