APriority76
Artificial Analysis (X)
2 sourcesGrok 4.6 Shows Big Gains on AA-Briefcase Benchmark
Grok 4.6 achieved significant improvements on the AA-Briefcase benchmark, which tests long-horizon agentic knowledge work tasks. The model performs on par with Claude Fable 5 but at a much lower cost, at $4.42 per task compared to Claude Fable 5's $22.30. This release highlights xAI's competitive edge in both performance and cost efficiency.