SpaceXAI released Grok 4.6 on August 13. Official benchmark tables show it scored 1753 on GDPVal-AA v2, surpassing GPT-5.6 Sol's 1728 and Fable 5 Max's 1741; it also ranked first on AA-Briefcase and Harvey LAB with scores of 1577 and 15.8%, respectively. On the comprehensive AA Intelligence Index, Grok 4.6 scored 61, tying with GPT-5.6 Sol, 5 points higher than Grok 4.5, and only 1 point lower than Fable 5 Max.
In terms of pricing, Grok 4.6 costs $2 per million input tokens and $6 per million output tokens, matching Grok 4.5 and far below GPT-5.6 Sol Max's $5/$30 and Claude Opus 5's $5/$25. However, note that for single requests exceeding 200k tokens in context, the input and output prices double to $4/$12, and the entire request is billed at the higher rate. Even with the doubling, it remains cheaper than competitors. There is also a low-latency version priced at twice the standard rate.
However, on Terminal-Bench v3.0, Grok 4.6 only achieved 26%, below GPT-5.6 Sol's 34.6%. This benchmark tests agent operations in a pure command-line environment, where the model must execute multi-step instructions without a graphical interface, and any error interrupts the process. This indicates that Grok 4.6 is more adept at knowledge-work agentic tasks, while pure terminal operations have not yet caught up.