Back to feed
News Story
APriority81
量子位
2 sources

Musk's Grok 4.6 Returns to Top Tier! Lower Price Beats Fable 5, Cursor Acquisition Pays Off

SpaceXAI released Grok 4.6, which outperforms GPT-5.6 Sol and Fable 5 Max on several benchmarks while being cheaper. The new model is integrated into Grok Build, Cursor, Grok Bot, and API, with a focus on long-horizon agent tasks. The Grok Bot agent product, released the day before, now supports Grok 4.6.

SynthePulse Insight · AI deep readingMembers

Grok 4.6 Back in the Game: Low-Price Overtake, Long-Horizon Agents, and Musk's Combined Moves

Version 1 · 1 source

SpaceXAI releases Grok 4.6, overtaking GPT-5.6 Sol and Fable 5 Max on several benchmarks at a lower price, and simultaneously launches Grok Bot, targeting long-horizon agent tasks. This marks the first coordinated effort after Musk integrated xAI and Cursor.

  • Grok 4.6 scores 1753 on GDPVal-AA v2, surpassing GPT-5.6 Sol's 1728 and Fable 5 Max's 1741.
  • Priced at only $2 per million input tokens and $6 per million output tokens, lower than GPT-5.6 Sol Max and Claude Opus 5, but doubles for contexts over 200k tokens.
  • Grok 4.6 scores only 26% on Terminal-Bench v3.0, below GPT-5.6 Sol's 34.6%, with pure terminal operations still a weak point.
Open section navigationLow Price and Benchmark Scores: Grok 4.6's Competitiveness

Low Price and Benchmark Scores: Grok 4.6's Competitiveness

SpaceXAI released Grok 4.6 on August 13. Official benchmark tables show it scored 1753 on GDPVal-AA v2, surpassing GPT-5.6 Sol's 1728 and Fable 5 Max's 1741; it also ranked first on AA-Briefcase and Harvey LAB with scores of 1577 and 15.8%, respectively. On the comprehensive AA Intelligence Index, Grok 4.6 scored 61, tying with GPT-5.6 Sol, 5 points higher than Grok 4.5, and only 1 point lower than Fable 5 Max.

In terms of pricing, Grok 4.6 costs $2 per million input tokens and $6 per million output tokens, matching Grok 4.5 and far below GPT-5.6 Sol Max's $5/$30 and Claude Opus 5's $5/$25. However, note that for single requests exceeding 200k tokens in context, the input and output prices double to $4/$12, and the entire request is billed at the higher rate. Even with the doubling, it remains cheaper than competitors. There is also a low-latency version priced at twice the standard rate.

However, on Terminal-Bench v3.0, Grok 4.6 only achieved 26%, below GPT-5.6 Sol's 34.6%. This benchmark tests agent operations in a pure command-line environment, where the model must execute multi-step instructions without a graphical interface, and any error interrupts the process. This indicates that Grok 4.6 is more adept at knowledge-work agentic tasks, while pure terminal operations have not yet caught up.

Free for now

Read the full analysis

3 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article's information primarily comes from reports by QbitAI, a secondary source. All benchmark scores, prices, and acquisition amounts are based on official SpaceXAI releases or reported figures, not independently verified. The release timeline and performance claims for Grok 4.7 are solely Musk's statements and carry uncertainty.

Primary report

量子位

Primary source

Same-event coverage

Also covered by 1 sources