Back to feed
News Story
SSignal87
Latent Space
1 sources

SpaceXAI Releases Grok 4.6 and Grok @Bot

SpaceXAI has released Grok 4.6, a 1.5T parameter model focused on long-running agents and interactive visual work, along with the AI teammate product Grok @Bot, which has received positive reviews. The model is considered the second best knowledge work model in the world and the most efficient.

SynthePulse Insight · AI deep reading

Grok 4.6 and Grok Bot: A Watershed Moment in the AI Teammate Arena

Version 1 · 1 source

SpaceXAI releases Grok 4.6 model and Grok Bot product, entering the knowledge work arena with efficiency and price advantages, but capability ceilings and training details remain in question.

  • Grok 4.6 is positioned as the world's second-best knowledge work model, but the most efficient, with prices far below comparable frontier models.
  • Grok Bot enters early testing as an AI teammate that can log into tools and complete real work, receiving positive reviews.
  • Grok 4.6 is a 1.5T parameter model, with training focused on long-running agents and interactive visual work.
  • Independent evaluations show Grok 4.6's intelligence index close to GPT-5.6 Sol Max, but behind Claude Opus/Fable.
  • Grok 4.7 is already in training, with plans for supplementary training using SpaceX internal data.
Open section navigationRelease Context and Market Positioning

Release Context and Market Positioning

Amid mixed reviews of Claude Tag and Block's Buzz requiring more technical users, the AI teammate arena still lacks a clear leader. The team that moved from Cursor to SpaceX has launched Grok Bot, entering early testing with very positive reviews.

Grok Bot is described as an AI teammate that can log into user tools, operate like the user, and return completed work. The product is powered by the latest model, Grok 4.6, released the same day, which is acknowledged by competitor Cognition and Elon himself as the world's second-best knowledge work model, but undoubtedly the most efficient.

Model Capabilities and Evaluation Data

Grok 4.6 is a confirmed 1.5T parameter model, with official statements noting it 'builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.' Training details reveal it underwent longer supplementary training than Grok 4.5, using curated model-generated data, high-quality engineering data, and improved optimizer and training recipe.

Independent evaluator Artificial Analysis shows Grok 4.6 scores 61 on the intelligence index, roughly on par with GPT-5.6 Sol Max, but behind Claude Opus/Fable. On agent tasks, it achieves 88.4% on Terminal-Bench v2.1, GDPval-AA v2 Elo of 1753, and competitive performance on AA-Briefcase, but at a much lower cost than other leading models.

Pricing for Grok 4.6 is $2/$6 per million tokens for input/output, significantly lower than frontier peers, and practitioners immediately view it as the new default for coding and bug-finding workloads.

Training Methodology and Self-Testing Behavior

Officially disclosed training methods include: using Grok 4.5 to regenerate SFT trajectories covering reasoning, agent tools, STEM, software engineering, and knowledge work, with model-based checks to filter problematic trajectories. This is followed by training on a broad range of agent RL tasks, including knowledge work, general coding, kernel optimization, web development, and computer-aided design.

Notably, xAI reports the model exhibits more self-testing behavior during long tasks. However, it remains unclear whether sandbox escape incidents occurred during training, which commentators jokingly refer to as a test for infrastructure engineers or researchers.

Competitive Landscape and Future Outlook

The same day saw other frontier model releases: Qwen3.8-Max as an open-weight model with 2.4T total parameters and 95B activated, but the initial version is text-only; DeepSeek V4 Pro GA enters the market at extremely low prices ($0.435/M input, $0.87/M output) but with mixed capability reviews; Microsoft introduces its first reasoning model, MAI-Thinking-1, emphasizing tool-use feedback.

Elon states Grok 4.7 is already in development with initial training complete, and plans to use SpaceX internal data for supplementary training. This indicates SpaceXAI is iterating rapidly and may leverage unique data advantages.

Credibility boundary

This report is based on Latent Space's AINews summary, which includes official disclosures, independent evaluations, and community commentary. Official training details are confirmed facts; evaluation data comes from third parties like Artificial Analysis but has not been independently verified. Some assessments (e.g., 'second best') come from competitor and Elon's acknowledgment, constituting source claims.

Insight takeaway

The release of Grok 4.6 and Grok Bot marks a new phase in the AI teammate arena, capturing market share with efficiency and price advantages, but capability ceilings and training details still require further verification.

Primary report

Latent Space

Primary source