xAI officially released Grok 4.6, which scored 61 on the AI Intelligence Index, tying with GPT-5.6 Sol (Max), and ranking just behind Claude Opus 5 (63 points) and Claude Fable 5 (62 points). This achievement places Grok 4.6 in the global top-tier models, directly competing with the flagship products of OpenAI and Anthropic. The model is built upon Grok 4.5, with core improvements including reasoning and advanced technical concept training using curated data generated by the model, combined with high-quality engineering data and an improved optimizer.

A key change in the training approach is the data loop: Grok 4.5 was used to regenerate supervised fine-tuning trajectories, covering areas such as reasoning, agent tool usage, and fields like STEM, software engineering, and knowledge work. The model was also trained on a wide range of agent-based reinforcement learning tasks, including knowledge work, general coding, kernel optimization, web development, and computer-aided design, a strategy clearly aimed at the core workloads of the "agent era."

Agent benchmarks have surged dramatically, significantly reducing the effort required for long-term tasks

In agent-related benchmark tests, Grok 4.6 performed exceptionally well. GDPval-AA v2 reached 1753 Elo, second only to Claude Opus 5; DeepSWE v1.1 increased to 65.9%, a significant rise from 54% in Grok 4.5; APEX-Agents jumped from 47.1% to 57.5%; and Terminal-Bench v3.0 rose from 15.7% to 26%. Additionally, it achieved 69.9% on CursorBench v3.2 and 61.3% on FrontierCode v1.1, achieving significant breakthroughs in coding and agent capabilities.

More notably, Grok 4.6 has a cost-efficiency advantage. Its pricing remains at $2 per million input tokens and $6 per million output tokens, which is over 60% lower than that of Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). On the AA-Briefcase long-term knowledge work benchmark, Grok 4.6 completed tasks with an average of 53 rounds and about 500 million input tokens, while Claude Opus 5 required approximately 103 rounds and about 2 billion tokens—achieving the same task with less than a quarter of the resources, making the cost-performance gap obvious. The model is now available through channels such as Cursor, Grok Build, API, and OpenRouter, Vercel, Cloudflare. During the first week, there are double usage offers on Grok Build and Cursor. The context window remains at 500,000 tokens, and a cost-performance battle centered around "equal intelligence, lower prices" has already begun.