xAI
Grok 4.6: what xAI's new model is actually for
Not the smartest model on the market, and not trying to be. It's built to work through long jobs without giving up — and it's a third of the price.
The answer
xAI's Grok 4.6, out 12 August 2026, is built for long multi-step tasks at $2/$6 per million tokens.
Most new AI models are announced as being cleverer than the last one. Grok 4.6, released by Elon Musk's xAI on 12 August 2026, was announced as being better at sticking with a job — which is a more useful thing than it sounds, and a good excuse to explain what 'agentic' actually means.
What 'long-running agent' means
When you ask a chatbot a question, it answers once. When you give an AI a job — 'research these twenty companies and build me a comparison', or 'work through this codebase and fix the failing tests' — it has to take many steps in a row, deciding what to do next each time, without a human checking between steps. That is an agent, and the difficulty is not intelligence so much as stamina and self-correction.
stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase
xAI says Grok 4.6 does more checking of its own work on long tasks. That matters enormously. If a model makes a mistake at step 12 of a 60-step job and never notices, everything after that is built on the error. A model that catches its own slips is worth more than one that is marginally smarter but blunders on.
How good is it, and what does it cost?
On a composite score made from nine different tests, Grok 4.6 reached 61 — level with OpenAI's GPT-5.6 Sol, and two points behind Anthropic's Claude Opus 5. That is genuine top-tier company.
The price is where it gets interesting: $2 per million tokens of input and $6 for output, against $5 and $30 for GPT-5.6 Sol. A million tokens is around 750,000 words. So you are paying roughly a third for a very similar score.
Pricing starts at $2 per million input tokens and $6 per million output tokens.
The catch
Grok 4.6 is notably weak at command-line work — the text-based way developers control computers — scoring 26% on a standard test for it. If your use involves an AI running commands and managing files, that is the thing to check carefully. For research, writing, analysis and building first drafts of software, the weakness matters much less.
There is also a reliability point worth knowing if you follow the news. xAI has been promising a successor, Grok 4.7, since late July, and it still had not arrived by mid-September. Musk said it needed more tuning because the model stopped too early on difficult tasks and did not check its own work carefully enough. All of that came from his posts rather than official documentation.
The practical advice: judge what you can use today. Grok 4.6 is a capable, well-priced model for long multi-step work that does not lean heavily on the command line. That is a real and useful thing, regardless of what is promised next month.
Why 'steps' matter more than speed
xAI makes one claim that is worth understanding properly, because it is the most useful way to compare AI models for real work. On long jobs, Grok 4.6 reportedly reached the finish in about 53 steps, where Anthropic's Claude Opus 5 took 103 for comparable work. Fewer steps means less time, less cost, and — most importantly — fewer chances to go off course.
Think of it like asking two people to assemble flat-pack furniture. One reads the instructions carefully, makes a plan, and works steadily. The other starts immediately, gets halfway, realises a panel is backwards, and undoes fifteen minutes of work. Both may finish. One of them finishes in half the time with fewer scratches. Measuring AI by how quickly it starts talking misses this entirely.
The honest caveat is that these step counts are measured by the company selling the model, using its own setup. That does not make them false, but it does mean the number to trust is the one you get when you run your own real task on it. Most providers offer enough free or cheap usage to do exactly that before committing.
Frequently asked questions
What is Grok 4.6 good for?
How much does Grok 4.6 cost?
Is Grok 4.6 as good as ChatGPT or Claude?
Is Grok 4.7 available?
Sources
- Introducing Grok 4.6 — xAI, 12 August 2026
- Grok 4.6: Features, Benchmarks, Pricing & Comparisons — DataCamp, 13 August 2026
- xAI's Grok 4.6: Benchmarks, Pricing, and the Full Developer Ecosystem — Better Stack, 14 August 2026
- xAI delays Grok 4.7 as Musk points to reinforcement-learning and self-checking problems — DataStudios, 14 September 2026
- Space Exploration Technologies Corp — Form S-1 (FY2026) — U.S. Securities and Exchange Commission, 20 May 2026
- Anthropic will pay xAI $1.25B per month for compute — TechCrunch, 20 May 2026