
Grok 4.7 Teardown: Cheap, Mid-Pack, and Still Behind
xAI's newest model undercuts Claude and GPT-6 on price but trails both by seven points on the Artificial Analysis Intelligence Index — and the agentic-coding gap depends heavily on which benchmark harness is doing the measuring.
Bottom line: Skip it for agentic coding and frontier-reasoning work. Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index versus 53 for both Claude Fable 5.1 and GPT-6 — the one number that decides this call. Its sole edge, $2/$6 per million tokens, only pays off for cost-sensitive, non-agentic tasks where mid-pack reasoning is acceptable.
The Weights Desk · 4 min read- Bottom line: Skip Grok 4.7 for agentic coding and frontier-reasoning workloads — the Intelligence Index gap (46 vs. 53) is the one number that decides it.
- Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index, seven points behind Claude Fable 5.1 and GPT-6, both at 53.
- The Decoder's comparison puts Grok 4.7 at 26% on Terminal-Bench 4.0, versus 55% for Claude Fable 5.1 and 60% for GPT-6 Astra — a wider gap than the general index implies.
- Artificial Analysis's own article reports a different 33% Terminal-Bench 4.0 component score when Grok 4.7 runs inside xAI's Grok Build agentic harness at 'xhigh' effort, illustrating how harness choice moves agentic benchmark numbers.
- Pricing is aggressive at $2 per million input tokens and $6 per million output tokens, but Grok 4.7 burns roughly 81,000 output tokens per Intelligence Index task versus 36,000 for Grok 4.6, eating into the headline price advantage.
xAI shipped Grok 4.7 on September 21, 2026, pitching it as the company's most capable model yet for coding and knowledge work, and pricing it at $2 per million input tokens and $6 per million output tokens — well under Western frontier rates. But on Artificial Analysis's Intelligence Index, the model lands at 46, seven points behind both Claude Fable 5.1 and GPT-6, which score 53. Bottom line: Skip it for agentic coding and frontier-reasoning work; the discount only pencils out for tasks that don't need frontier judgment.
The headline number: mid-pack, not frontier
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index (v4.3.2), a composite of reasoning, knowledge, and math benchmarks. That's a two-point gain over its predecessor, Grok 4.6, but still seven points shy of Claude Fable 5.1 and GPT-6, which both score 53, The Decoder reported. Artificial Analysis's own write-up confirms the 46 figure and notes it pulls xAI into the 'top 4 AI labs' by its ranking — a claim that says more about how thin the field above 46 is than about Grok 4.7 closing real ground on the leaders.
Agentic coding: the gap depends on who's measuring
Here the sourcing gets genuinely messy, which matters more than the headline score. The Decoder's comparison table puts Grok 4.7 at 26% on Terminal-Bench 4.0, against 55% for Claude Fable 5.1 and 60% for GPT-6 Astra — a much wider gap than the Intelligence Index alone suggests. Artificial Analysis's own benchmarking article, however, reports a materially different 33% for the same Terminal-Bench 4.0 component, measured as part of a 56-point Coding Agent Index score when Grok 4.7 runs inside xAI's Grok Build harness at 'xhigh' reasoning effort — still trailing the top three, but a smaller gap than The Decoder's figure implies.
Why the discrepancy is the real story
Two numbers for the same benchmark, from two credible outlets citing the same evaluator, is not a rounding error — it's a methodology tell. Terminal-Bench 4.0 scores swing depending on whether a model runs bare or wrapped in an agentic scaffold with tool access and extended reasoning budget; Grok Build's xhigh configuration lifted Grok 4.7's component score from 18% (on Grok 4.6) to 33%, per Artificial Analysis. Readers comparing agentic coding claims across vendors should ask which harness produced the number before trusting a single headline percentage — a point this desk could not fully resolve without both outlets' raw run logs.
Price is real, but so is the token bill
At $2 per million input tokens and $6 per million output tokens, Grok 4.7 is priced closer to Chinese open-weight models than to Western frontier rates, The Decoder noted. But Artificial Analysis also flagged that Grok 4.7 burns roughly 81,000 output tokens per Intelligence Index task, more than double Grok 4.6's 36,000 — meaning the effective cost per completed task narrows some of the headline discount, even before accounting for the accuracy gap against Claude Fable 5.1 and GPT-6.
The bottom line
Bottom line: Skip Grok 4.7 for agentic coding pipelines or any workload where frontier reasoning quality is the point. The one number that decides it — 46 on the Artificial Analysis Intelligence Index versus 53 for Claude Fable 5.1 and GPT-6 — holds regardless of which Terminal-Bench figure you trust. The one scenario worth a pilot: high-volume, cost-sensitive, non-agentic tasks — bulk classification, drafting, summarization — where mid-pack reasoning is tolerable and the per-token price still clears the bar after accounting for its heavier token usage.
- Use it, wait, or skip it?
- Skip it for agentic coding and frontier-reasoning work. The decisive number: Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index against 53 for Claude Fable 5.1 and GPT-6. Consider it only for high-volume, cost-sensitive, non-agentic tasks where the $2/$6-per-million-token price offsets mid-pack reasoning.
- Is Grok 4.7 a frontier-class model?
- No. It scores 46 on the Artificial Analysis Intelligence Index, seven points below Claude Fable 5.1 and GPT-6, which both score 53 — mid-pack, not frontier.
- Why do Grok 4.7's coding benchmark numbers differ between origins?
- The Decoder cites 26% on Terminal-Bench 4.0; Artificial Analysis's own article reports 33% for the same benchmark component when measured through xAI's Grok Build agentic harness at higher reasoning effort — the gap reflects harness configuration, not a factual error.
- Who should actually consider using Grok 4.7 today?
- Teams with high-volume, cost-sensitive, non-agentic workloads — bulk classification, drafting, summarization — where mid-pack reasoning is acceptable. Teams building agentic coding pipelines should stay with Claude Fable 5.1, GPT-6 Astra, or Claude Opus 5.
- xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6 — The Decoder
- Benchmarking Grok 4.7 — Artificial Analysis
- Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index — Artificial Analysis