
Gemini 4 Argon ties GPT-6 Astra on Artificial Analysis
Google's first frontier model in over seven months lands in the top tier of independent testing without taking the lead. It uses more than twice the output tokens of GPT-6 Astra per task, and its cost advantage depends on introductory pricing.
Gemini 4 Argon is a credible but not leading frontier model: it scores 53 on the Artificial Analysis Intelligence Index, level with GPT-6 Astra and behind Claude Opus 5.5 at 58. It uses more than twice Astra's output tokens per task, yet at promotional prices costs less per task. Access remains restricted.
The Weights Desk · 3 min read- Artificial Analysis lists Gemini 4 Argon (High) at 53 on its Intelligence Index, ranked #8 of 223 models. The Decoder reports GPT-6 Astra at the same 53 and Claude Opus 5.5 ahead at 58.
- The Decoder reports Argon averages 62,000 output tokens per task against 27,000 for GPT-6 Astra, more than twice as many.
- At introductory pricing, Artificial Analysis puts Argon at $1.99 per Intelligence Index task. The Decoder reports $3.26 for GPT-6 Astra, so Argon is cheaper per task for now.
- Standard pricing is double the introductory rate ($4 input and $20 output per million tokens), so the per-task advantage should be recomputed once promotional pricing ends.
- Google's headline claims (DeepSWE 77.9%, CWE-bench tie for first at 68%, Gray Swan prompt-injection leadership) are vendor-reported and were not independently checked here.
Google's Gemini 4 Argon reaches the top tier of frontier models without taking first place. Artificial Analysis scores it 53 on its Intelligence Index, and The Decoder reports that GPT-6 Astra matches that score while Claude Opus 5.5 leads at 58. It is Google's first frontier model in over seven months. The model uses more than twice as many output tokens per task as Astra, but at introductory prices it still costs less per task. Access is limited to select testers for now.
Where Argon lands on the independent index
Artificial Analysis lists Argon (High) at 53 on its Intelligence Index, ranked eighth of 223 models. The Decoder reports GPT-6 Astra at the same 53 and Claude Opus 5.5 at 58, a five-point lead. The Decoder also reports Argon at 77.5 percent on AutomationBench-AA and first place on the Arena.ai text leaderboard with 1,525 points. These figures come from the Artificial Analysis page and The Decoder's report. We did not reproduce the benchmarks.
Google's own results are narrower and vendor-chosen
Google's announcement cites 77.9 percent on DeepSWE v1.1, a tie for first at 68 percent on CWE-bench v1 for vulnerability remediation, and leadership on the Vals Index. It also claims leadership on the Gray Swan prompt-injection benchmark but gives no score. These are vendor-selected results, not independent measurements. They do not contradict the index tie, because a vendor picks the benchmarks it publishes and an aggregate weights many tasks. For security teams, the Gray Swan claim needs a published score before it can inform a decision.
Token use is high, but promotional pricing still wins per task
The Decoder reports that Argon averages 62,000 output tokens per task while GPT-6 Astra needs 27,000. Even so, Artificial Analysis lists $1.99 per Intelligence Index task for Argon at introductory rates, and The Decoder reports $3.26 for Astra. Argon is therefore cheaper per task today. That depends on the promotion: standard pricing is $4 per million input tokens and $20 per million output tokens, double the $2 and $10 introductory rates, so the comparison should be redone when pricing changes.
Access is restricted and the cyber positioning is unproven
Google says Argon is rolling out to trusted cyber defenders through its Fairwind Program, with developers, enterprises and consumers to follow, starting with paid API customers and Google AI Ultra subscribers. The post gives no general-availability date. The cyber-defense framing rests on Google's own CWE-bench and Gray Swan claims. No independent red-team evaluation of Argon appears in the origins reviewed. Buyers in security-sensitive deployments should treat the positioning as a vendor claim until reproducible results are published.
Verdict: Signal. Argon puts Google back in the frontier group, which matters for multi-model routing and vendor negotiation. It does not beat Claude Opus 5.5 on the independent index, and its per-task cost edge over Astra is tied to introductory pricing. Once access opens, run your own evaluation with token counts logged and compare cost per completed task at standard rates, not at the promotional rate.
- Does Gemini 4 Argon beat OpenAI and Anthropic?
- Not on the independent aggregate. Artificial Analysis scores Argon at 53, and The Decoder reports GPT-6 Astra at the same score and Claude Opus 5.5 at 58. Google claims leading results on some benchmarks it chose to publish, which is a narrower claim than overall leadership.
- Is Gemini 4 Argon cheaper to run than GPT-6 Astra?
- At promotional rates, yes. Artificial Analysis lists $1.99 per Intelligence Index task for Argon, and The Decoder reports $3.26 for GPT-6 Astra, even though Argon uses about 62,000 output tokens per task against Astra's 27,000. Regular pricing is double the introductory rate, so the gap is likely to narrow.
- Who can use Gemini 4 Argon today?
- Google says it is rolling out to trusted cyber defenders through its Fairwind Program. Developers, enterprises and consumers follow, starting with paid API customers and Google AI Ultra subscribers. Google's post gives no firm general-availability date.