Google announced Gemini 4 Argon on September 30, 2026. It is the company's first new high-end Gemini model since February.
Independent testing places it on a level with GPT-6 Astra. Google chose its own benchmarks, so those results need careful reading.
Access is limited to trusted cyber defenders for now. No public model card has appeared, and the price doubles after the introductory period.
Google finally has a new top-tier AI model. Most people cannot use it yet. Gemini 4 Argon arrived on September 30, 2026, after months of delay. Early scores look strong. Several key details remain unpublished.
Gemini 3.1 Pro Preview arrived in February. Gemini 3.5 Pro was planned for May. It never shipped. Google released cheaper Flash models instead. Rivals kept shipping new models in the meantime. Gemini 4 Argon is now Google's first new high-end Gemini model since February.
Gemini 4 Argon is the first model in Google's Gemini 4 generation. Google calls it a frontier model, placing it among its most advanced AI systems. It is a reasoning model designed for long, multi-step tasks.
Examples include legal drafting, financial research, and large code migrations. It reads text and images, then writes text. Google says thousands of its staff already use Argon. In one project, Argon agents made a video decoder run 2.7 times faster than the earlier version.
Independent testing offers a useful counterpoint to Google's own charts. Google chose its benchmarks. Those numbers deserve careful reading.
| Measure | Result | Source |
|---|---|---|
| Intelligence Index | 53, level with GPT-6 Astra | Artificial Analysis |
| Gap to Gemini 3.1 Pro Preview | 23 points higher | Office Chai summary |
| Hallucination rate | 15%, lowest among models scoring 45 or higher | Artificial Analysis via OfficeChai |
| DeepSWE v1.1 (software engineering) | 77.90% | |
| LVBench (long video) | 91.70% | |
| Automation Bench | 51.3%, first place | |
| CWE-bench v1 (security fixes) | 68%, tied for first |
Argon trailed rivals on terminal-based tasks. Launch coverage counts wins in 12 of 18 tests in Google's table. A hallucination is a confident answer that is false. Benchmarks measure narrow skills. They do not guarantee results on real work.
A token is a small piece of text. Cached input tokens cost 95% less.
| Period | Input per million tokens | Output per million tokens |
|---|---|---|
| Introductory | USD 2 | USD 10 |
| After the introductory period | USD 4 | USD 20 |
Google has not said when the introductory period ends. Artificial Analysis puts Argon's cost per task at about 60% of GPT-6 Astra's under the discounted price. A doubled price would likely erase that edge if usage stays the same. That is an inference, not a company figure.
Argon is rolling out to trusted cyber defenders through Google's Fairwind Program. Google is also taking part in the U.S. government's voluntary pre-release access process. Paid API customers and Google AI Ultra subscribers come next. Developers, enterprises, and consumers follow. Google gives no date.
Google says security firm Wiz used Argon to find a critical flaw in hospital software that earlier models had missed. Trusted defenders and Google's internal teams will get a version with the cyber guardrails removed. The announcement does not explain what those guardrails cover.
Google's announcement does not link to a public model card. Launch reports reviewed for this article did not find one either. A model card normally lists a model's abilities, limits, test results, and safety details. Google has also kept the model's size private. The company does describe safeguards.
These include monitoring of the model's reasoning and actions. They also include defenses against prompt injection, where hidden instructions try to hijack a model. Google says internal and external red teams tested them. No public report of the results turned up.
Also Read: Google Unveils Gemini 4 Argon, its Most Powerful AI Model Yet
Google confirms a 1 million token output limit. The earlier limit was 64,000 tokens. This is how much the model can write in one response. It is not the context window, which is how much text a model can read at once. Google has not stated that figure. Third-party listings that show 1 million are unverified.
Artificial Analysis also tested Argon using long decode continuation. This Gemini API feature pauses long responses and resumes them across calls. Google's announcement does not mention it. Both details suggest a focus on long, multi-step work rather than one-shot answers. That reading is my own.
Argon puts Google back among the leading AI labs. Independent scoring supports that view. The next few months will show whether the lead holds. Rivals are likely to answer with new releases. Developers will test Argon on real codebases.
Companies will compare their cost per task once the price rises. Researchers will look for the model card and the red team results. The cyber-first rollout also sets a pattern. Powerful defensive tools go to vetted users before the public. Other labs may follow the same path.
The next phase of the AI race may depend less on raw ability. Buyers will ask who can prove that an agent stays inside its limits. Labs that publish clear evidence of control will win the largest contracts. Argon is an early test of that idea.
How to Use Gemini for Coding: Debugging, Explaining Code, and Scripts
Is Google Gemini AI Safe? Security, Privacy & Ethical Concerns Explained
1. What is Gemini 4 Argon?
Gemini 4 Argon is the first model in Google's Gemini 4 generation, announced on September 30, 2026. It is a reasoning model designed for long, multi-step tasks such as software engineering, legal and financial work, and cyber defense. It reads text and images and writes text.
2. How well does it perform against rivals?
Artificial Analysis scored it 53 on its Intelligence Index, level with GPT-6 Astra at its top setting. Google also reports first place on several benchmarks, including AutomationBench. Google chose those tests, so the results should be read with care.
3. How much does Gemini 4 Argon cost?
The introductory API price is $2 per million input tokens and $10 per million output tokens. Google says the rates rise to $4 and $20 after the introductory period. It has not said when that period ends.
4. Who can use it right now?
Access is limited to trusted cyber defenders through Google's Fairwind Program. Paid API customers and Google AI Ultra subscribers come next. Developers, enterprises and consumers follow, with no date given.
5. Does the 1 million token limit mean it can read 1 million tokens?
No. Google confirms a 1 million token output limit, which is how much the model can write in one response. The context window, which is how much it can read at once, has not been stated by Google.