Google Launches Gemini 4 Argon AI Model: Features, Pricing, Availability and Benchmark Concerns
Google has introduced Gemini 4 Argon, a new artificial intelligence model designed for complex tasks involving software development, cybersecurity, law and finance. Announced on September 30, 2026,...
Google has introduced Gemini 4 Argon, a new artificial intelligence model designed for complex tasks involving software development, cybersecurity, law and finance. Announced on September 30, 2026, the model is described by Google as its most powerful AI model to date. However, access remains restricted, and independent evaluations have raised questions about some of its benchmark results.
Gemini 4 Argon is Google’s first flagship model since Gemini 3.1 Pro, which was introduced in February. The company had previously indicated that Gemini 3.5 Pro would arrive in June, but the model was never released. According to Reuters, Google no longer plans to launch it.
Who Can Access Gemini 4 Argon?
Google is initially making Gemini 4 Argon available to selected cybersecurity experts through its Fairwind Program, which began on September 2, 2026, and has more than 650 partners.
The next phase is expected to include paid API customers and Google AI Ultra subscribers. Google plans to expand access to developers, businesses and general consumers later.
The company has not announced a public release date or published a public API model ID. Google is also participating in a voluntary US government pre-release review process for advanced AI models.
Gemini 4 Argon: Key Features and Performance
Google says Gemini 4 Argon has been designed to handle demanding, long-running tasks, particularly in software engineering and cybersecurity.
Its key claimed capabilities include:
- One million token context window: The model supports up to one million tokens of input and output, compared with the previous limit of 64,000 tokens cited in the supplied report.
- DeepSWE v1.1: Argon achieved a reported score of 77.9% on this benchmark for evaluating complex software engineering tasks.
- AutomationBench: Google reported a score of 51.3%, placing the model first on the benchmark.
- LVBench: The model achieved 91.7% on a benchmark designed to evaluate understanding of long videos.
- CWE-bench v1: Argon reportedly achieved 68% on a benchmark measuring the ability to identify and address software vulnerabilities, matching the leading score.
A token is a small unit of text processed by an AI model. Depending on the language and content, a token may represent part of a word, a complete word or another text element.
Google also claims that Argon can improve infrastructure efficiency. The company says the model helped free up more than 300 TiB of memory in its data centres. In another reported engineering task, the model converted the code of the libgav1 video decoder to Rust and produced a version that ran 2.7 times faster than the existing Rust implementation.
These performance figures are company-reported results and should be considered alongside independent evaluations.
How Gemini 4 Argon Is Being Used for Cybersecurity
Cybersecurity is one of the main areas Google is targeting with Gemini 4 Argon. The model has been trained to identify software vulnerabilities and assist with fixing them.
Google says selected security researchers and its internal teams will be able to use the model without the usual cybersecurity guardrails. Such unrestricted access is intended for trusted experts working on security research and vulnerability remediation.
According to Google, cybersecurity company Wiz used Argon to discover a serious vulnerability in healthcare software used by hospitals around the world. The flaw reportedly put sensitive personal information at risk of exposure.
The example highlights the potential benefits of advanced AI in vulnerability discovery. However, the details provided do not independently establish the full scope of the vulnerability or its remediation status.
Independent Testing: How Does Gemini 4 Argon Compare With Other AI Models?
According to the Artificial Analysis Intelligence Index results cited in the report, Gemini 4 Argon received a score of 53.
| AI model | Intelligence Index score |
|---|---|
| Claude Opus 5.5 | 58 |
| Claude Sonnet 5.5 | 56 |
| Gemini 4 Argon | 53 |
| GPT-6 Astra | 53 |
These figures place Argon alongside GPT-6 Astra while trailing the listed Claude models on this evaluation.
Artificial Analysis also reported a hallucination rate of 15% for Argon, which it described as the lowest among the major models it evaluated. Hallucination occurs when an AI system generates information that is incorrect or unsupported but presents it as factual.
Benchmark scores do not necessarily predict how a model will perform in every real-world situation. Results can vary depending on the task, evaluation method and amount of computing used.
Gemini 4 Argon Pricing
Google has announced introductory API pricing for Gemini 4 Argon, with lower rates available during an initial promotional period.
| API usage | Introductory price | Standard price |
|---|---|---|
| Input | $2 per million tokens | $4 per million tokens |
| Output | $10 per million tokens | $20 per million tokens |
| Cached input | 95% cheaper than regular input | Subject to applicable pricing |
Google has not announced when the introductory pricing will end.
The cost of using Argon may also be higher than the per-token rates alone suggest. Tasks requiring extensive reasoning or lengthy responses can consume significantly more tokens than simpler requests.
Artificial Analysis estimated an average cost of $1.99 per task for Argon at introductory prices and $3.98 at standard prices. For comparison, the reported average costs were $3.26 for GPT-6 Astra, $5.98 for Claude Opus 5.5 and $0.72 for GPT-6.1 Sol.
According to the report, Argon used approximately 62,000 output tokens per task, compared with around 27,000 for GPT-6 Astra. This higher token consumption helps explain why the overall cost per task can differ substantially between models.
These figures reflect the specific tasks and evaluation methods used by Artificial Analysis and should not be treated as universal costs for every workload.
Why Are Questions Being Raised About Gemini 4 Argon’s Benchmarks?
Despite its reported performance, Gemini 4 Argon has faced questions about the reliability of some of its benchmark results.
One of its most prominent coding results comes from DeepSWE v1.1. Epoch AI has described the benchmark as flawed and reported confirming evaluation problems in at least 23 of 113 tasks, representing approximately 20.3% of the examined tasks.
The benchmark had also been publicly available since May, raising questions about whether the evaluation adequately measures performance under conditions that prevent models from benefiting from prior exposure to the test.
Another limitation is that the independent evaluations cited in the report were conducted using pre-release access. Researchers outside Google’s Fairwind programme have not yet been able to independently test the model themselves, according to the supplied information.
These concerns do not establish that Argon’s reported results are invalid. They do mean that additional independent testing will be important for determining how reliably the model performs across real-world coding and cybersecurity tasks.
What Gemini 4 Argon Means for AI Users
Gemini 4 Argon represents Google’s push towards AI systems capable of handling complex, multi-step technical work. Its reported context window, software engineering capabilities and cybersecurity applications could make it useful for developers, enterprise teams and security researchers.
However, most consumers cannot access it directly yet. Its final real-world value will depend on wider availability, independent verification of its benchmark results, reliability across different tasks and the overall cost of running it.
For now, Gemini 4 Argon is best viewed as a promising but restricted AI model whose reported capabilities still require broader independent evaluation.
Frequently Asked Questions
1. What is Google Gemini 4 Argon?
Gemini 4 Argon is Google’s AI model designed for complex tasks such as coding, cybersecurity research, software vulnerability analysis, law and finance. Google announced it on September 30, 2026.
2. Is Gemini 4 Argon available to the public?
No public release date has been announced. Access is initially limited to selected cybersecurity experts through Google’s Fairwind Program, with paid API customers and Google AI Ultra subscribers expected to be included in a later phase.
3. What is the context window of Gemini 4 Argon?
Google reports a context window of up to one million tokens. This enables the model to process large amounts of text and other supported content, although the practical limit depends on the task and how the model handles its input and output.
4. How much does Gemini 4 Argon cost?
Introductory API pricing is $2 per million input tokens and $10 per million output tokens. The standard rates are expected to be $4 and $20 per million tokens, respectively. Google has not specified when the introductory pricing will end.
5. Is Gemini 4 Argon better than Claude Opus 5.5?
Not across every evaluation. The Artificial Analysis Intelligence Index results cited in the report give Claude Opus 5.5 a score of 58, compared with 53 for Gemini 4 Argon. The overall comparison depends on the task, cost, accuracy and other performance measures.
6. Why are experts questioning Gemini 4 Argon’s coding scores?
Epoch AI has reported flaws in the DeepSWE v1.1 benchmark, including evaluation problems in at least 23 of 113 examined tasks. Broader independent testing will help establish how well Argon performs outside the reported benchmark conditions.





