Google makes Gemini 4 AI model available to a trusted few
CSO Online ·

Google has unveiled a new frontier AI model after months of delay. Gemini 4 Argon is designed to handle complex, long-horizon workloads spanning software engineering, enterprise knowledge work such as legal and financial analysis, and cybersecurity. But only a few organizations can get their hands on it for now. Argon is “rolling out to a set of trusted cyber defenders through our Fairwind Program ,” Google wrote in a blog post announcing the release. That limitation, it said, is in order to comply with the US government’s voluntary process for granting early access to models in order to test and improve their safety guardrails before making them generally available. Google DeepMind head Koray Kavukcuoglu recently said that Gemini 4 should be released “much earlier” than the end of this year. The limited launch comes after Google delayed and ultimately skipped the release of Gemini 3.5 Pro , which the company had initially promised would arrive in June. “Google took about seven months without releasing a major top-tier model. It is three to four months behind its own Gemini roadmap as Gemini 3.5 Pro, the model Google planned for June, never shipped. The delay was linked to model-development challenges, particularly around coding and reasoning,” said Pareekh Jain , CEO of EIIRTrend and Pareekh Consulting. With Argon, Google has increased the model’s token capacity, while introductory pricing is set at $2 per million input tokens and $10 per million output tokens. More tokens, no clear lead The headline change in Argon is the model’s output limit, which has been increased from 64,000 tokens in previous Gemini models to 1 million, allowing room for more reasoning and for completing long, multi-step tasks within a single trajectory. Google said Argon is designed to handle enterprise workflows spanning coding, reasoning and multimodal tasks. The company is backing those claims with examples from its internal use. It has used Argon to identify memory optimizations across its data centers that could free up more than 300 TB of memory once deployed, with estimated total savings of 500 TB to 1 PB. On benchmarks, Argon scores 68.9% on the Vals Index, ahead of Claude Opus 5.5 at 67.0%, and 77.9% on DeepSWE v1.1, compared with 74.2% for Opus 5.5. But these results do not translate into a lead across every type of workload. In PostTrainBench for ML engineering, Claude Opus 5.5 leads at 49.3% compared to 45.3% for Argon. The company is also aiming to close the gap with Anthropic and OpenAI in cyberdefence by enabling Argon to autonomously find, validate, and patch critical software vulnerabilities. On CWE-bench v1, a benchmark for vulnerability remediation, Google reports a 68% score, tied for the top position with Grok 4.7, GPT-6 Astra and Claude Opus 5.5. While Google has published a broad set of benchmark results to support its claims, Jain cautioned that the scores should be treated as hints, not proof. “Argon looks to be good and beating rivals at using automated tools to complete multi-step tasks without getting confused, and it sticks closely to real facts. Its everyday coding abilities are basically average and tied with others. It still trails others when it comes to creative writing, nuanced explanations, and running command-line computer terminals,” he said. Jain noted Argon shows that Google has regained much of the capability gap, but it has lost some momentum and developer mindshare. It is now competitive at the frontier, but not clearly ahead across all areas. Price, performance and switching Argon’s capabilities come with an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced 95% lower. Those rates will rise to $4 per million input tokens and $20 per million output tokens after the introductory period, although Google has not specified when the new rates will take effect. In comparison, Anthropic’s Claude Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, while OpenAI’s GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens. “This makes the introductory pricing very competitive. The eventual pricing is roughly comparable to other premium frontier models. Enterprises should therefore build business cases using the $4/$20 pricing, rather than assuming the introductory price will continue,” said Jain. While Argon is worth evaluating for enterprises already using competing models, it is not necessarily enough to justify switching outright. CIOs would also need to weigh the cost of migrating existing applications and developer workflows and integrations. Jain noted CIOs should evaluate models on their own workloads, focusing on task success rate, accuracy and hallucinations, agent reliability, cost per successful task, latency, security and governance, data privacy, integration with existing systems, and vendor lock-in. The key metric is not cost per token, but the cost per successful business outcome. Argon gives CIOs a strong reason to evaluate it but not switch to it automatically. The smart move, he said, is for enterprises to add Argon for the jobs it does best, like legal, finance, and long documents, and keep their current AI for things like coding. If a company already runs all its tech on Google Cloud, switching can be considered if Argon clearly wins in the tests and the costs still make sense at the full price, said Jain. This article first appeared on InfoWorld .
Google has unveiled a new frontier AI model after months of delay. Gemini 4 Argon is designed to handle complex, long-horizon workloads spanning software engineering, enterprise knowledge work such as legal and financial analysis, and cybersecurity. But only a few organizations can get their hands on it for now. Argon is “rolling out to a set of trusted cyber defenders through our Fairwind Program ,” Google wrote in a blog post announcing the release. That limitation, it said, is in order to comply with the US government’s voluntary process for granting early access to models in order to test and improve their safety guardrails before making them generally available. Google DeepMind head Koray Kavukcuoglu recently said that Gemini 4 should be released “much earlier” than the end of this year. The limited launch comes after Google delayed and ultimately skipped the release of Gemini 3.5 Pro , which the company had initially promised would arrive in June. “Google took about seven months without releasing a major top-tier model. It is three to four months behind its own Gemini roadmap as Gemini 3.5 Pro, the model Google planned for June, never shipped. The delay was linked to model-development challenges, particularly around coding and reasoning,” said Pareekh Jain , CEO of EIIRTrend and Pareekh Consulting. With Argon, Google has increased the model’s token capacity, while introductory pricing is set at $2 per million input tokens and $10 per million output tokens. More tokens, no clear lead The headline change in Argon is the model’s output limit, which has been increased from 64,000 tokens in previous Gemini models to 1 million, allowing room for more reasoning and for completing long, multi-step tasks within a single trajectory. Google said Argon is designed to handle enterprise workflows spanning coding, reasoning and multimodal tasks. The company is backing those claims with examples from its internal use. It has used Argon to identify memory optimizations across its data centers that could free up more than 300 TB of memory once deployed, with estimated total savings of 500 TB to 1 PB. On benchmarks, Argon scores 68.9% on the Vals Index, ahead of Claude Opus 5.5 at 67.0%, and 77.9% on DeepSWE v1.1, compared with 74.2% for Opus 5.5. But these results do not translate into a lead across every type of workload. In PostTrainBench for ML engineering, Claude Opus 5.5 leads at 49.3% compared to 45.3% for Argon. The company is also aiming to close the gap with Anthropic and OpenAI in cyberdefence by enabling Argon to autonomously find, validate, and patch critical software vulnerabilities. On CWE-bench v1, a benchmark for vulnerability remediation, Google reports a 68% score, tied for the top position with Grok 4.7, GPT-6 Astra and Claude Opus 5.5. While Google has published a broad set of benchmark results to support its claims, Jain cautioned that the scores should be treated as hints, not proof. “Argon looks to be good and beating rivals at using automated tools to complete multi-step tasks without getting confused, and it sticks closely to real facts. Its everyday coding abilities are basically average and tied with others. It still trails others when it comes to creative writing, nuanced explanations, and running command-line computer terminals,” he said. Jain noted Argon shows that Google has regained much of the capability gap, but it has lost some momentum and developer mindshare. It is now competitive at the frontier, but not clearly ahead across all areas. Price, performance and switching Argon’s capabilities come with an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced 95% lower. Those rates will rise to $4 per million input tokens and $20 per million output tokens after the introductory period, although Google has not specified when the new rates will take effect. In comparison, Anthropic’s Claude Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, while OpenAI’s GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens. “This makes the introductory pricing very competitive. The eventual pricing is roughly comparable to other premium frontier models. Enterprises should therefore build business cases using the $4/$20 pricing, rather than assuming the introductory price will continue,” said Jain. While Argon is worth evaluating for enterprises already using competing models, it is not necessarily enough to justify switching outright. CIOs would also need to weigh the cost of migrating existing applications and developer workflows and integrations. Jain noted CIOs should evaluate models on their own workloads, focusing on task success rate, accuracy and hallucinations, agent reliability, cost per successful task, latency, security and governance, data privacy, integration with existing systems, and vendor lock-in. The key metric is not cost per token, but the cost per successful business outcome. Argon gives CIOs a strong reason to evaluate it but not switch to it automatically. The smart move, he said, is for enterprises to add Argon for the jobs it does best, like legal, finance, and long documents, and keep their current AI for things like coding. If a company already runs all its tech on Google Cloud, switching can be considered if Argon clearly wins in the tests and the costs still make sense at the full price, said Jain. This article first appeared on InfoWorld .