In short
- Google unveiled Gemini 4 Argon on Wednesday, scoring 77.9% on DeepSWE v1.1 and main 12 of 18 benchmarks in its personal comparability desk.
- It posted a 0.7% assault success price on Grey Swan’s immediate injection take a look at, forward of Claude Opus 5.5 and Claude Fable 5.1, which each scored 1.0%.
- Argon goes first to vetted cyber defenders by means of the Fairwind Program, with out cyber guardrails, earlier than reaching paid API prospects and Google AI Extremely subscribers.
Gemini 4 is lastly right here, one week after the discharge of Claude Opus 5.5 and in the future after GPT 6.1 Sol, proving American labs are very a lot dedicated to slowing down AI improvement. Please excuse our sarcasm.
Google unveiled Gemini 4 Argon on Wednesday, calling it its frontier mannequin, which means its most succesful, for coding, workplace work and cyber protection.
On DeepSWE v1.1, a take a look at of whether or not an AI can end lengthy, messy, real-world software program engineering jobs, scored as a proportion, Argon hit 77.9%. Claude Opus 5.5 bought 74.2%, GPT-6 Astra 74.1% and Claude Fable 5.1 67.4%.
For scale, Gemini 3.6 Flash managed 49% on the identical take a look at in July. Argon may write as much as 1 million tokens in a single reply, up from 64,000. A token is a piece of textual content, roughly three-quarters of a phrase, so that’s about 750,000 phrases versus about 48,000.
Take these numbers with a grain of salt, although. Google computed its personal DeepSWE rating, whereas rivals’ numbers got here from a public leaderboard and firm reviews. Its desk additionally concedes floor. Argon leads on 12 of 18 benchmarks, ties one and trails on 5, a mixture of coding, science and computer-control assessments.
However the mannequin’s flashy characteristic is cyber capabilities. Conceal a secret instruction inside an e mail, watch for an AI assistant to learn it, and see if the AI obeys the stranger as a substitute of you. That’s oblique immediate injection, and it’s the nightmare for anybody who needs handy an AI their inbox or purchasing cart.
On Grey Swan’s Oblique Immediate Injection benchmark, which hides malicious directions in content material brokers learn and scores how usually the assaults work inside 15 tries, Argon landed at 0.7%. Decrease is best. Claude Opus 5.5 and Claude Fable 5.1 each scored 1.0%.
GPT-6 Astra got here in at 8.5%. Grok 4.6 and Kimi K3 bought tricked simply over half the time, at 51.8% and 52.7%.
Argon goes to vetted safety groups by means of the Fairwind Program, Google’s limited-access cyber protection initiative, which launched September 2 with greater than 650 companions together with governments and significant infrastructure operators. And it ships “with out cyber guardrails,” the built-in refusals that usually cease a mannequin from serving to with hacking.
BitcoinBTC · USD
$84,148+0.04%
Sep 24Sep 26Sep 27Sep 29Oct 1
$85.3k$84.4k$83.5k$82.7k
24h ExcessiveExcessive$85,518
24h LowLow$82,951
VolVol$1.6B
Market projectionsOdds by Myriad
The logic behind such a transfer is that defenders want a mannequin that may assume like an attacker to patch holes earlier than criminals discover them. The catch is that the identical ability cuts each methods, so Google says a phased rollout is the one secure path. It is usually participating within the U.S. authorities’s voluntary course of for pre-release mannequin entry.
Google is not the primary to place a cyber mannequin behind a velvet rope. An early model of Anthropic’s Claude Mythos helped find 271 vulnerabilities in Firefox, which means 271 safety holes Mozilla then patched. OpenAI has taken the same route with its Trusted Access for Cyber program.
Argon’s cyber scores soar over Gemini 3.8 Flash Cyber, the restricted mannequin Google launched with Fairwind. On the Wiz Penetration Check Benchmark, an inner Google take a look at that asks an AI to write down working exploits in opposition to actual web-application flaws with out seeing the code, and scores the share solved on the primary strive, Argon hit 70.9% in opposition to 58.2%.
Google additionally says Argon helped safety agency Wiz discover a crucial flaw in healthcare software program utilized by hospitals worldwide, one earlier frontier fashions had missed.
The launch follows a tough summer season for Google. In July it shipped smaller Flash fashions however skipped the promised Gemini 3.5 Professional, and Alphabet shares fell about 4.4%. Argon additionally landed the identical day President Trump unveiled a voluntary, penalty-free AI accord that Google’s management signed.
Google says wider launch comes as quickly as potential, beginning with paid API prospects and Google AI Extremely subscribers.
Introductory pricing is $2 per million enter tokens and $10 per million output tokens. Google hasn’t stated when that interval ends, solely that normal charges are $4 and $20.
Day by day Debrief Publication
Begin on daily basis with the highest information tales proper now, plus authentic options, a podcast, movies and extra.
You might also like
More from Web3
Minecraft, Candy Crush Among 11 Games in EU Virtual Currency Crackdown
Briefly EU client authorities have opened coordinated actions towards 9 video games corporations over how they promote in-game currencies. The rules …
Project Tapestry Gains Momentum for Sovereign AI During UN General Assembly Week
First Technical Milestones Achieved Alongside Expanded Collaborations with Vietnam and IndiaNEW YORK, Oct. 1, 2026 /PRNewswire/ — The AI …
Dogecoin Is Getting Apps as DogeOS Opens Its Public Testnet
In short DogeOS opened the general public testnet of its EVM-compatible software layer for Dogecoin on Sept. 30. Charges on the …





