Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

Google’s overdue Gemini 4 Argon reaches the frontier with mixed results for science

By Brian Buntz | October 1, 2026

After seemingly fading in the LLM model race, Google again has a frontier model with Gemini 4 Argon, and it could do more to put pricing pressure on the frontier than redefine the frontier benchmarks. As Google touts leads in biology research and several coding tests, independent evaluations suggest more of a competitive entrant. Argon matches GPT-6 Astra on Artificial Analysis’s overall index, behind Claude Opus 5.5, and ranks eighth on Arena’s Agent Arena leaderboard.

Google is rolling Argon out first to cybersecurity defenders in Fairwind, a program it opened Sept. 2 alongside 3.8 Flash Cyber for governments, national cyber authorities, critical infrastructure operators and core technology platforms. Google says the paid Gemini API and Google AI Ultra come next. It has not yet provided a date.

The continued rapid pace of frontier model development appears to have caught Google by surprise. At its I/O developer conference in May, Google promised Gemini 3.5 Pro within a month. But it was delayed, and then delayed some more, and was apparently scrapped for a new model. Bloomberg reported in July that 3.5 Pro was months behind schedule, in part because its coding performance fell short of internal goals.

The months in between belonged to Google’s less expensive Flash tier. Gemini 3.5 Flash arrived at I/O, followed by 3.6 Flash on July 21, 3.7 Flash on Aug. 13 and 3.8 Flash on Sept. 2.

While Gemini 2.5 Pro helped reshape the AI race in March 2025, Google has broadly trailed in AI coding since then. Menlo Ventures, an Anthropic investor, estimated last December that Anthropic held 54% of enterprise spending on coding models, versus 21% for OpenAI and 11% for Google. While OpenAI has made gains in coding traction since, Gemini largely hasn’t. In September, Google opened Anthropic’s Claude Opus 5 to its engineers through its Antigravity coding environment, while saying Gemini remains its primary model for internal development. Some Google employees questioned Argon’s performance on practical coding tasks, while others believed it had caught up with leading rivals, according to Bloomberg.

Google disputed the characterization. The company claims that Gemini 4 Argon leads the frontier on a majority of benchmarks. Independent testing shows competitive, if not chart-topping, performance.

How Gemini 4 Argon compares

Results from Google, Artificial Analysis and Arena as of Oct. 1, 2026. Best result in each row.

Category Benchmark Gemini 4 Argon GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5
Google’s reported results (methodology)
Knowledge work Vals Index 68.9% 63.1% 65.8% 67.0%
AutomationBenchScore 51.3% 41.4% 31.4% 42.5%
Vals Finance Agent v2 65.4% 53.5% 58.9% 58.6%
Harvey’s Legal Agent Benchmark 19.6% 5.4% 6.7% 3.8%
Agentic coding DeepSWE v1.1 77.9% 74.1% 67.4% 74.2%
FrontierSWE v2 55.0% 65.5% 56.3% 62.3%
Vibe Code Bench 91.9% 89.6% 90.3% 90.3%
Terminal-Bench 4.0 57.4% 58.2% 57.9% 66.4%
ML engineering PostTrainBench 45.3% 44.3% 40.2% 49.3%
Science and math Terminal-Bench Science 0.1 57.6% 68.1% 52.6% 63.3%
LABBench2 88.8% 85.4% 68.6% 73.1%
RiemannBench 76.0% 72.0% 65.6% 69.6%
Long context GraphWalksUp to 128K, BFS (F1) 99.7% 98.7% 91.4% 90.6%
GraphWalks256K to 1M, BFS (F1) 84.2% 71.8% 65.0% 66.8%
Computer use Agent’s Last ExamPass rate 39.5% 34.2% — 38.2%
OSWorld-2.0Offline subset, partial score 69.2% 72.6% — —
Multimodal Chartography 71.6% 71.0% 46.2% 66.3%
LVBench 91.7% 87.5% 79.7% 83.7%
Cybersecurity CWE-bench v1 68.0% 68.0% 58.0% 67.0%
Independent: Artificial Analysis (analysis)
Overall Intelligence IndexScore; Argon at highest reasoning setting 53 53 53 58
Cost per index taskArgon at introductory price; $3.98 at standard price $1.99 $3.26 — —
Output tokens per index taskLower is more efficient ~62,000 ~27,000 — —
Independent: Arena (Agent Arena leaderboard)
Agents Agent Arena overall rankArgon: 3,417 sessions; rank range 3rd–16th 8th — — —
Google’s figures come from its own testing. For Terminal-Bench Science, Google ran Argon itself and took competitor scores from the benchmark’s public leaderboard. On LABBench2, all four models had terminal access, scientific software and the internet.

A dash means no score was reported. Artificial Analysis and Arena figures are live and may change.

What independent tests show

Artificial Analysis provides independent support for Google’s progress. Argon, using its highest reasoning setting, scores 53 on the firm’s Intelligence Index, matching the reported scores of Astra and Claude Fable 5.1. That is 23 points above Gemini 3.1 Pro Preview, the last Pro model Google shipped. The current index chart places Claude Opus 5.5, tested at maximum reasoning with fallback, at 58.

Arena’s Agent Arena leaderboard offers a second view, focused on tool-using performance. Argon ranked eighth overall when checked Oct. 1, based on 3,417 sessions. Its reported rank range stretched from third to 16th.

The science results split

In Google’s comparison, Argon scores 88.8% on LABBench2, ahead of Astra at 85.4% and Opus 5.5 at 73.1%. On Terminal-Bench Science 0.1, Argon’s 57.6% trails Astra’s 68.1% and Opus 5.5’s 63.3%. Argon beats Claude Fable 5.1 on both.

LABBench2, the benchmark from Edison Scientific, tests tasks such as finding evidence in scientific literature, working with biological sequences and troubleshooting experimental protocols. Its scores speak to capabilities a biologist might use while preparing or interpreting an experiment. 

Terminal-Bench Science asks agents to complete 70 bounded research workflows through a computer terminal, producing outputs that can be checked. It probes whether a system can carry a computational assignment through to a verifiable result. 

Google’s methodology says it ran all four models on LABBench2 with terminal access, scientific software and the internet. For Terminal-Bench Science, Google tested Argon itself and took the competing scores from the benchmark’s leaderboard. 

Frontier performance for a fraction of the token price

Google set introductory API prices of $2 per million input tokens and $10 per million output tokens, half its standard rates. Argon can also generate up to 1 million output tokens in a response, up from a 64,000-token ceiling on earlier Gemini releases, which leaves room for the long, multi-step runs that research agents require.

Artificial Analysis found that Argon consumed about 62,000 output tokens per index task, versus Astra’s 27,000. Its lower introductory prices nevertheless brought the measured cost to $1.99 per task, against Astra’s $3.26. At standard prices, Argon’s cost would rise to $3.98. Google’s 50% promotion therefore changes which model is cheaper on this particular workload; its end date has not been confirmed.

Arena also placed Argon on its cost-performance frontier, identifying a competitive trade-off at launch prices.

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

OpenAI and Synopsys plan a specialized model for chip design
OpenAI reportedly seeks $30 billion after delaying IPO
Hansa Biopharma taps Cradle’s AI to design next-generation autoimmune therapeutics
R&D World parent company Arrowfly launches AI for Engineers, definitive platform for engineers navigating AI
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE