Models & research
Gemini 4 Argon Delivers 1M Output Tokens and Deep Code Migrations
Google announced Gemini 4 Argon, a frontier reasoning model featuring an unprecedented 1 million output token limit. Designed for long-horizon engineering and autonomous refactoring, it introduces 95% prompt caching discounts and sets new state-of-the-art benchmarks in SWE tasks.
October 1, 2026 5 min read
Curated by Oleksandr Kuzmenko, AI Product EngineerUpdated October 1, 2026Sources cited on every story
AI-assisted · editor-reviewedHow we use AI

Impact: High
Why it matters
Plan for long-running refactoring agents capable of emitting complete codebases without splitting outputs or losing multi-step reasoning context.
TL;DR
- 01Argon introduces an industry-leading 1M token output window, opening true single-pass whole-module refactoring.
- 02Prompt caching offers 95% savings ($0.10/M tokens), making massive static code contexts financially viable.
- 03Google reported 77.9% on DeepSWE v1.1 and demonstrated a 2.7x speedup rewriting libgav1 C++ SIMD into Rust.
Key facts
- Input Price per 1M Tokens
- $2.00
- Output Price per 1M Tokens
- $10.00
- Cached Input Discount
- 95% ($0.10 per 1M tokens)
- Output Token Limit
- 1,000,000 tokens
- DeepSWE v1.1 Score
- 77.9% (self-reported)
- CWE-bench v1 Score
- 68% (self-reported)
Million-Token Generation Engine Google expanded the single-turn generation envelope from 64K to 1M output tokens. This enables deep reasoning agents to run hundreds of thousands of thinking and generation steps without truncation, eliminating brittle agent harness loops designed to stitch fragmented responses together. ### Production Infrastructure Refactoring In internal trials, Argon tackled large-scale codebase migrations, converting C/C++ repositories up to 800K lines (such as Fuchsia's Zircon kernel) into Rust. For video decoder libgav1, Argon agents replaced 32K lines of SIMD assembly with safe, auto-vectorizing Rust that benchmarks 2.7x faster than prior manual ports. Agents also tuned datacenter memory profiles, reclaiming over 300 TiB of RAM across Google's fleet. ### Pricing and Benchmark Metrics Argon introduces an introductory rate of $2.00 per million input tokens and $10.00 per million output tokens. Prompt caching reduces input costs by 95% down to $0.10 per million tokens. The model achieved 77.9% on DeepSWE v1.1, 51.3% on Zapier's AutomationBench, and 68% on security patch suite CWE-bench v1.
✓ When to use
- Autonomous multi-file repository migrations where output token limits in standard models cause context fragmentation.
- Security audit agents that require deep reasoning loops to find, reproduce, and verify vulnerability patches.
- Complex quantitative or algorithmic code restructuring requiring profile-guided experiment cycles.
What to do today
- Audit large legacy C/C++ modules suitable for automated memory-safe Rust translation once developer API access opens.
- Structure agent architectures to leverage prompt caching headers to benefit from the 95% input token discount.
What the community says
“More Cartmanland marketing. It's the best park ever, and you can't come! Get lost.”
“Likely just trying to appear relevant in the news cycle. It's also kinda wild how the competition being at v6.1 makes 3.x feel ancient, at least saying you are at v4 now changes public perception a bit imo.”
Sources