Token & cost optimization
Coding Agent Index Shows Tenfold Price Gap for Comparable Benchmark Scores
Artificial Analysis published its Coding Agent Index evaluating model and harness combinations. Claude Sonnet 5.5 in Claude Code leads with 68 at $14.19 per task, while GPT-6.1 Sol in Codex achieves 63 for $1.04.
October 2, 2026 6 min read
Curated by Oleksandr Kuzmenko, AI Product EngineerUpdated October 2, 2026Sources cited on every story
AI-assisted · editor-reviewedHow we use AI

Impact: High
Why it matters
Developers can lower agent orchestration bills by up to 93% by routing routine implementation tasks to Codex instead of Claude Code.
TL;DR
- 01Claude Sonnet 5.5 in Claude Code leads the index at 68 but averages $14.19 per completed task.
- 02GPT-6.1 Sol in Codex scores 63 while costing $1.04 per task, offering 13x higher cost efficiency.
- 03Gemini 4 Argon in Antigravity CLI scores 64 at $5.84 per task under promotional pricing.
Key facts
- Claude Sonnet 5.5 (Claude Code) score
- 68
- Claude Sonnet 5.5 (Claude Code) cost per task
- $14.19
- Gemini 4 Argon (Antigravity CLI) score
- 64
- Gemini 4 Argon (Antigravity CLI) cost per task
- $5.84
- GPT-6.1 Sol (Codex) score
- 63
- GPT-6.1 Sol (Codex) cost per task
- $1.04
Benchmark Evaluation and Score Distribution The Artificial Analysis Coding Agent Index evaluates autonomous coding agents across three distinct coding benchmarks, assessing how model capabilities translate when paired with specific agent harnesses. In the latest benchmark sweep, Claude Sonnet 5.5 running in Claude Code captured the top position with an overall score of 68. Gemini 4 Argon running in Antigravity CLI followed with a score of 64, while GPT-6.1 Sol configured with extra high effort in Codex achieved a score of 63. ### Severe Cost Discrepancy Across Agent Harnesses Although benchmark performance clustered within a narrow five-point spread across all three systems, execution costs diverged significantly. Claude Sonnet 5.5 in Claude Code registered the highest measured expense at $14.19 per completed task. Gemini 4 Argon in Antigravity CLI averaged $5.84 per task, representing less than half the cost of Sonnet 5.5 (utilizing Google promotional rates). GPT-6.1 Sol in Codex recorded $1.04 per task, operating at roughly one-sixth the cost of Argon and one-fourteenth the cost of Sonnet 5.5. ### Practical Routing Recommendations for Teams For daily development and continuous agent execution, running Claude Code with maximum reasoning on repetitive chores creates compounding token invoices. Tech leads can implement multi-tier harness policies: dispatch high-volume bug fixing, refactoring, and test writing to GPT-6.1 Sol in Codex at $1.04 per task, and escalate failed trajectories or multi-file architectural overhauls to Claude Sonnet 5.5.
Try it in 2 minutes
codex exec --model gpt-6.1-sol --effort xhigh --task 'Fix failing unit tests in auth module'bash
✓ When to use
- When choosing between Claude Code, Codex, and Antigravity CLI for automated workflows.
- When configuring multi-agent harnesses where token budget dictates execution frequency.
✕ When NOT to use
- When a project requires single-token deterministic zero-shot transformations without tools.
- When enterprise security restrictions prohibit external hosted execution harnesses.
What to do today
- Audit background agent spend across Claude Code and switch routine CI tasks to Codex.
- Benchmark repository test-fixing tasks against GPT-6.1 Sol before standardizing on Claude Sonnet 5.5.
- Track promotional pricing expirations when benchmarking against Gemini 4 Argon.
Sources