Javlon Baxtiyorov
← Writing

GPT-5.6 Sol: The Agent-First Flagship (Full Benchmarks)

OpenAI's GPT-5.6 Sol is GA. #1 on the Coding Agent Index, 91.9 on Terminal-Bench, one point behind Fable 5 on intelligence — at a price that bites. Every number, and the catch.

GPT-5.6 Sol: The Agent-First Flagship (Full Benchmarks)
Image: FelloAI

On 9 July 2026 OpenAI moved GPT-5.6 to general availability and, with it, its flagship agentic mode: Sol. The GPT-5.6 family ships in three tiers — Sol (the heavy agentic model), Terra, and Luna — and Sol is the one every coding-agent team spent the week benchmarking. I pulled the numbers from the trackers and the model card so you can see exactly where it lands.

Solagentic flagship tier
1Mtoken context
$5 / $30in / out per 1M
#1AA Coding Agent Index

What Sol actually is

GPT-5.6 is OpenAI's mid-cycle bump on the GPT-5 line, and Sol is its top reasoning-and-tools tier — built for "difficult professional, coding, research, computer-use, and tool-heavy work." It carries a 1M-token context window, and it's the engine behind ChatGPT Work, OpenAI's new agent-based product that pairs Sol with Codex to actually do multi-step jobs rather than just answer.

It landed in limited preview on 26 June and went GA across ChatGPT, Codex, and the API on 9 July 2026. There's also an Ultra setting that trades latency and cost for a few more points on the hardest benchmarks.

The benchmarks

Sol's whole pitch is agentic and tool-use performance, and that's where the headline numbers are. These are the ones doing the talking:

GPQA Diamond94.6
BrowseComp92.2
Terminal-Bench 2.091.9
FrontierMath v2 (T1–3)89.0
MMMU-Pro (w/ Python)84.6
CyberGym84.5

A 94.6 on GPQA Diamond (PhD-level science) and a 91.9 on Terminal-Bench 2.0 — can the model drive a shell to finish real tasks — put Sol squarely at the frontier. On Artificial Analysis's Coding Agent Index it takes the #1 spot at 80 points, and on their overall Intelligence Index it sits at 59 — one single point behind Claude Fable 5, at roughly a third of Fable's price.

Agentic

Benchmark GPT-5.6 Sol
BrowseComp 92.2
Terminal-Bench 2.0 91.9
CyberGym 84.5
OSWorld 2.0 62.6
Toolathlon 58.0
ExploitGym 33.7

Coding

Benchmark GPT-5.6 Sol
DeepSWE v1.1 72.7
CursorBench 3.2 67.2
SWE-bench Pro 64.6
Terminal-Bench 2.1 (Ultra) 91.9
Terminal-Bench 2.1 (standard) 88.8

Knowledge, math & multimodal

Benchmark GPT-5.6 Sol
GPQA Diamond 94.6
FrontierMath v2 (Tiers 1–3) 89.0
FrontierMath v2 (Tier 4) 83.0
MMMU-Pro (with Python) 84.6
MMMU-Pro 83.0
HealthBench Professional 60.5

Where it sits in the market

Sol's positioning is unusually clean: near-Fable-5 intelligence, best-in-class coding-agent scores, at $5 in / $30 out per 1M tokens. That output price is the one wrinkle — at $30/M out it's more expensive on output than Claude Opus 4.8 ($25) and far above the new wave of cheap frontier models. OpenAI is betting that if Sol finishes agentic tasks in fewer, better steps, the higher per-token price still wins on total job cost. For tool-heavy pipelines that's a real argument; for chatty, high-output workloads it's a harder sell.

The honest caveats

  • Terminal-Bench "91.9" is the Ultra setting. Standard GPT-5.6 lands at 88.8 on Terminal-Bench 2.1. Both are excellent — just make sure you're comparing the tier you'll actually run.
  • Output is pricey. $30/M out is the highest of the mainstream frontier tier right now. Budget by task completion cost, not sticker price.
  • Benchmarks aren't your pipeline. A #1 Coding Agent Index is a strong signal, not a guarantee your agents will behave on your codebase. Run it against your own evals.

My take

Sol is the most convincing "agent-first" flagship yet — the Terminal-Bench and Coding Agent Index numbers are exactly what you want if you're building autonomous coding or computer-use pipelines. But 2026's real story isn't any one model; it's that the frontier got crowded in a single fortnight — Grok 4.5 on the 8th, Sol on the 9th, Kimi K3 on the 16th. I keep everything behind one interface and let models compete for my traffic on merit. Right now Sol earns the agentic slots; the cheaper models earn the high-volume ones. Wire for portability and you get to say that instead of being locked into whoever's on top this week.


Sources

Numbers are as reported around GA (9 July 2026) and will move as independent boards settle. Spot a corrected figure? Tell me and I'll update.

Read next All writing →
← All writing Get in touch →