Gemini 3.6 Flash: Google's Speed Play (Full Benchmarks)
Google's Gemini 3.6 Flash is the fastest model on the board (#1 of 187, 303 tok/s) at $1.50/$7.50 and 50 on the Intelligence Index — a workhorse, not a crown-taker, while its flagship stays late. Every number.
On 21 July 2026, while everyone was still arguing about who owns the top of the leaderboard, Google shipped Gemini 3.6 Flash — and quietly took a number nobody else has: the fastest frontier-class model on the board. It isn't the smartest model in the world. It was never trying to be. It's the best workhorse released this year, and it lands at a moment when Google's actual flagship — a "3.5 Pro" — is conspicuously, awkwardly late. Here's every number that matters, and the honest read on where it fits.
What Google actually shipped
Gemini 3.6 Flash is the newest entry in Google's fast tier — the line built for high-volume agentic work, coding loops, and long-context jobs where you're paying per million tokens and latency compounds. The headline spec is throughput: on Artificial Analysis's independent testing it streams at 303.6 tokens per second — ranked #1 of 187 models. Nothing else at the frontier is faster.
It pairs that with a 1M-token context window, $1.50 / $7.50 input/output pricing, and — the number most people miss — a cached-input price of $0.15 per 1M, a 90% discount. For anything with a stable system prompt or a re-read corpus, that cache line is the whole economic story (I break the math down in a companion post).
What it is not is a leaderboard-topper. On the Artificial Analysis Intelligence Index it scores 50, good for #21 of 187 overall — roughly ten points behind the flagship crown. Read that correctly: a Flash-class model is now within ten points of the smartest models on earth, at a fraction of their price and twice their speed. That's the story.
The benchmarks
Google leaned into agentic and long-context evals, and the numbers back the pitch:
OSWorld-Verified at 83.0 and Terminal-Bench 2.1 at 78.0 are the ones to weight if you build agents — that's computer-use and drive-a-real-shell-to-done, and a fast model that scores here is a fast model you can actually loop. GDM-MRCR v2 at 54.0 across the full 1M window is the long-context number: multi-round coreference resolution at a million tokens, which is where most "big context" models quietly fall apart.
The generational jump is the real news
Benchmarks in isolation flatter everyone. The honest way to read a .6 release is against the .5 it replaces — and against the previous Pro, because that's the model you were paying more for last month:
| Benchmark | 3.6 Flash | 3.5 Flash | 3.1 Pro (old flagship) |
|---|---|---|---|
| SWE-Bench Pro | 58.7 | 55.1 | 54.2 |
| Terminal-Bench 2.1 | 78.0 | 76.2 | 73.8 |
| MLE-Bench | 63.9 | 49.7 | 42.6 |
| DeepSWE v1.1 | 49.0 | 37.0 | — |
| GDPVal-AA v2 (Elo) | 1421 | 1349 | 965 |
| GDM-MRCR v2 @ 1M | 54.0 | <27 | — |
Two lines jump off that table. MLE-Bench went from 49.7 to 63.9 — a fourteen-point machine-learning-engineering gain in one generation. And long-context MRCR more than doubled, from under 27 to 54. But the quiet humiliation is the last column: the new Flash now beats the old Pro on every shared row. Whatever "Gemini 3.5 Pro" turns out to be, it has to clear a bar its own cheap sibling just raised.
Where it sits in the field
Here's the part Google's blog won't tell you. Six labs now field a model scoring above 50 on the Intelligence Index — Anthropic, OpenAI, Moonshot, xAI, Z AI, and Meta. Gemini 3.6 Flash sits at exactly 50, just outside that club:
| Model | Lab | AA Intelligence Index |
|---|---|---|
| Claude Fable 5 | Anthropic | 60 |
| GPT-5.6 Sol | OpenAI | 59 |
| Kimi K3 | Moonshot | 57 |
| Claude Opus 4.8 | Anthropic | 56 |
| Grok 4.5 | xAI | 54 |
| GPT-5.6 Luna | OpenAI | 51 |
| Gemini 3.6 Flash | 50 |
The comparison that actually matters isn't Flash-vs-flagships — it's Flash-vs-the other cheap-and-fast models, and there Gemini is genuinely competitive. I put it head to head with GPT-5.6 Luna, Grok 4.5, Kimi K3, and Claude in a dedicated comparison.
The elephant: where is Gemini 3.5 Pro?
Google's flagship reasoning model was promised for June, slipped to July, and as of this writing has no model card, no official benchmarks, and no confirmed pricing — only reporting of a ~2M-token context and premium pricing in the ballpark of $15/$60 per 1M. Treat all of that as rumor until Google publishes.
The strategic read: while OpenAI shipped three flagship tiers (Sol, Terra, Luna) and Anthropic fields two models in the top four, Google's answer to the July frontier wave was a Flash model. A very good Flash model — but a mid-tier one. The crown fight is happening without them, and every week the Pro stays unshipped, that's a louder statement than any benchmark.
The honest caveats
- 50 is not 60. For the hardest reasoning, math-proof, and research-grade work, the flagship tier is a real step up. Flash is the model you route the other 80% of traffic to — not the 20% that needs the smartest brain in the room.
- Fast throughput ≠ fast first token. That 303 tok/s is sustained streaming speed; time-to-first-token measured ~11.5s at the high-reasoning setting, because the model thinks before it streams. For chat UIs that's fine; for hard-latency SLAs, measure it yourself.
- Some evals are Google-reported. OSWorld, Terminal-Bench, and SWE-Bench Pro are broadly corroborated; a few figures are first-party or normalized. The generational deltas are more trustworthy than any single absolute score.
My take
Gemini 3.6 Flash is the most useful model Google has shipped in a year — precisely because it stopped chasing the crown. It's the fastest thing on the board, it's cheap, it holds a real million-token context, and it's a genuinely strong agentic model. If you run high-volume coding or computer-use agents, it belongs in your router today.
But "our best new model is a Flash" is a strange sentence for the company that once defined the frontier. The product is excellent. The absence behind it — a flagship that keeps not arriving — is the thing to actually watch.
Sources
- Artificial Analysis — Gemini 3.6 Flash and Four frontier launches in eight days
- OfficeChai — Gemini 3.6 Flash benchmarks
- BenchLM — Gemini 3 Flash · AA Intelligence Index leaderboard
- eesel — Google Gemini 3 pricing 2026 · AIToolsReview — Gemini 3.5 Pro: what's confirmed
Numbers reflect independent trackers and Google's launch figures as of 22 July 2026 and will move as boards settle. Spot a corrected figure? Tell me and I'll update.


