Grok 4.5: The Value Frontier ($2/$6, Full Benchmarks)
xAI's Grok 4.5 lands 4th on the intelligence index at one-eighth of Fable 5's output price, sitting right inside Cursor. Frontier-adjacent quality, cheap. The full breakdown.
xAI shipped Grok 4.5 on 8 July 2026, and it did something the rest of the frontier hasn't: it undercut everyone on price without falling off the leaderboard. A mixture-of-experts model trained on tens of thousands of NVIDIA GB300 GPUs — and, notably, on trillions of Cursor tokens — it launched straight into Cursor on every plan. Here are the numbers, and where the trade-offs actually are.
What xAI shipped
Grok 4.5 is SpaceXAI/xAI's smartest model to date — a mixture-of-experts system aimed squarely at coding, agentic work, and knowledge tasks. The headline isn't a single benchmark; it's the slope: frontier-adjacent quality at $2 per 1M input / $6 per 1M output, the cheapest entry in the current top 10. It ships in Grok Build, the SpaceXAI console, and — the distribution masterstroke — in Cursor on all plans, where a big chunk of its training data came from in the first place.
The one spec to keep in mind: the context window is 500K tokens, half of the 1M that Sol, Kimi K3, and the Claude tier offer. For most coding sessions that's plenty; for whole-repo or long-document work it's a real ceiling.
The benchmarks
Grok 4.5 leads with coding and agentic scores, and it holds its own on knowledge:
A 93.1 on GPQA Diamond is within a whisker of Sol (94.6) and Kimi K3 (93.5) — for a model priced at $2/M input, that's the number that reorders spreadsheets. On Artificial Analysis's Intelligence Index it scores ~53.8 and lands 4th overall, behind only Fable 5, GPT-5.6 Sol, and Opus 4.8.
Coding
| Benchmark | Grok 4.5 |
|---|---|
| SWE Multilingual | 78.0 |
| AA Coding Index | 72.5 |
| CursorBench 3.2 | 66.7 |
| SWE-bench Pro | 64.7 |
| AA-SciCode | 54.1 |
| DeepSWE | 53.0 |
Agentic
| Benchmark | Grok 4.5 |
|---|---|
| Terminal-Bench 2.0 | 83.3 |
| GDPval-AA (normalized) | 51.7 |
| AA AutomationBench | 51.4 |
| AA Agentic Index | 45.7 |
| AA EnterpriseOps-Gym | 40.8 |
| AA Tau3 Banking | 32.6 |
Knowledge & multimodal
| Benchmark | Grok 4.5 |
|---|---|
| GPQA Diamond | 93.1 |
| MMMU-Pro | 80.4 |
| AA Long Context Reasoning | 67.7 |
| AA Intelligence Index | 53.8 |
| Humanity's Last Exam | 40.3 |
| Design Arena Website Elo | 1328 |
Why it matters
The story of Grok 4.5 is cost per unit of quality. At $2/$6 it's roughly a third of Opus 4.8's price and a fraction of Fable 5's, while trailing the leaders by only a handful of index points. Bundle that with default availability inside Cursor — where a lot of the world's coding assistance already happens — and xAI has a genuine volume play: good-enough frontier reasoning at a price that makes "just route everything through it" a defensible default for a lot of teams.
The honest caveats
- Many scores are AA-normalized or first-party. The GPQA and coding numbers are strong, but treat the full table as "vendor + early-tracker" until neutral boards fully settle.
- 500K context is the real limitation. If your workflow leans on 1M-token whole-repo prompts, this is where Grok 4.5 gives ground to Sol and Kimi K3.
- Trained on Cursor tokens cuts both ways. It's superb inside Cursor; whether that advantage transfers cleanly to your own harness is worth testing.
My take
Grok 4.5 is the clearest "value frontier" pick of the July wave — the one I'd reach for on high-volume coding traffic where $30/M output tokens would bleed the budget dry. It doesn't top the charts, but it doesn't need to: it's close enough, cheap enough, and sitting right inside the editor. That's exactly the kind of model a portability-first setup loves — cheap capacity you can pour volume into, while the pricier flagships handle the hard 10%.
Sources
- xAI — Introducing Grok 4.5
- BenchLM — Grok 4.5 benchmarks
- ExplainX — Grok 4.5 launch, benchmarks & pricing
- FelloAI — Grok 4.5 is here
- Apidog — Grok 4.5 benchmarks: what xAI published
Numbers are as reported at launch (8 July 2026) and will move as independent benchmarks land. Spot a corrected figure? Tell me and I'll update.


