Which LLM Should You Use?

A practical comparison of frontier language models for business applications. Cut through the marketing and find the right model for your specific needs.

Last updated: August 2026

Model Comparison

ModelContextAPI Pricing (in/out)Consumer Access
Claude Fable 5
Anthropic
1M tokens$10 / $50 per 1M tokensClaude Pro $20/mo (Pro/Max/Team/Enterprise)
Claude Opus 5
Anthropic
1M tokens$5 / $25 per 1M tokensClaude Pro $20/mo (strongest on Pro); new default on Claude Max
Claude Haiku 4.5
Anthropic
200K tokens$1 / $5 per 1M tokensClaude Pro $20/mo
GPT-5.6 Sol
OpenAI
1.05M tokens$4 / $20 per 1M tokens (promotional through Nov 21, 2026; reverts to $5/$30)ChatGPT Plus $20/mo
GPT-5.6 Terra
OpenAI
1.05M tokens$2 / $12 per 1M tokensChatGPT Plus $20/mo
GPT-5.6 Luna
OpenAI
1.05M tokens$0.20 / $1.20 per 1M tokensFree tier: unlimited text chats from Aug 7, 2026; also ChatGPT Plus $20/mo
Gemini 3.1 Pro
Google
1M tokens$2 / $12 per 1M tokens (up to 200K; $4/$18 above)Google AI Pro $19.99/mo (Deep Think needs AI Ultra $99.99/mo)
Gemini 3.7 Flash
Google
1M tokens (64K max output)$0.75 / $3.75 per 1M tokens (intro rate through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027)Google AI Pro $19.99/mo; runs on the free Gemini app tier
GLM-5.3
Z.ai
1M tokens$1.40 / $4.40 per 1M tokens (official GLM-5.3 rate; $0.26/M cached input)GLM Coding Plan from $12.60/mo (annual) or $18/mo; 753B-param open weights released Aug 28, 2026 under a bespoke license, not MIT
Grok 4.6
SpaceXAI (formerly xAI)
500K tokens$2 / $6 per 1M tokens (<200K prompt); $4 / $12 above 200K ($0.50/M cached input)SuperGrok $30/mo (Heavy $300/mo for top rate limits)
Muse Glimmer
Meta
131K tokensSelf-hosted / ~$0.30–$1.20 per 1M via third-party hosts (no official Meta API price)Open weights (Apache 2.0)
DeepSeek V4 Pro
DeepSeek
1M tokens$1.32 / $3.96 per 1M peak, $0.66 / $1.98 off-peak (up from $0.435/$0.87 flat)API only

Strengths & Limitations

Claude Fable 5

Anthropic

Strengths

  • State-of-the-art on nearly all capability benchmarks: coding, vision, research
  • Full 1M context at standard pricing, no long-context surcharge
  • Leads the EQ-Bench Creative Writing leaderboard (Elo 2230)
  • 90% prompt-caching discount on cached reads ($1 per 1M)

Limitations

  • Priciest Anthropic model - several times Sonnet 5's per-token cost
  • Self-reported 80.3% SWE-bench Pro score contested after OpenAI raised audit-methodology concerns
  • Safeguards route flagged cybersecurity/bio/chem queries to other models (under 5% of sessions)
  • Anthropic has begun a restricted preview of Claude Mythos, a still more capable model limited to roughly 50 critical-infrastructure partners under Project Glasswing - not broadly available

Claude Opus 5

Anthropic

Strengths

  • Replaced Opus 4.8 on July 24, 2026 at the same price, with a full generation of capability gains
  • 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro, more than 10 points above Opus 4.8
  • Leads Anthropic's new agentic-coding benchmark, ahead of both Fable 5 and GPT-5.6 Sol
  • Full 1M context at standard pricing, with adaptive thinking on by default and a five-level effort toggle

Limitations

  • At $5/$25 per 1M tokens, still costs roughly 1.7x Sonnet 5's now-permanent $2/$10 rate for high-volume, well-scoped work
  • Trails Claude Mythos 5 (80.3%) and Claude Fable 5 (80.0%) on SWE-bench Pro by under a point
  • Anthropic's own evaluations still put Mythos 5 ahead on cybersecurity tasks
  • Released July 24, 2026, so independent reliability data beyond Anthropic's own benchmarks is still thin

Claude Haiku 4.5

Anthropic

Strengths

  • Fastest and cheapest Anthropic model
  • Low latency for real-time and high-volume applications
  • Strong on classification, extraction, and summarisation
  • 90% cheaper than Fable 5 on input tokens

Limitations

  • Smaller context window (200K vs 1M for Sonnet/Fable)
  • Less capable on complex multi-step reasoning
  • Not suited for nuanced creative or deep analytical work

GPT-5.6 Sol

OpenAI

Strengths

  • Replaced GPT-5.5 as OpenAI's flagship on July 9, 2026, with the family's strongest benchmark performance
  • API pricing cut over 20% on Aug 21, 2026 (input $5→$4, output $30→$20 per 1M tokens) - the promotional rate now undercuts Claude Opus 5 on both input and output
  • Strong agentic coding: terminal workflows and multi-step tool coordination (88.8% Terminal-Bench 2.1)
  • Native text-and-image multimodal input across the full 1.05M token context window
  • August 6, 2026 ChatGPT update folded Instant and reasoning into one Sol experience with a thinking slider; OpenAI's internal evaluation found 68% fewer factually-erroneous responses than GPT-5.5 Instant

Limitations

  • 64.6% on SWE-bench Pro - OpenAI has publicly disputed the audit methodology behind rival scores, adding noise to head-to-head comparisons
  • The Aug 21, 2026 price cut is promotional, running through Nov 21, 2026, after which list price reverts to $5/$30 per 1M tokens
  • OpenAI's own System Card flagged a 6.3x higher rate of unauthorized file-deletion behavior than GPT-5.5 (0.019% vs 0.003%); multiple developers reported Sol deleting files unprompted after the July 9 launch
  • Output capped at 128K tokens regardless of input context size

GPT-5.6 Terra

OpenAI

Strengths

  • Practical center of the GPT-5.6 lineup - roughly GPT-5.5-class performance at a fraction of the price
  • Repriced from $2.50/$15 to $2/$12 on July 30, 2026 as part of an OpenAI-wide GPT-5.6 price cut
  • OpenAI's own recommended default tier for most production teams
  • Full 1.05M token context window, same as Sol, with the same predictable prompt-caching improvements

Limitations

  • Noticeably behind Sol on the hardest multi-step reasoning and coding benchmarks
  • Launched July 9, 2026, so third-party benchmark and reliability data is still thin
  • Still pricier than GLM-5.2 or DeepSeek V4 Pro for comparable coding workloads

GPT-5.6 Luna

OpenAI

Strengths

  • Cost champion of the family - cut 80% on July 30, 2026 from its $1/$6 launch price to $0.20/$1.20
  • Now roughly 25x cheaper than Sol on both input and output, the widest price gap in the lineup
  • Full 1.05M token context window, same as Sol and Terra
  • Became the default model for Free and Go ChatGPT users on August 7, 2026, with unlimited text chats and a Think button for harder questions

Limitations

  • Weak long-context recall - 41.3% on the MRCR benchmark
  • Gap to Sol widens significantly on the hardest reasoning tasks
  • Not suited to nuanced creative or deep analytical work

Gemini 3.1 Pro

Google

Strengths

  • Adjustable reasoning depth (Low/Medium/High) — High mode runs as a "Deep Think Mini"
  • 77.1% on ARC-AGI-2, leading abstract reasoning benchmarks
  • Natively multimodal: text, images, video, and audio
  • Strong agentic tool-use with fewer wasted tool calls

Limitations

  • Pricing doubles above 200K tokens ($4/$18 per 1M)
  • Still labelled preview as of August 2026, with some output inconsistency pre-GA
  • Flatter creative/emotional tone versus earlier Gemini releases

Gemini 3.7 Flash

Google

Strengths

  • Replaced Gemini 3.6 Flash on Aug 13, 2026 as Google's workhorse Flash-tier model for coding and agents, just three weeks after 3.6 Flash shipped
  • Introductory pricing is half the list rate through the end of 2026 - Google temporarily applied the same $0.75/$3.75 rate to 3.6 Flash during the promotion
  • Built for agentic workloads: tool calling, subagent orchestration, multi-step workflows, with built-in Computer Use
  • Native multimodal: text, images, video, audio, and PDFs

Limitations

  • Introductory pricing expires Jan 1, 2027, when both input and output rates double to $1.50/$7.50
  • Max output still capped at 64K tokens
  • Not ideal for long-form nuanced writing or deep analytical work
  • Fewer third-party integrations than OpenAI

GLM-5.3

Z.ai

Strengths

  • Released Aug 14, 2026 on the same GLM-5.2 base model, with every capability gain from a 10x expansion of long-horizon post-training
  • Largest gains on long-horizon coding: Terminal-Bench 3.0 jumped from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9
  • Rolled out immediately to all existing GLM Coding Plan subscribers, with API access live from day one
  • Open weights (753B-param MoE) now live on Hugging Face — free to download, fine-tune, redistribute and use commercially for virtually all users

Limitations

  • Weights ship under a bespoke "GLM-5.3 License," not the MIT license GLM-5.2 used — Model-as-a-Service operators above $10B/12-month revenue need a Z.ai security review before commercial use
  • Post-training produced an offensive-security capability jump the company says outgrew its own expectations, surfacing 1,097 critical/high-severity vulnerabilities in evaluated software
  • Trails GPT-5.6 Sol by 6.3 points on Terminal-Bench 3.0 (34.6 vs 28.3) and 5.8 points on DeepSWE v1.1 (72.7 vs 66.9)
  • Z.ai-hosted API still routes through China-based infrastructure — a concern for regulated or sensitive workloads

Grok 4.6

SpaceXAI (formerly xAI)

Strengths

  • Released Aug 12, 2026, just 35 days after Grok 4.5, live across Grok Build, Cursor, Grok Bot and the SpaceXAI API
  • Matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index (~61), about two points behind Claude Opus 5, for roughly 60% less than Sol's API price
  • Terminal-Bench 3.0 score of 26%, nearly double Grok 4.5's 15.7%
  • Pricing held flat versus Grok 4.5 despite the capability gain

Limitations

  • Weak knowledge calibration: 48.2% accuracy on AA-Omniscience with only a 65.7% non-hallucination rate — fabricates an answer roughly one time in three when it doesn't know one
  • Terminal-Bench 3.0 score of 26% still lags the top coding-focused frontier models
  • Context window unchanged at 500K, smaller than the 1M+ windows most rivals now offer
  • Cached-input pricing rose to $0.50/M from Grok 4.5's $0.30/M

Muse Glimmer

Meta

Strengths

  • Released Aug 10, 2026 under Apache 2.0 — Meta's most permissive license yet, with no commercial-use or EU restrictions (unlike Llama 4's license)
  • 29.6B-parameter dense model distilled from Meta's Muse Spark, natively multimodal (text and images) and tuned for local agent and coding workflows
  • Runs on consumer hardware — a 17GB 4-bit quant fits a single consumer GPU or Mac
  • Replaces Llama 4 Maverick as Meta's current open-weight offering

Limitations

  • Much smaller context window (131K) than Llama 4 Maverick's 1M
  • Meta's actual flagship remains the closed, paid Muse Spark 1.2 ($1.25/$4.25 per 1M, or $0.10/$0.20 on a data-sharing "Contributor" tier) — Glimmer is a distilled, smaller sibling, not directly comparable on capability
  • Meta says it will also open-weight Muse Spark 1.2 itself, but has confirmed no date or license terms as of this refresh
  • No official Meta-hosted API — total cost depends on self-hosting or third-party providers

DeepSeek V4 Pro

DeepSeek

Strengths

  • Reached general availability on July 20, 2026, after a preview period that began April 24
  • 80.6% on SWE-bench Verified — within 0.2 points of Claude Opus 4.6, at a fraction of the output price even after the August price rise
  • 1.6T-parameter Mixture-of-Experts architecture (~49B active params), 1M token context window, unchanged from preview
  • Open weights available under MIT for self-hosted deployment; a smaller V4-Flash sibling now runs $0.44/$1.32 peak, $0.22/$0.66 off-peak per 1M tokens

Limitations

  • Data stored on China-based servers — significant privacy risk
  • New peak/off-peak surge pricing took effect Aug 16, 2026, roughly tripling to quadrupling rates versus the prior flat price, with peak hours running 01:00–04:00 and 06:00–10:00 UTC
  • Off-peak rates are still higher than the old flat rate in every category — a price increase with a time-of-day discount, not a genuine cut
  • Legacy deepseek-chat and deepseek-reasoner model aliases were retired on July 24, 2026 — calls using those names now fail

Use Case Recommendations

Different tasks demand different trade-offs. Here are Agathon's recommendations based on common business scenarios.

Use CaseRecommendedAlternativesNotes
Complex Analysis & ResearchClaude Fable 5
GPT-5.6 SolGemini 3.1 Pro
When accuracy and depth matter more than speed or cost
Production ApplicationsClaude Opus 5
GPT-5.6 TerraGLM-5.3
Anthropic's new default for quality-sensitive production workloads
Long Document ProcessingGemini 3.1 Pro
Claude Fable 5GPT-5.6 Sol
Native 1M context; watch the pricing step above 200K tokens
Reasoning, Math & ScienceGemini 3.1 Pro
Claude Fable 5GPT-5.6 Sol
Deep Think mode leads on abstract reasoning (77.1% ARC-AGI-2)
Customer Service & ChatbotsGemini 3.7 Flash
Claude Haiku 4.5GPT-5.6 Luna
Half price through end of 2026 on an introductory rate; handles complex queries with tool use
Budget-Conscious ProjectsDeepSeek V4 Pro
GLM-5.3GPT-5.6 Luna
Near-frontier performance well below frontier prices, even after the August price rise
On-Premise / Air-GappedMuse Glimmer
GLM-5.3 (self-hosted)DeepSeek V4 Pro (self-hosted)
When data cannot leave your infrastructure; Apache 2.0 licensed with no EU restrictions, and light enough for a single GPU
Creative WritingClaude Fable 5
Claude Opus 5GPT-5.6 Sol
Leads EQ-Bench Creative Writing leaderboard (Elo 2230)
Code GenerationClaude Opus 5
GLM-5.3Grok 4.6
Leads Anthropic's agentic-coding benchmark at half of Fable 5's price; GLM-5.3 and Grok 4.6 for cheaper agentic coding
Multimodal (Images/Documents/Video)GPT-5.6 Sol
Gemini 3.1 ProGrok 4.6
Broadest native multimodal understanding across formats and file types
Real-Time & Social ContextGrok 4.6
GPT-5.6 SolGemini 3.7 Flash
Native access to real-time X data for current-events tasks via SpaceXAI's X integration

The Model is Only Part of the Equation

Choosing the right LLM matters, but how you architect your system, design your prompts, and integrate AI into your workflows determines success. Agathon helps organisations move from model selection to production deployment.

Discuss Your AI Project

* Pricing reflects August 2026 rates and may change. Check provider websites for current pricing.

* GPT-5.6 Sol's $4/$20 per 1M token API rate is a promotional price running through Nov 21, 2026; list price reverts to $5/$30 afterward.

* Model capabilities and context windows are based on publicly available documentation.

* Recommendations reflect Agathon's experience across client engagements. Your specific requirements may differ.

Get AI insights in your inbox
Practical analysis on AI strategy, products, and technical leadership
No more than one briefing a month