Which LLM Should You Use?

A practical comparison of frontier language models for business applications. Cut through the marketing and find the right model for your specific needs.

Last updated: July 2026

Model Comparison

ModelContextAPI Pricing (in/out)Consumer Access
Claude Fable 5
Anthropic
1M tokens$10 / $50 per 1M tokensClaude Pro $20/mo (Pro/Max/Team/Enterprise)
Claude Opus 5
Anthropic
1M tokens$5 / $25 per 1M tokensClaude Pro $20/mo (strongest on Pro); new default on Claude Max
Claude Haiku 4.5
Anthropic
200K tokens$1 / $5 per 1M tokensClaude Pro $20/mo
GPT-5.6 Sol
OpenAI
1.05M tokens$5 / $30 per 1M tokensChatGPT Plus $20/mo
GPT-5.6 Terra
OpenAI
1.05M tokens$2.50 / $15 per 1M tokensChatGPT Plus $20/mo
GPT-5.6 Luna
OpenAI
1.05M tokens$1 / $6 per 1M tokensChatGPT Plus $20/mo
Gemini 3.1 Pro
Google
1M tokens$2 / $12 per 1M tokens (up to 200K; $4/$18 above)Google AI Pro $19.99/mo (Deep Think needs AI Ultra $99.99/mo)
Gemini 3.6 Flash
Google
1M tokens (64K max output)$1.50 / $7.50 per 1M tokensGoogle AI Pro $19.99/mo; runs on the free Gemini app tier
GLM-5.2
Z.ai
1M tokens$1.40 / $4.40 per 1M tokens ($0.26/M cached input)Open weights (MIT) or Z.ai plans from $12.60/mo
Grok 4.5
SpaceXAI (formerly xAI)
500K tokens$2 / $6 per 1M tokens ($0.50/M cached input)SuperGrok $30/mo (Heavy $300/mo for top rate limits)
Llama 4 Maverick
Meta
1M tokensSelf-hosted / ~$0.30–$0.49 per 1M (blended)Open weights (API varies by provider)
DeepSeek V4 Pro
DeepSeek
1M tokens$0.435 / $0.87 per 1M tokens off-peak; 2× (to $0.87/$1.74) during Beijing peak hours (9am–12pm & 2pm–6pm)API only

Strengths & Limitations

Claude Fable 5

Anthropic

Strengths

  • State-of-the-art on nearly all capability benchmarks: coding, vision, research
  • Full 1M context at standard pricing, no long-context surcharge
  • Leads the EQ-Bench Creative Writing leaderboard (Elo 2230)
  • 90% prompt-caching discount on cached reads ($1 per 1M)

Limitations

  • Priciest Anthropic model - several times Sonnet 5's per-token cost
  • Self-reported 80.3% SWE-bench Pro score contested after OpenAI raised audit-methodology concerns
  • Safeguards route flagged cybersecurity/bio/chem queries to other models (under 5% of sessions)
  • Anthropic has begun a restricted preview of Claude Mythos, a still more capable model limited to roughly 50 critical-infrastructure partners under Project Glasswing - not broadly available

Claude Opus 5

Anthropic

Strengths

  • Replaced Opus 4.8 on July 24, 2026 at the same price, with a full generation of capability gains
  • 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro, more than 10 points above Opus 4.8
  • Leads Anthropic's new agentic-coding benchmark, ahead of both Fable 5 and GPT-5.6 Sol
  • Full 1M context at standard pricing, with adaptive thinking on by default and a five-level effort toggle

Limitations

  • At $5/$25 per 1M tokens, still costs roughly 1.7x Sonnet 5's introductory rate ($2/$10 through Aug 31, 2026) for high-volume, well-scoped work
  • Trails Claude Mythos 5 (80.3%) and Claude Fable 5 (80.0%) on SWE-bench Pro by under a point
  • Anthropic's own evaluations still put Mythos 5 ahead on cybersecurity tasks
  • Released July 24, 2026, so independent reliability data beyond Anthropic's own benchmarks is still thin

Claude Haiku 4.5

Anthropic

Strengths

  • Fastest and cheapest Anthropic model
  • Low latency for real-time and high-volume applications
  • Strong on classification, extraction, and summarisation
  • 90% cheaper than Fable 5 on input tokens

Limitations

  • Smaller context window (200K vs 1M for Sonnet/Fable)
  • Less capable on complex multi-step reasoning
  • Not suited for nuanced creative or deep analytical work

GPT-5.6 Sol

OpenAI

Strengths

  • Replaced GPT-5.5 as OpenAI's flagship on July 9, 2026, with the family's strongest benchmark performance
  • Strong agentic coding: terminal workflows and multi-step tool coordination (88.8% Terminal-Bench 2.1)
  • Native text-and-image multimodal input across the full 1.05M token context window
  • More predictable prompt caching, with explicit cache breakpoints and a 30-minute minimum cache life

Limitations

  • 64.6% on SWE-bench Pro - OpenAI has publicly disputed the audit methodology behind rival scores, adding noise to head-to-head comparisons
  • Explicit chain-of-thought reasoning trades higher latency and token usage for accuracy on math/reasoning tasks
  • OpenAI's own System Card flagged a 6.3x higher rate of unauthorized file-deletion behavior than GPT-5.5 (0.019% vs 0.003%); multiple developers reported Sol deleting files unprompted after the July 9 launch
  • Output capped at 128K tokens regardless of input context size

GPT-5.6 Terra

OpenAI

Strengths

  • Practical center of the GPT-5.6 lineup - roughly GPT-5.5-class performance at half the price
  • OpenAI's own recommended default tier for most production teams
  • Full 1.05M token context window, same as Sol
  • Same predictable prompt-caching improvements as Sol

Limitations

  • Noticeably behind Sol on the hardest multi-step reasoning and coding benchmarks
  • Launched July 9, 2026, so third-party benchmark and reliability data is still thin
  • Still pricier than GLM-5.2 or DeepSeek V4 Pro for comparable coding workloads

GPT-5.6 Luna

OpenAI

Strengths

  • Cost champion of the family - about 24 benchmark points per estimated API dollar on DeepSWE, versus 3.2 for Claude Fable 5
  • Roughly 85% of Sol's quality at about one-fifth the price for high-volume pipelines
  • Full 1.05M token context window, same as Sol and Terra
  • Cheapest way to access OpenAI's July 2026 model generation

Limitations

  • Weak long-context recall - 41.3% on the MRCR benchmark
  • Gap to Sol widens significantly on the hardest reasoning tasks
  • Not suited to nuanced creative or deep analytical work

Gemini 3.1 Pro

Google

Strengths

  • Adjustable reasoning depth (Low/Medium/High) — High mode runs as a "Deep Think Mini"
  • 77.1% on ARC-AGI-2, leading abstract reasoning benchmarks
  • Natively multimodal: text, images, video, and audio
  • Strong agentic tool-use with fewer wasted tool calls

Limitations

  • Pricing doubles above 200K tokens ($4/$18 per 1M)
  • Still labelled preview as of July 2026, with some output inconsistency pre-GA
  • Flatter creative/emotional tone versus earlier Gemini releases

Gemini 3.6 Flash

Google

Strengths

  • Replaced Gemini 3.5 Flash on July 21, 2026 as Google's default Flash-tier model
  • Same input price as 3.5 Flash but 17% cheaper output pricing ($7.50 vs $9.00 per 1M), with 17% fewer output tokens on comparable tasks
  • Built for agentic workloads: tool calling, subagent orchestration, multi-step workflows, with built-in Computer Use
  • Native multimodal: text, images, video, audio, and PDFs

Limitations

  • Max output capped at 64K tokens
  • Not ideal for long-form nuanced writing or deep analytical work
  • Fewer third-party integrations than OpenAI

GLM-5.2

Z.ai

Strengths

  • Led long-horizon open-weight coding benchmarks at launch (SWE-bench Pro 62.1)
  • Fully open-weight under an unrestricted MIT licence — free to download, fine-tune, self-host
  • Strong tool use: 77.0 on MCP-Atlas
  • A fraction of the API cost of frontier closed models for comparable coding performance

Limitations

  • Benchmarks were released alongside the open weights rather than at initial launch — an unusual sequencing that drew scrutiny
  • Z.ai-hosted API routes data through China-based infrastructure — a concern for regulated or sensitive workloads
  • Trails Claude Fable 5 and GPT-5.6 Sol on the hardest frontier reasoning benchmarks

Grok 4.5

SpaceXAI (formerly xAI)

Strengths

  • Built on a 1.5T-parameter "V9" foundation trained partly on Cursor coding-agent data; pitched as "Opus-class" at a fraction of the cost
  • 4.2x more token-efficient than Opus 4.8 on SWE-bench Pro tasks (15,954 vs 67,020 average output tokens)
  • Aggressive pricing relative to other frontier-tier models
  • Leads on Terminal-Bench 2.1 and provider-harness DeepSWE 1.0 scores

Limitations

  • Hallucination rate roughly doubled versus Grok 4.3 (54% on AA-Omniscience, up from 25%), higher than Claude models
  • Smaller 500K context window, down from Grok 4.3's 1M
  • Not available in the EU at launch (July 8, 2026) pending EU AI Act compliance review
  • Trails Opus 4.8 on neutral SWE-bench Pro and DeepSWE 1.1 runs despite the "Opus-class" framing

Llama 4 Maverick

Meta

Strengths

  • Open weights under a commercial-friendly license for most companies, for complete data sovereignty
  • Natively multimodal (128-expert MoE architecture)
  • Full fine-tuning and deployment flexibility across local stacks (llama.cpp, vLLM, Ollama)
  • Still the only current Meta model available as open weights for self-hosting

Limitations

  • Requires own infrastructure to deploy at scale
  • Resource-intensive (400B total parameters)
  • EU-domiciled companies and users are prohibited from using or distributing the model under its license
  • Frozen, unmaintained lineage — Meta's July 9, 2026 Muse Spark 1.1 launch (closed-weight, paid API at $1.25/$4.25 per 1M tokens) is its actual current flagship, with no open-weight successor to Llama 4 announced

DeepSeek V4 Pro

DeepSeek

Strengths

  • Reached general availability on July 20, 2026, after a preview period that began April 24
  • 80.6% on SWE-bench Verified — within 0.2 points of Claude Opus 4.6, at roughly a seventh of the output price
  • 1.6T-parameter Mixture-of-Experts architecture (~49B active params), 1M token context window, unchanged from preview
  • Open weights available under MIT for self-hosted deployment

Limitations

  • Data stored on China-based servers — significant privacy risk
  • Peak/off-peak surge pricing is now confirmed and in effect, doubling cost during Beijing business hours
  • Legacy deepseek-chat and deepseek-reasoner model aliases were retired on July 24, 2026 — calls using those names now fail

Use Case Recommendations

Different tasks demand different trade-offs. Here are our recommendations based on common business scenarios.

Use CaseRecommendedAlternativesNotes
Complex Analysis & ResearchClaude Fable 5
GPT-5.6 SolGemini 3.1 Pro
When accuracy and depth matter more than speed or cost
Production ApplicationsClaude Opus 5
GPT-5.6 TerraGLM-5.2
Anthropic's new default for quality-sensitive production workloads
Long Document ProcessingGemini 3.1 Pro
Claude Fable 5GPT-5.6 Sol
Native 1M context; watch the pricing step above 200K tokens
Reasoning, Math & ScienceGemini 3.1 Pro
Claude Fable 5GPT-5.6 Sol
Deep Think mode leads on abstract reasoning (77.1% ARC-AGI-2)
Customer Service & ChatbotsGemini 3.6 Flash
Claude Haiku 4.5GPT-5.6 Luna
17% cheaper output than the prior Flash generation; handles complex queries with tool use
Budget-Conscious ProjectsDeepSeek V4 Pro
GLM-5.2GPT-5.6 Luna
Near-frontier performance at a fraction of the cost
On-Premise / Air-GappedLlama 4 Maverick
GLM-5.2 (self-hosted)DeepSeek V4 Pro (self-hosted)
When data cannot leave your infrastructure; note Llama 4 is off-limits for EU-domiciled deployments
Creative WritingClaude Fable 5
Claude Opus 5GPT-5.6 Sol
Leads EQ-Bench Creative Writing leaderboard (Elo 2230)
Code GenerationClaude Opus 5
GLM-5.2Grok 4.5
Leads Anthropic's agentic-coding benchmark at half of Fable 5's price; GLM-5.2 and Grok 4.5 for cheaper agentic coding
Multimodal (Images/Documents/Video)GPT-5.6 Sol
Gemini 3.1 ProGrok 4.5
Broadest native multimodal understanding across formats and file types
Real-Time & Social ContextGrok 4.5
GPT-5.6 SolGemini 3.6 Flash
Native access to real-time X data for current-events tasks via SpaceXAI's X integration

The Model is Only Part of the Equation

Choosing the right LLM matters, but how you architect your system, design your prompts, and integrate AI into your workflows determines success. We help organisations move from model selection to production deployment.

Discuss Your AI Project

* Pricing reflects July 2026 rates and may change. Check provider websites for current pricing.

* Model capabilities and context windows are based on publicly available documentation.

* Recommendations reflect our experience across client engagements. Your specific requirements may differ.

Get AI insights in your inbox
Practical analysis on AI strategy, products, and technical leadership
No more than one newsletter a month