Which LLM Should You Use?

A practical comparison of frontier language models for business applications. Cut through the marketing and find the right model for your specific needs.

Last updated: July 2026

Model Comparison

ModelContextAPI Pricing (in/out)Consumer Access
Claude Fable 5
Anthropic
1M tokens$10 / $50 per 1M tokensClaude Pro $20/mo (Pro/Max/Team/Enterprise)
Claude Opus 4.8
Anthropic
1M tokens$5 / $25 per 1M tokensClaude Pro $20/mo
Claude Sonnet 5
Anthropic
1M tokens$2 / $10 per 1M tokens (intro, to Aug 2026; $3/$15 standard)Claude Pro $20/mo
Claude Haiku 4.5
Anthropic
200K tokens$1 / $5 per 1M tokensClaude Pro $20/mo
GPT-5.5
OpenAI
1M tokens$5 / $30 per 1M tokensChatGPT Plus $20/mo
Gemini 3.1 Pro
Google
1M tokens$2 / $12 per 1M tokens (up to 200K; $4/$18 above)Google AI Pro $19.99/mo (Deep Think needs AI Ultra $99.99/mo)
Gemini 3.5 Flash
Google
1M tokens$1.50 / $9.00 per 1M tokensGoogle AI Pro $19.99/mo
GLM-5.2
Z.ai
1M tokens$1.40 / $4.40 per 1M tokens ($0.26/M cached input)Open weights (MIT) or Z.ai plans from $12.60/mo
Grok 4.3
xAI
1M tokens$1.25 / $2.50 per 1M tokensSuperGrok $30/mo (Heavy $300/mo for top rate limits)
Llama 4 Maverick
Meta
1M tokensSelf-hosted / ~$0.30–$0.49 per 1M (blended)Open weights (API varies by provider)
DeepSeek V4 Pro
DeepSeek
1M tokens$0.435 / $0.87 per 1M tokens (2× during Beijing peak hours)API only

Strengths & Limitations

Claude Fable 5

Anthropic

Strengths

  • State-of-the-art on nearly all capability benchmarks: coding, vision, research
  • Full 1M context at standard pricing, no long-context surcharge
  • Exceptional long-horizon agentic coding (migrated a 50M-line codebase for Stripe in a day)
  • 90% prompt-caching discount on cached reads ($1 per 1M)

Limitations

  • Exactly double the cost of Opus 4.8 on both input and output
  • Self-reported 80.3% SWE-bench Pro score contested by independent evaluators
  • Safeguards route flagged cybersecurity/bio/chem queries to Opus 4.8 (under 5% of sessions)
  • Briefly pulled worldwide in June 2026 under US export controls before returning July 1

Claude Opus 4.8

Anthropic

Strengths

  • Second only to Fable 5 on raw capability, at half the cost
  • Long-horizon coding with high autonomy (69.2% SWE-bench Pro)
  • Full 1M context at standard pricing
  • Leads EQ-Bench Creative Writing (Elo 2216)

Limitations

  • Higher API cost than Sonnet 5
  • Fast Mode adds further cost ($10/$50 per 1M tokens)
  • No image generation

Claude Sonnet 5

Anthropic

Strengths

  • Replaced Sonnet 4.6 on June 30, 2026 with adaptive thinking and better instruction-following
  • Near-Opus coding quality at a fraction of the cost
  • Full 1M context at standard pricing
  • Best speed-to-quality ratio for production workloads

Limitations

  • Newer tokenizer produces ~30% more tokens for the same text, affecting effective cost
  • Less depth on highly complex multi-step reasoning than Opus 4.8 or Fable 5
  • No image generation

Claude Haiku 4.5

Anthropic

Strengths

  • Fastest and cheapest Anthropic model
  • Low latency for real-time and high-volume applications
  • Strong on classification, extraction, and summarisation
  • 90% cheaper than Opus 4.8 on input tokens

Limitations

  • Smaller context window (200K vs 1M for Opus/Sonnet/Fable)
  • Less capable on complex multi-step reasoning
  • Not suited for nuanced creative or deep analytical work

GPT-5.5

OpenAI

Strengths

  • Strong multimodal capabilities (images, documents, audio)
  • Large ecosystem of tools and integrations
  • Broad general knowledge and creative tasks
  • Batch and Flex processing at 50% discount

Limitations

  • Input pricing doubles above 272K tokens ($10/M)
  • GPT-5.5 Pro variant very expensive ($30/$180 per 1M)
  • GPT-5.6 Sol/Terra/Luna previewed to select partners in July 2026, not yet broadly available

Gemini 3.1 Pro

Google

Strengths

  • Adjustable reasoning depth (Low/Medium/High) — High mode runs as a "Deep Think Mini"
  • 77.1% on ARC-AGI-2, leading abstract reasoning benchmarks
  • Natively multimodal: text, images, video, and audio
  • Strong agentic tool-use with fewer wasted tool calls

Limitations

  • Pricing doubles above 200K tokens ($4/$18 per 1M)
  • Still labelled preview as of July 2026, with some output inconsistency pre-GA
  • Flatter creative/emotional tone versus earlier Gemini releases

Gemini 3.5 Flash

Google

Strengths

  • Google's current default Gemini model — outperforms Gemini 3.1 Pro on coding and agentic benchmarks
  • 4× faster token output than competing frontier models
  • Built for agentic workloads: tool calling, subagent orchestration, multi-step workflows
  • Native multimodal: text, images, video, audio, and PDFs

Limitations

  • 5× more expensive input than the outgoing Gemini 2.5 Flash
  • Not ideal for long-form nuanced writing or deep analytical work
  • Fewer third-party integrations than OpenAI

GLM-5.2

Z.ai

Strengths

  • Beats GPT-5.5 on multiple long-horizon coding benchmarks (SWE-bench Pro 62.1 vs 58.6)
  • Fully open-weight under an unrestricted MIT licence — free to download, fine-tune, self-host
  • Strong tool use: 77.0 on MCP-Atlas versus GPT-5.5’s 75.3
  • Roughly 1/6th the API cost of GPT-5.5 for comparable coding performance

Limitations

  • Benchmarks were released alongside the open weights rather than at initial launch — an unusual sequencing that drew scrutiny
  • Z.ai-hosted API routes data through China-based infrastructure — a concern for regulated or sensitive workloads
  • Trails Claude Opus 4.8 and Fable 5 on the hardest frontier reasoning benchmarks

Grok 4.3

xAI

Strengths

  • Native video input — an xAI first among frontier models
  • Aggressive pricing relative to other frontier-tier models
  • Real-time access to X/Twitter data for current-events context
  • Large 1M token context window

Limitations

  • Smaller enterprise tooling and integration ecosystem than the Big Three
  • Consumer pricing fragmented across five separate tiers
  • Benchmarks less independently verified than Anthropic/OpenAI/Google

Llama 4 Maverick

Meta

Strengths

  • Open weights for complete data sovereignty
  • Natively multimodal (128-expert MoE architecture)
  • Competitive with GPT-4o and Gemini 2.0 Flash on benchmarks
  • Full fine-tuning and deployment flexibility

Limitations

  • Requires own infrastructure to deploy at scale
  • Resource-intensive (400B total parameters)
  • Meta’s newer Muse Spark model is closed-weight, so Llama 4 remains the open-source option

DeepSeek V4 Pro

DeepSeek

Strengths

  • Extraordinary value — a fraction of the cost of Claude Sonnet 5
  • 1M token context window with thinking and standard modes
  • Strong coding and reasoning performance
  • Open weights available for self-hosted deployment

Limitations

  • Data stored on China-based servers — significant privacy risk
  • New peak-hour pricing doubles cost during Beijing business hours from the official V4 launch
  • Not suitable for sensitive, regulated, or enterprise data

Use Case Recommendations

Different tasks demand different trade-offs. Here are our recommendations based on common business scenarios.

Use CaseRecommendedAlternativesNotes
Complex Analysis & ResearchClaude Fable 5
Claude Opus 4.8GPT-5.5
When accuracy and depth matter more than speed or cost
Production ApplicationsClaude Sonnet 5
GPT-5.5GLM-5.2
Balance of quality, speed, and cost for real workloads
Long Document ProcessingGemini 3.1 Pro
Claude Fable 5Claude Opus 4.8
Native 1M context; watch the pricing step above 200K tokens
Reasoning, Math & ScienceGemini 3.1 Pro
Claude Fable 5GPT-5.5
Deep Think mode leads on abstract reasoning (77.1% ARC-AGI-2)
Customer Service & ChatbotsGemini 3.5 Flash
Claude Haiku 4.5GLM-5.2
4× faster token output than competing frontier models; handles complex queries with tool use
Budget-Conscious ProjectsDeepSeek V4 Pro
GLM-5.2Llama 4 Maverick
Near-frontier performance at a fraction of the cost
On-Premise / Air-GappedLlama 4 Maverick
GLM-5.2 (self-hosted)DeepSeek V4 Pro (self-hosted)
When data cannot leave your infrastructure
Creative WritingClaude Opus 4.8
Claude Fable 5GPT-5.5
Leads EQ-Bench Creative Writing leaderboard (Elo 2216)
Code GenerationClaude Sonnet 5
GLM-5.2Claude Fable 5
Strong SWE-bench results at lower cost than Fable 5; GLM-5.2 for open-weight budget builds
Multimodal (Images/Documents/Video)GPT-5.5
Gemini 3.1 ProGrok 4.3
Native multimodal understanding across formats and file types
Real-Time & Social ContextGrok 4.3
GPT-5.5Gemini 3.5 Flash
Native access to real-time X/Twitter data for current-events tasks

The Model is Only Part of the Equation

Choosing the right LLM matters, but how you architect your system, design your prompts, and integrate AI into your workflows determines success. We help organisations move from model selection to production deployment.

Discuss Your AI Project

* Pricing reflects July 2026 rates and may change. Check provider websites for current pricing.

* Model capabilities and context windows are based on publicly available documentation.

* Recommendations reflect our experience across client engagements. Your specific requirements may differ.

Get AI insights in your inbox
Practical analysis on AI strategy, products, and technical leadership
No more than one newsletter a month