Est. 2026
← All dispatches

GPT-4o Mini vs Claude Haiku vs Gemini Flash: The Budget Model Showdown

Every major AI provider now has a "budget" model. OpenAI has GPT-4o Mini. Anthropic has Claude Haiku. Google has Gemini Flash.

But which one is actually the best value? We analyzed pricing across 80 providers to find out.

The Contenders

Three models dominate the budget tier:

Model Provider Context Window Key Strength
GPT-4o Mini OpenAI 128K Speed, broad compatibility
Claude Haiku 4.5 Anthropic 200K Largest context, coding
Gemini 2.0 Flash Google 1M+ Massive context, multimodal

Direct API Pricing Comparison

Prices per 1 million tokens, direct from each provider:

Model Input Output Cached Input
Gemini 2.0 Flash $0.10 $0.40 $0.025
GPT-4o Mini $0.15 $0.60 $0.08
Claude 3 Haiku $0.25 $1.25 $0.03
Claude 3.5 Haiku $0.80 $4.00 $0.08
Claude Haiku 4.5 $1.00 $5.00 $0.10

Winner on base price: Gemini 2.0 Flash at $0.10 input.

But wait—there's more to this story.

The Gemini Flash Family (It's Confusing)

Google has seven different Flash variants. Here's the breakdown:

Model Input Output Best For
Gemini 1.5 Flash-8B $0.0375 $0.15 Absolute cheapest
Gemini 2.0 Flash-Lite $0.075 $0.30 Budget + speed
Gemini 2.0 Flash $0.10 $0.40 General purpose
Gemini 2.5 Flash-Lite $0.10 $0.40 Newer, same price
Gemini 2.5 Flash $0.30 $2.50 Latest capabilities
Gemini 3 Flash Preview $0.50 $3.00 Cutting edge

The actual cheapest option: Gemini 1.5 Flash-8B at $0.0375/1M input tokens—4x cheaper than GPT-4o Mini.

Cross-Provider Pricing

The same model costs different amounts depending on where you access it:

GPT-4o Mini Across Providers

Provider Input Output Notes
OpenAI (direct) $0.15 $0.60 Standard pricing
Azure OpenAI $0.15 $0.60 Same as direct
OpenRouter $0.15 $0.60 Same as direct
Poe $0.14 $0.54 7% cheaper
GitHub Models $0.00 $0.00 Free (rate limited)
Helicone $0.15 $0.60 Passthrough

Claude Haiku 4.5 Across Providers

Provider Input Output Notes
Anthropic (direct) $1.00 $5.00 Standard pricing
Amazon Bedrock $1.00 $5.00 Same as direct
Azure OpenAI $1.00 $5.00 Same as direct
Google Vertex $1.00 $5.00 Same as direct
Poe $0.85 $4.30 15% cheaper
AIHubMix $1.10 $5.50 10% markup
Firmware $0.00 $0.00 Free tier

Gemini 2.0 Flash Across Providers

Provider Input Output Notes
Google (direct) $0.10 $0.40 Standard pricing
Vertex AI $0.10 $0.40 Same as direct
OpenRouter ~$0.10 ~$0.40 Passthrough

The Hidden Cost: Output Tokens

Here's what most comparisons miss: output tokens cost 4-5x more than input tokens.

If your application is output-heavy (generating content, code, long responses), the output price matters more than input.

Output-heavy workload example (1M input, 2M output):

Model Input Cost Output Cost Total
Gemini 2.0 Flash $0.10 $0.80 $0.90
GPT-4o Mini $0.15 $1.20 $1.35
Claude 3 Haiku $0.25 $2.50 $2.75
Claude Haiku 4.5 $1.00 $10.00 $11.00

Gemini Flash is 12x cheaper than Claude Haiku 4.5 for output-heavy work.

Caching Changes Everything

All three providers offer input caching at steep discounts:

Model Standard Input Cached Input Savings
Gemini 2.0 Flash $0.10 $0.025 75% off
GPT-4o Mini $0.15 $0.08 47% off
Claude Haiku 4.5 $1.00 $0.10 90% off

Claude's cache discount is the steepest. If you can structure prompts to maximize cache hits, Claude Haiku 4.5's effective cost drops dramatically.

Cached input comparison:

Model Cached Input
Gemini 2.0 Flash $0.025
Claude 3 Haiku $0.03
GPT-4o Mini $0.08
Claude Haiku 4.5 $0.10

With caching, Gemini and Claude 3 Haiku are nearly identical.

Which Should You Choose?

Choose Gemini 2.0 Flash if:

  • Cost is your primary concern
  • You need massive context windows (1M+ tokens)
  • You're doing multimodal work (images, video)
  • Output generation is heavy

Choose GPT-4o Mini if:

  • You need broad ecosystem compatibility
  • You're already on OpenAI's platform
  • You need the fastest response times
  • You want free access via GitHub Models

Choose Claude Haiku if:

  • You need the best coding performance in the budget tier
  • 200K context is sufficient
  • You can leverage aggressive caching (90% discount)
  • You need Claude's safety/style for your use case

The Surprise Winner

For pure cost optimization: Gemini 1.5 Flash-8B at $0.0375 input.

For best balance of capability and cost: Gemini 2.0 Flash at $0.10 input.

For coding tasks: Claude 3.5 Haiku at $0.80 input (not the newer 4.5).

For ecosystem and speed: GPT-4o Mini at $0.15 input.

Real-World Cost Calculator

Processing 10 million tokens per month (5M in, 5M out):

Model Monthly Cost
Gemini 1.5 Flash-8B $0.94
Gemini 2.0 Flash $2.50
GPT-4o Mini $3.75
Claude 3 Haiku $7.50
Claude 3.5 Haiku $24.00
Claude Haiku 4.5 $30.00

The gap between cheapest and most expensive: 32x.


All pricing data from Subquery's real-time database covering 80 providers. Compare models →