LLM API Pricing — The 2026 Comparison

Published 2026-04-13 · Reviewed 2026-08-27, corrections outstanding · 19 providers compared · Figures compiled 2026-04-13, last checked 2026-08-27 · 22 pricing changes tracked

19
Providers Compared
9
Generous Free Tier
9
Credits / Limited / Trial
4
Provider Categories

LLM API pricing in April 2026: 19 providers across four categories — frontier labs, inference providers, open-source hosts, and specialized services. The biggest story: Claude Opus 4.6 pricing dropped 67% ($5/$25 per M tokens, down from $15/$75). Google restricted Gemini free tier to Flash-only models. DeepSeek V4 delivers 1M context at $0.30/M input — the cheapest long-context option. Groq and Cerebras offer genuinely free inference at thousands of tokens/second.

Key trends: Inference providers (Groq, Cerebras, OpenRouter) are commoditizing open-source model access — free tiers with no credit card required. xAI Grok 4.1 is the cheapest frontier model at $0.20/M input. The gap between frontier and open-source quality is narrowing, making the price delta harder to justify for many use cases.

This guide covers: pricing tables, provider breakdowns, free tier analysis, cheapest-per-token rankings, pricing gotchas, recent changes, and best-for-use-case recommendations — compiled by hand from vendor pricing pages.

Cheapest per Million Tokens

Frontier: xAI Grok 4.1 Fast ($0.20/M input, $0.50/M output) · Open-source: Groq Llama 4 Scout ($0.11/M input) · Long context (1M): DeepSeek V4 ($0.30/M input) · Reasoning: DeepSeek R1 ($0.55/M input) vs Claude Opus 4.6 ($5/M) vs OpenAI o3 ($10/M)

Best Free Tiers

Most generous: Groq (30 RPM, 500K tok/day) · GitHub Models (100+ models free) · Mistral (1B tok/month) · Best for prototyping: OpenRouter (~30 free models) · Cerebras (1M tok/day) · Completely free: LLM7.io (donor-supported) · Ollama (self-hosted)

Jump to section

  1. Pricing Comparison Table
  2. Provider Breakdown (Frontier, Inference, Open-Source, Specialized)
  3. What You Actually Get for Free
  4. Pricing Gotchas
  5. Recent Pricing Changes
  6. Best-for-Use-Case Recommendations
  7. FAQ

Pricing Comparison Table

All prices verified as of April 2026. Per-million-token pricing for flagship models. Hover rows to highlight. Click provider names for full vendor profiles.

Provider Free Tier Flagship Model Input /M Output /M Context
OpenAIGPT-3.5 onlyGPT-4o$2.50/M$10/M128K
AnthropicConsole accessClaude Opus 4.6$5/M$25/M200K
Google GeminiFlash models onlyGemini 2.5 Pro$1.25/M$10/M1M
Mistral AI2 RPM, 1B tok/moMistral Large$2/M$6/M128K
Cohere1K calls/moCommand R+$2.50/M$10/M128K
xAI (Grok)$25 signup creditsGrok 4.1 Fast$0.20/M$0.50/M128K
Groq30 RPM freeLlama 4 Scout 17B$0.11/M$0.18/M128K
Cerebras1M tok/dayLlama 3.1 70B$0.60/M$0.60/M128K
OpenRouter~30 free modelsMulti-model gatewayVariesVariesVaries
NVIDIA NIM1K creditsLlama 3.1 70BCredit-basedCredit-based128K
SiliconFlow100 req/day + $1DeepSeek-R1$0.14/M$0.14/M64K
Hugging Face$0.10/mo credits200+ modelsProvider-dependentProvider-dependentVaries
ReplicateFree runs (curated)Llama, Stable DiffusionPer-second billingPer-second billingVaries
Baseten$30 free creditsCustom deploymentsPer-minute GPUPer-minute GPUVaries
Cloudflare Workers AI10K neurons/dayLlama, Mistral, SDXLNeuron-basedNeuron-basedVaries
DeepSeek5M tokens freeDeepSeek V4$0.30/M$0.50/M1M
GitHub Models50-150 req/dayGPT-4o, Llama, MistralFree (rate-limited)Free (rate-limited)Varies
LLM7.ioNo published limits30+ modelsFreeFreeVaries
OllamaFree (light usage)Llama, Mistral, GemmaFree (self-host)Free (self-host)Varies
The price floor: xAI Grok 4.1 Fast at $0.20/M input is the cheapest frontier model. For open-source, Groq and SiliconFlow serve Llama and DeepSeek models at $0.10–0.15/M. DeepSeek V4 offers 1M context at $0.30/M input — 4x cheaper than Gemini 2.5 Pro for long-context work. The batch APIs from OpenAI and Anthropic offer 50% discounts for non-real-time workloads.

Provider Breakdown

LLM API providers fall into four categories. Each optimized for different use cases, budgets, and requirements.

Frontier Labs

The AI labs that train and serve their own flagship models. Higher prices, cutting-edge capabilities, proprietary architectures.

OpenAI Limited free tier

Limited to GPT-3.5 Turbo only, 3 requests/minute rate limit. No free trial credits for new accounts (discontinued mid-2025). Paid tiers unlock GPT-4o, GPT-4.1, o3/o4-mini reasoning models. Batch processing at 50% discount.

Key differentiator: Widest model selection; ecosystem leader with function calling, vision, and structured outputs

Anthropic Pay-as-you-go

Limited access via console with rate limits. Opus 4.6: $5/$25 per MTok (67% below previous Opus pricing). Sonnet 4.6: $3/$15 per MTok. Haiku 4.5: $0.80/$4 per MTok. Batch API at 50% discount. Extended thinking for complex reasoning.

Key differentiator: Largest context window (200K); extended thinking; 67% Opus price drop makes frontier reasoning affordable

Google Gemini Limited free tier

Free tier restricted to Flash models only (April 2026). Flash: 10 RPM, Flash-Lite: 15 RPM. 1M token context window. Pro models now require paid billing. Mandatory spend caps enforced since April 1, 2026.

Key differentiator: 1M token context window; cheapest frontier input tokens; Flash models genuinely free

Mistral AI Generous free tier

Experiment tier: 2 RPM, 1 billion tokens/month. No credit card required. Access to all Mistral models including Large, Codestral, Pixtral. Le Chat consumer app included.

Key differentiator: European AI lab; strong open-weight models (Mixtral); Codestral for code generation

Cohere Trial only

Trial key: 1,000 API calls/month across all endpoints (Chat, Embed, Rerank). Non-commercial use only. Access to Command R+, Rerank 3.5, Embed 4.

Key differentiator: Best-in-class RAG pipeline (Embed + Rerank + Chat); enterprise-focused; non-commercial free tier

xAI (Grok) Free credits

$25 in free API credits on signup. Additional $150/month via data sharing program (opt-in, requires $5 minimum spend first). Grok 4.1 Fast: $0.20/M input, $0.50/M output — cheapest frontier model.

Key differentiator: Cheapest frontier pricing ($0.20/M input); $175/month possible in free credits via data sharing

Inference Providers

Platforms that serve open-source models on optimized hardware. Cheaper than frontier labs, focused on speed and cost efficiency.

Groq Generous free tier

30 RPM, 100K-500K tokens/day depending on model. No credit card required. Custom LPU hardware delivers thousands of tokens/second. Models: Llama 4 Scout 17B, Llama 3.3 70B, Qwen3 32B, Whisper.

Key differentiator: Fastest inference speed (custom LPU hardware); generous free tier with no credit card

Cerebras Generous free tier

1M tokens/day, 10-30 requests/min. Custom wafer-scale chips for multi-thousand tokens/sec inference. Models: Llama 3.1 8B/70B, Qwen 3 235B, GPT-OSS 120B.

Key differentiator: Wafer-scale chip inference; competitive speeds with Groq; 1M tokens/day free

OpenRouter Generous free tier

~30 free models including DeepSeek R1, Llama 3.3, Qwen3, Gemma 3. ~20 RPM per model. OpenAI-compatible API. Routes to cheapest provider automatically. One API key for 200+ models.

Key differentiator: Universal gateway to 200+ models; automatic provider routing; OpenAI-compatible API

NVIDIA NIM Free credits

~40 RPM, 1,000 free API credits. No credit card required for development. Models: Llama 3.1, Mistral, and NVIDIA models. Optimized with TensorRT-LLM for NVIDIA GPUs.

Key differentiator: NVIDIA-optimized inference; deploy on your own NVIDIA GPUs; enterprise support

SiliconFlow Limited free tier

100 requests/day and $1 free credits. Models: DeepSeek-R1, DeepSeek-V3, QwQ-32B, other open-source models. China-based provider.

Key differentiator: China-based with competitive pricing; strong DeepSeek model support

Open-Source Model Hosts

Platforms for deploying and serving any open-source model. Pay for compute, bring your own model.

Hugging Face Limited free tier

$0.10/month free inference credits, 200+ models via Inference Providers, unlimited model hosting on Hub. Routes to multiple inference providers (AWS, GCP, etc.).

Key differentiator: Largest model hub (800K+ models); Inference Providers route to optimal backend; community ecosystem

Replicate Generous free tier

Free runs on curated model collection without billing. No credit card required to start. Pay-per-second billing by hardware type (CPU/GPU) after free allowance.

Key differentiator: Run any open-source model; per-second GPU billing; one-click model deployment

Baseten Free credits

$30 in free credits for new accounts. Basic plan: $0/month with pay-as-you-go billing after credits. Per-minute GPU/CPU billing for custom deployments, per-token for Model APIs.

Key differentiator: Deploy custom models with autoscaling; $30 free credits; optimized Truss framework

Cloudflare Workers AI Generous free tier

10,000 neurons/day free across text generation, image classification, translation, speech-to-text models. Runs on Cloudflare's 300+ edge locations. No cold starts.

Key differentiator: Edge inference on 300+ locations; zero cold starts; bundled with Workers ecosystem

Specialized & Free Providers

Providers with unique positioning — free tiers, self-hosting, or niche model access.

DeepSeek Free credits

5M free tokens for new accounts (30-day validity). DeepSeek V4: $0.30/M input, $0.50/M output (1M context). R1 reasoning: $0.55/M input, $2.19/M output. Cache hits 90% cheaper. Off-peak discounts up to 75% off. China-based.

Key differentiator: Cheapest 1M-context model; cache-hit discounts (90% off); R1 reasoning model competitive with frontier

GitHub Models Generous free tier

10-15 RPM, 50-150 requests/day depending on model tier. OpenAI-compatible API endpoint. 100+ models including GPT-4o, Llama, Mistral, Phi. Free for GitHub users.

Key differentiator: Free access to 100+ models for GitHub users; OpenAI-compatible; great for prototyping

LLM7.io Generous free tier

No published rate limits on free tier. Supported by donors. 30+ models including DeepSeek R1, Qwen2.5 Coder, text, image, and speech-to-text models. UK-based.

Key differentiator: Completely free donor-supported inference; 30+ models; UK-based

Ollama Generous free tier

Free tier for light usage with 1 concurrent model. Ollama Cloud provides hosted API. Local installation runs models on your hardware at zero cost. Models: Llama, Mistral, Gemma, and hundreds more.

Key differentiator: Local-first with optional cloud; run models on your own hardware; largest model library

What You Actually Get for Free

Free tiers range from genuinely production-viable (Groq, GitHub Models) to token giveaways that run out in hours. Here's the honest breakdown.

Provider Free Tier Type Rate Limit Context Window When You Hit the Wall
OpenAILimited free tier3 RPM (free)128KDays of moderate use
AnthropicPay-as-you-goRate-limited (free)200KImmediately (pay per token)
Google GeminiLimited free tier10-15 RPM (free)1MDays of moderate use
Mistral AIGenerous free tier2 RPM (free)128KWeeks to months of real usage
CohereTrial only1K calls/month128K1K calls/month limit
xAI (Grok)Free creditsStandard128KWhen credits run out
GroqGenerous free tier30 RPM, 100K-500K tok/day128KWeeks to months of real usage
CerebrasGenerous free tier10-30 RPM, 1M tok/day128KWeeks to months of real usage
OpenRouterGenerous free tier~20 RPM per modelVariesWeeks to months of real usage
NVIDIA NIMFree credits~40 RPM128KWhen credits run out
SiliconFlowLimited free tier100 req/day64KDays of moderate use
Hugging FaceLimited free tierVariesVariesDays of moderate use
ReplicateGenerous free tierStandardVariesWeeks to months of real usage
BasetenFree creditsStandardVariesWhen credits run out
Cloudflare Workers AIGenerous free tier10K neurons/dayVariesWeeks to months of real usage
DeepSeekFree creditsStandard1MWhen credits run out
GitHub ModelsGenerous free tier10-15 RPM, 50-150 req/dayVariesWeeks to months of real usage
LLM7.ioGenerous free tierNo published limitsVariesWeeks to months of real usage
OllamaGenerous free tier1 concurrent modelVariesWeeks to months of real usage

Pricing Gotchas

Token pricing is rarely the full story. These are the costs and limits that surprise developers.

OpenAI Reasoning Tokens: 5–10x Cost Multiplier

OpenAI o3 and o4-mini models generate internal "reasoning tokens" that count toward output pricing but aren't visible in the response. A simple query can generate 10x more reasoning tokens than output tokens. Monitor usage carefully — your bill reflects total tokens, not just visible output.

Anthropic Extended Thinking: Output Tokens Add Up

Claude's extended thinking generates visible thinking tokens billed at output rates ($25/M for Opus 4.6). A complex reasoning task can produce 10K+ thinking tokens before the actual answer. Budget for 2–5x the output tokens you'd expect from a non-thinking request.

Gemini Free Tier: Flash Only Since April 2026

Google restricted the Gemini free tier to Flash and Flash-Lite models only. Pro models (Gemini 2.5 Pro) now require a paid billing account with mandatory spend caps. If you were using Pro models for free — that's over.

Context Window ≠ Effective Context

A 128K context window doesn't mean the model performs well at 128K tokens. Quality degrades in the middle of long contexts ("lost in the middle" problem). For retrieval-heavy tasks, expect effective context of 30–50% of the advertised window. DeepSeek and Gemini's 1M windows are more susceptible to this.

Rate Limits Scale with Spend, Not Plan

OpenAI and Anthropic increase rate limits based on cumulative spend, not the plan you're on. A new account with $100 in credits still starts at Tier 1 limits. Groq's free tier limits are per-model, so switching models resets your quota.

Recent Pricing Changes

LLM API pricing is the most volatile in the developer tools space. Here are the changes we've tracked. See full change timeline for all tracked changes.

Date Vendor Change Impact
Sep 1, 2026Cloudflare Workers AIThe pricing has been restructured. While a free tier of 10,000 Neurons per day still exists, the pricing is now more granular and based on per-model unit pricing, billed in Neurons at $0.011 / 1,000 Neurons. Some models now require a paid plan or AI Gateway credits.HIGH
Aug 26, 2026OpenAIAssistants API deprecated, full shutdown August 26, 2026. Developers must migrate to Responses API + Conversations APIHIGH
Jun 1, 2026Google Gemini APIGemini 2.0 Flash and 2.0 Flash-Lite will be deprecated June 1, 2026. Developers must migrate to Gemini 2.5 Flash or 3.x Flash models. Google is consolidating the model lineup — 2.0 generation reaching end of life as 2.5 and 3.x become production-ready.HIGH
May 12, 2026OpenAIDALL-E 2 and DALL-E 3 API access discontinued. Developers must migrate to gpt-image-1 (different pricing model, quality tiers changed from standard/hd to low/medium/high) or switch to free alternatives like Pollinations.AI or Lumenfall.aiHIGH
May 7, 2026OpenAIRealtime API beta endpoints deprecated. Developers must remove OpenAI-Beta header, use new client_secrets endpoint, specify session_type, and update event names. GA Realtime API is the direct replacement.HIGH
Apr 17, 2026DeepSeekDeepSeek V3.2 replaces V4 branding. Pricing dropped: chat model $0.30→$0.28/M input, $0.50→$0.42/M output. Reasoner model pricing unified with chat at $0.28/$0.42 (was $0.55/$2.19). Free token welcome package appears removed.MEDIUM
Apr 13, 2026xAI$25/month free API credits no longer offered — Grok API paid-onlyLOW
Apr 12, 2026GemRemoved: source page no longer accessible or deal program discontinuedLOW
Apr 12, 2026ModeRemoved: source page no longer accessible or deal program discontinuedLOW
Apr 8, 2026Google Gemini APIGemini API free tier restricted to Flash and Flash-Lite models only (April 2026). Gemini 2.5 Pro and other Pro models now require a paid billing account. Previously free users could access Pro models with rate limits. Combined with April 1 spend cap enforcement, this significantly narrows what is available at $0.HIGH
Apr 3, 2026OpenAI CodexSwitched from per-seat subscription to pay-as-you-go token-based pricing. Teams can add Codex-only seats billed on token consumption with no rate limits. ChatGPT Business price cut from $25 to $20/month (annual). $100 credit per new Codex team member (up to $500/team, limited time)MEDIUM
Apr 1, 2026Google Gemini APIBilling-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing.MEDIUM
Mar 30, 2026GoogleGoogle Developer Program annual subscription ending March 30. Replaced by Google AI Pro ($10/mo) and AI Ultra ($100/mo) tiers with bundled cloud credits and AI model access.HIGH
Mar 19, 2026xAI (Grok)Grok Imagine free image/video generation locked behind SuperGrok subscription ($30/mo). Over 1B generations in first month overwhelmed capacityMEDIUM
Mar 13, 2026Anthropic ClaudeTemporary usage promotion: double five-hour usage limits during off-peak hours, March 13-27, 2026. Applies to Claude Pro and Team subscribersLOW
Mar 3, 2026Google Gemini 2.0 FlashGemini 2.0 Flash and 2.0 Flash-Lite models deprecated, scheduled for retirement. Free tier continues with Gemini 2.5 models. Developers on 2.0 must migrateMEDIUM
Feb 9, 2026OpenAIAds launched in ChatGPT Free and Go ($8/mo) tiers. Sponsored units from major brands appear below responses on first prompt. $60 CPM, $200K minimum ad commitmentHIGH
Feb 5, 2026AnthropicClaude Opus 4.6 API pricing at $5/$25 per MTok (input/output) — 67% below previous Opus 4/4.1 pricing of $15/$75. Frontier AI model at mid-tier pricesHIGH
Feb 4, 2026CloudflareQueues added to Workers free plan — message queuing now freeMEDIUM
Jan 27, 2026GoogleGoogle Developer Program Premium merged into Google One AI Pro ($19.99/mo, includes $10 Cloud credits) and AI Ultra ($100 Cloud credits). Developer benefits now bundled into consumer AI subscriptionsMEDIUM
The trend: Frontier model pricing is in freefall. Anthropic dropped Opus pricing 67% in 2026. Google is aggressively undercutting on input tokens ($1.25/M for Gemini Pro). Open-source inference is approaching zero — Groq and Cerebras give away millions of tokens daily. The implication: if you're paying more than $5/M input tokens, you should evaluate whether a cheaper model handles your use case.

Best-for-Use-Case Recommendations

Pick the Right LLM API

Best for prototyping

Groq (free, fast, no credit card) or OpenRouter (~30 free models, try different providers). GitHub Models for accessing GPT-4o free with a GitHub account.

Best for production chat / assistants

OpenAI GPT-4o ($2.50/M in, $10/M out) — widest ecosystem, function calling, structured outputs. Claude Sonnet 4.6 ($3/$15/M) for nuanced conversation and longer context.

Best for complex reasoning

Claude Opus 4.6 ($5/$25/M) with extended thinking — 67% cheaper than before. DeepSeek R1 ($0.55/$2.19/M) for budget reasoning. OpenAI o3 for math/science benchmarks.

Best for high-volume / cost-sensitive

xAI Grok 4.1 Fast ($0.20/$0.50/M) — cheapest frontier model. DeepSeek V4 ($0.30/$0.50/M) with 90% cache-hit discounts. OpenAI/Anthropic batch APIs at 50% off for async workloads.

Best for long-context (100K+ tokens)

DeepSeek V4 (1M context, $0.30/M) or Google Gemini 2.5 Pro (1M context, $1.25/M). Claude (200K) for highest quality within context window.

Best for self-hosting / privacy

Ollama (free, run locally) for development. Replicate or Baseten for hosted open-source models with dedicated infrastructure. Cloudflare Workers AI for edge inference.

Frequently Asked Questions

Which LLM API has the best free tier in 2026?
Groq offers the most generous free tier: 30 RPM with 100K-500K tokens/day, no credit card required, with fast LPU-accelerated inference. GitHub Models gives free access to 100+ models (GPT-4o, Llama, Mistral) for GitHub users. OpenRouter provides ~30 free open-source models. For frontier models specifically, Mistral's Experiment tier (1B tokens/month, 2 RPM) is the most generous.
How much does GPT-4o cost per token?
GPT-4o costs $2.50 per million input tokens and $10 per million output tokens. For reference, 1 million tokens is roughly 750,000 words. The batch API offers 50% discount ($1.25/$5 per M tokens). GPT-4o-mini is significantly cheaper at $0.15/$0.60 per M tokens.
How much does Claude cost per token?
Claude Opus 4.6 costs $5/M input and $25/M output tokens — a 67% price drop from previous Opus pricing ($15/$75). Sonnet 4.6 is $3/$15 per M tokens. Haiku 4.5 is the budget option at $0.80/$4 per M tokens. The Batch API offers 50% discount on all models.
What is the cheapest LLM API for production use?
For frontier-quality models: xAI Grok 4.1 Fast at $0.20/M input, $0.50/M output. For open-source models: DeepSeek V4 at $0.30/M input, $0.50/M output with cache-hit discounts up to 90%. Groq and Cerebras offer free tiers that can handle moderate production traffic. Google Gemini Flash models are free with rate limits.
Should I use a frontier lab API or an inference provider?
Use frontier lab APIs (OpenAI, Anthropic, Google) when you need their proprietary models (GPT-4o, Claude, Gemini Pro) or specific features (function calling, vision, extended thinking). Use inference providers (Groq, Cerebras, OpenRouter) when running open-source models — they're 5-10x cheaper and often faster. Many apps work well with Llama 3.3 70B or DeepSeek R1 at a fraction of frontier pricing.

Data Source & Methodology

Powered by AgentDeals. The tables on this page were compiled by hand from official vendor pricing pages and have not been re-checked since. Pricing changes are tracked via our deal changes timeline (466 total changes tracked). The pricing changes we track are updated continuously; the tables above are not.

Query this data programmatically via /api/llm-pricing (JSON), our MCP tools, or REST API — search for LLM providers, compare pricing, or track changes from your AI coding assistant.

Get this data in your AI editor

Compare LLM API pricing, search free tiers, and track pricing changes — all from your AI coding assistant.

claude mcp add agentdeals -- npx -y agentdeals

Related Guides

Explore all 1,580 developer tool deals → Browse the full index or connect via MCP