Published 2026-04-13 · Reviewed 2026-08-27, corrections outstanding · 19 providers compared · Figures compiled 2026-04-13, last checked 2026-08-27 · 22 pricing changes tracked
LLM API pricing in April 2026: 19 providers across four categories — frontier labs, inference providers, open-source hosts, and specialized services. The biggest story: Claude Opus 4.6 pricing dropped 67% ($5/$25 per M tokens, down from $15/$75). Google restricted Gemini free tier to Flash-only models. DeepSeek V4 delivers 1M context at $0.30/M input — the cheapest long-context option. Groq and Cerebras offer genuinely free inference at thousands of tokens/second.
Key trends: Inference providers (Groq, Cerebras, OpenRouter) are commoditizing open-source model access — free tiers with no credit card required. xAI Grok 4.1 is the cheapest frontier model at $0.20/M input. The gap between frontier and open-source quality is narrowing, making the price delta harder to justify for many use cases.
This guide covers: pricing tables, provider breakdowns, free tier analysis, cheapest-per-token rankings, pricing gotchas, recent changes, and best-for-use-case recommendations — compiled by hand from vendor pricing pages.
Frontier: xAI Grok 4.1 Fast ($0.20/M input, $0.50/M output) · Open-source: Groq Llama 4 Scout ($0.11/M input) · Long context (1M): DeepSeek V4 ($0.30/M input) · Reasoning: DeepSeek R1 ($0.55/M input) vs Claude Opus 4.6 ($5/M) vs OpenAI o3 ($10/M)
Most generous: Groq (30 RPM, 500K tok/day) · GitHub Models (100+ models free) · Mistral (1B tok/month) · Best for prototyping: OpenRouter (~30 free models) · Cerebras (1M tok/day) · Completely free: LLM7.io (donor-supported) · Ollama (self-hosted)
All prices verified as of April 2026. Per-million-token pricing for flagship models. Hover rows to highlight. Click provider names for full vendor profiles.
| Provider | Free Tier | Flagship Model | Input /M | Output /M | Context |
|---|---|---|---|---|---|
| OpenAI | GPT-3.5 only | GPT-4o | $2.50/M | $10/M | 128K |
| Anthropic | Console access | Claude Opus 4.6 | $5/M | $25/M | 200K |
| Google Gemini | Flash models only | Gemini 2.5 Pro | $1.25/M | $10/M | 1M |
| Mistral AI | 2 RPM, 1B tok/mo | Mistral Large | $2/M | $6/M | 128K |
| Cohere | 1K calls/mo | Command R+ | $2.50/M | $10/M | 128K |
| xAI (Grok) | $25 signup credits | Grok 4.1 Fast | $0.20/M | $0.50/M | 128K |
| Groq | 30 RPM free | Llama 4 Scout 17B | $0.11/M | $0.18/M | 128K |
| Cerebras | 1M tok/day | Llama 3.1 70B | $0.60/M | $0.60/M | 128K |
| OpenRouter | ~30 free models | Multi-model gateway | Varies | Varies | Varies |
| NVIDIA NIM | 1K credits | Llama 3.1 70B | Credit-based | Credit-based | 128K |
| SiliconFlow | 100 req/day + $1 | DeepSeek-R1 | $0.14/M | $0.14/M | 64K |
| Hugging Face | $0.10/mo credits | 200+ models | Provider-dependent | Provider-dependent | Varies |
| Replicate | Free runs (curated) | Llama, Stable Diffusion | Per-second billing | Per-second billing | Varies |
| Baseten | $30 free credits | Custom deployments | Per-minute GPU | Per-minute GPU | Varies |
| Cloudflare Workers AI | 10K neurons/day | Llama, Mistral, SDXL | Neuron-based | Neuron-based | Varies |
| DeepSeek | 5M tokens free | DeepSeek V4 | $0.30/M | $0.50/M | 1M |
| GitHub Models | 50-150 req/day | GPT-4o, Llama, Mistral | Free (rate-limited) | Free (rate-limited) | Varies |
| LLM7.io | No published limits | 30+ models | Free | Free | Varies |
| Ollama | Free (light usage) | Llama, Mistral, Gemma | Free (self-host) | Free (self-host) | Varies |
LLM API providers fall into four categories. Each optimized for different use cases, budgets, and requirements.
The AI labs that train and serve their own flagship models. Higher prices, cutting-edge capabilities, proprietary architectures.
Limited to GPT-3.5 Turbo only, 3 requests/minute rate limit. No free trial credits for new accounts (discontinued mid-2025). Paid tiers unlock GPT-4o, GPT-4.1, o3/o4-mini reasoning models. Batch processing at 50% discount.
Key differentiator: Widest model selection; ecosystem leader with function calling, vision, and structured outputs
Limited access via console with rate limits. Opus 4.6: $5/$25 per MTok (67% below previous Opus pricing). Sonnet 4.6: $3/$15 per MTok. Haiku 4.5: $0.80/$4 per MTok. Batch API at 50% discount. Extended thinking for complex reasoning.
Key differentiator: Largest context window (200K); extended thinking; 67% Opus price drop makes frontier reasoning affordable
Free tier restricted to Flash models only (April 2026). Flash: 10 RPM, Flash-Lite: 15 RPM. 1M token context window. Pro models now require paid billing. Mandatory spend caps enforced since April 1, 2026.
Key differentiator: 1M token context window; cheapest frontier input tokens; Flash models genuinely free
Experiment tier: 2 RPM, 1 billion tokens/month. No credit card required. Access to all Mistral models including Large, Codestral, Pixtral. Le Chat consumer app included.
Key differentiator: European AI lab; strong open-weight models (Mixtral); Codestral for code generation
Trial key: 1,000 API calls/month across all endpoints (Chat, Embed, Rerank). Non-commercial use only. Access to Command R+, Rerank 3.5, Embed 4.
Key differentiator: Best-in-class RAG pipeline (Embed + Rerank + Chat); enterprise-focused; non-commercial free tier
$25 in free API credits on signup. Additional $150/month via data sharing program (opt-in, requires $5 minimum spend first). Grok 4.1 Fast: $0.20/M input, $0.50/M output — cheapest frontier model.
Key differentiator: Cheapest frontier pricing ($0.20/M input); $175/month possible in free credits via data sharing
Platforms that serve open-source models on optimized hardware. Cheaper than frontier labs, focused on speed and cost efficiency.
30 RPM, 100K-500K tokens/day depending on model. No credit card required. Custom LPU hardware delivers thousands of tokens/second. Models: Llama 4 Scout 17B, Llama 3.3 70B, Qwen3 32B, Whisper.
Key differentiator: Fastest inference speed (custom LPU hardware); generous free tier with no credit card
1M tokens/day, 10-30 requests/min. Custom wafer-scale chips for multi-thousand tokens/sec inference. Models: Llama 3.1 8B/70B, Qwen 3 235B, GPT-OSS 120B.
Key differentiator: Wafer-scale chip inference; competitive speeds with Groq; 1M tokens/day free
~30 free models including DeepSeek R1, Llama 3.3, Qwen3, Gemma 3. ~20 RPM per model. OpenAI-compatible API. Routes to cheapest provider automatically. One API key for 200+ models.
Key differentiator: Universal gateway to 200+ models; automatic provider routing; OpenAI-compatible API
~40 RPM, 1,000 free API credits. No credit card required for development. Models: Llama 3.1, Mistral, and NVIDIA models. Optimized with TensorRT-LLM for NVIDIA GPUs.
Key differentiator: NVIDIA-optimized inference; deploy on your own NVIDIA GPUs; enterprise support
100 requests/day and $1 free credits. Models: DeepSeek-R1, DeepSeek-V3, QwQ-32B, other open-source models. China-based provider.
Key differentiator: China-based with competitive pricing; strong DeepSeek model support
Platforms for deploying and serving any open-source model. Pay for compute, bring your own model.
$0.10/month free inference credits, 200+ models via Inference Providers, unlimited model hosting on Hub. Routes to multiple inference providers (AWS, GCP, etc.).
Key differentiator: Largest model hub (800K+ models); Inference Providers route to optimal backend; community ecosystem
Free runs on curated model collection without billing. No credit card required to start. Pay-per-second billing by hardware type (CPU/GPU) after free allowance.
Key differentiator: Run any open-source model; per-second GPU billing; one-click model deployment
$30 in free credits for new accounts. Basic plan: $0/month with pay-as-you-go billing after credits. Per-minute GPU/CPU billing for custom deployments, per-token for Model APIs.
Key differentiator: Deploy custom models with autoscaling; $30 free credits; optimized Truss framework
10,000 neurons/day free across text generation, image classification, translation, speech-to-text models. Runs on Cloudflare's 300+ edge locations. No cold starts.
Key differentiator: Edge inference on 300+ locations; zero cold starts; bundled with Workers ecosystem
Providers with unique positioning — free tiers, self-hosting, or niche model access.
5M free tokens for new accounts (30-day validity). DeepSeek V4: $0.30/M input, $0.50/M output (1M context). R1 reasoning: $0.55/M input, $2.19/M output. Cache hits 90% cheaper. Off-peak discounts up to 75% off. China-based.
Key differentiator: Cheapest 1M-context model; cache-hit discounts (90% off); R1 reasoning model competitive with frontier
10-15 RPM, 50-150 requests/day depending on model tier. OpenAI-compatible API endpoint. 100+ models including GPT-4o, Llama, Mistral, Phi. Free for GitHub users.
Key differentiator: Free access to 100+ models for GitHub users; OpenAI-compatible; great for prototyping
No published rate limits on free tier. Supported by donors. 30+ models including DeepSeek R1, Qwen2.5 Coder, text, image, and speech-to-text models. UK-based.
Key differentiator: Completely free donor-supported inference; 30+ models; UK-based
Free tier for light usage with 1 concurrent model. Ollama Cloud provides hosted API. Local installation runs models on your hardware at zero cost. Models: Llama, Mistral, Gemma, and hundreds more.
Key differentiator: Local-first with optional cloud; run models on your own hardware; largest model library
Free tiers range from genuinely production-viable (Groq, GitHub Models) to token giveaways that run out in hours. Here's the honest breakdown.
| Provider | Free Tier Type | Rate Limit | Context Window | When You Hit the Wall |
|---|---|---|---|---|
| OpenAI | Limited free tier | 3 RPM (free) | 128K | Days of moderate use |
| Anthropic | Pay-as-you-go | Rate-limited (free) | 200K | Immediately (pay per token) |
| Google Gemini | Limited free tier | 10-15 RPM (free) | 1M | Days of moderate use |
| Mistral AI | Generous free tier | 2 RPM (free) | 128K | Weeks to months of real usage |
| Cohere | Trial only | 1K calls/month | 128K | 1K calls/month limit |
| xAI (Grok) | Free credits | Standard | 128K | When credits run out |
| Groq | Generous free tier | 30 RPM, 100K-500K tok/day | 128K | Weeks to months of real usage |
| Cerebras | Generous free tier | 10-30 RPM, 1M tok/day | 128K | Weeks to months of real usage |
| OpenRouter | Generous free tier | ~20 RPM per model | Varies | Weeks to months of real usage |
| NVIDIA NIM | Free credits | ~40 RPM | 128K | When credits run out |
| SiliconFlow | Limited free tier | 100 req/day | 64K | Days of moderate use |
| Hugging Face | Limited free tier | Varies | Varies | Days of moderate use |
| Replicate | Generous free tier | Standard | Varies | Weeks to months of real usage |
| Baseten | Free credits | Standard | Varies | When credits run out |
| Cloudflare Workers AI | Generous free tier | 10K neurons/day | Varies | Weeks to months of real usage |
| DeepSeek | Free credits | Standard | 1M | When credits run out |
| GitHub Models | Generous free tier | 10-15 RPM, 50-150 req/day | Varies | Weeks to months of real usage |
| LLM7.io | Generous free tier | No published limits | Varies | Weeks to months of real usage |
| Ollama | Generous free tier | 1 concurrent model | Varies | Weeks to months of real usage |
Token pricing is rarely the full story. These are the costs and limits that surprise developers.
OpenAI o3 and o4-mini models generate internal "reasoning tokens" that count toward output pricing but aren't visible in the response. A simple query can generate 10x more reasoning tokens than output tokens. Monitor usage carefully — your bill reflects total tokens, not just visible output.
Claude's extended thinking generates visible thinking tokens billed at output rates ($25/M for Opus 4.6). A complex reasoning task can produce 10K+ thinking tokens before the actual answer. Budget for 2–5x the output tokens you'd expect from a non-thinking request.
Google restricted the Gemini free tier to Flash and Flash-Lite models only. Pro models (Gemini 2.5 Pro) now require a paid billing account with mandatory spend caps. If you were using Pro models for free — that's over.
A 128K context window doesn't mean the model performs well at 128K tokens. Quality degrades in the middle of long contexts ("lost in the middle" problem). For retrieval-heavy tasks, expect effective context of 30–50% of the advertised window. DeepSeek and Gemini's 1M windows are more susceptible to this.
OpenAI and Anthropic increase rate limits based on cumulative spend, not the plan you're on. A new account with $100 in credits still starts at Tier 1 limits. Groq's free tier limits are per-model, so switching models resets your quota.
LLM API pricing is the most volatile in the developer tools space. Here are the changes we've tracked. See full change timeline for all tracked changes.
| Date | Vendor | Change | Impact |
|---|---|---|---|
| Sep 1, 2026 | Cloudflare Workers AI | The pricing has been restructured. While a free tier of 10,000 Neurons per day still exists, the pricing is now more granular and based on per-model unit pricing, billed in Neurons at $0.011 / 1,000 Neurons. Some models now require a paid plan or AI Gateway credits. | HIGH |
| Aug 26, 2026 | OpenAI | Assistants API deprecated, full shutdown August 26, 2026. Developers must migrate to Responses API + Conversations API | HIGH |
| Jun 1, 2026 | Google Gemini API | Gemini 2.0 Flash and 2.0 Flash-Lite will be deprecated June 1, 2026. Developers must migrate to Gemini 2.5 Flash or 3.x Flash models. Google is consolidating the model lineup — 2.0 generation reaching end of life as 2.5 and 3.x become production-ready. | HIGH |
| May 12, 2026 | OpenAI | DALL-E 2 and DALL-E 3 API access discontinued. Developers must migrate to gpt-image-1 (different pricing model, quality tiers changed from standard/hd to low/medium/high) or switch to free alternatives like Pollinations.AI or Lumenfall.ai | HIGH |
| May 7, 2026 | OpenAI | Realtime API beta endpoints deprecated. Developers must remove OpenAI-Beta header, use new client_secrets endpoint, specify session_type, and update event names. GA Realtime API is the direct replacement. | HIGH |
| Apr 17, 2026 | DeepSeek | DeepSeek V3.2 replaces V4 branding. Pricing dropped: chat model $0.30→$0.28/M input, $0.50→$0.42/M output. Reasoner model pricing unified with chat at $0.28/$0.42 (was $0.55/$2.19). Free token welcome package appears removed. | MEDIUM |
| Apr 13, 2026 | xAI | $25/month free API credits no longer offered — Grok API paid-only | LOW |
| Apr 12, 2026 | Gem | Removed: source page no longer accessible or deal program discontinued | LOW |
| Apr 12, 2026 | Mode | Removed: source page no longer accessible or deal program discontinued | LOW |
| Apr 8, 2026 | Google Gemini API | Gemini API free tier restricted to Flash and Flash-Lite models only (April 2026). Gemini 2.5 Pro and other Pro models now require a paid billing account. Previously free users could access Pro models with rate limits. Combined with April 1 spend cap enforcement, this significantly narrows what is available at $0. | HIGH |
| Apr 3, 2026 | OpenAI Codex | Switched from per-seat subscription to pay-as-you-go token-based pricing. Teams can add Codex-only seats billed on token consumption with no rate limits. ChatGPT Business price cut from $25 to $20/month (annual). $100 credit per new Codex team member (up to $500/team, limited time) | MEDIUM |
| Apr 1, 2026 | Google Gemini API | Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. | MEDIUM |
| Mar 30, 2026 | Google Developer Program annual subscription ending March 30. Replaced by Google AI Pro ($10/mo) and AI Ultra ($100/mo) tiers with bundled cloud credits and AI model access. | HIGH | |
| Mar 19, 2026 | xAI (Grok) | Grok Imagine free image/video generation locked behind SuperGrok subscription ($30/mo). Over 1B generations in first month overwhelmed capacity | MEDIUM |
| Mar 13, 2026 | Anthropic Claude | Temporary usage promotion: double five-hour usage limits during off-peak hours, March 13-27, 2026. Applies to Claude Pro and Team subscribers | LOW |
| Mar 3, 2026 | Google Gemini 2.0 Flash | Gemini 2.0 Flash and 2.0 Flash-Lite models deprecated, scheduled for retirement. Free tier continues with Gemini 2.5 models. Developers on 2.0 must migrate | MEDIUM |
| Feb 9, 2026 | OpenAI | Ads launched in ChatGPT Free and Go ($8/mo) tiers. Sponsored units from major brands appear below responses on first prompt. $60 CPM, $200K minimum ad commitment | HIGH |
| Feb 5, 2026 | Anthropic | Claude Opus 4.6 API pricing at $5/$25 per MTok (input/output) — 67% below previous Opus 4/4.1 pricing of $15/$75. Frontier AI model at mid-tier prices | HIGH |
| Feb 4, 2026 | Cloudflare | Queues added to Workers free plan — message queuing now free | MEDIUM |
| Jan 27, 2026 | Google Developer Program Premium merged into Google One AI Pro ($19.99/mo, includes $10 Cloud credits) and AI Ultra ($100 Cloud credits). Developer benefits now bundled into consumer AI subscriptions | MEDIUM |
Groq (free, fast, no credit card) or OpenRouter (~30 free models, try different providers). GitHub Models for accessing GPT-4o free with a GitHub account.
OpenAI GPT-4o ($2.50/M in, $10/M out) — widest ecosystem, function calling, structured outputs. Claude Sonnet 4.6 ($3/$15/M) for nuanced conversation and longer context.
Claude Opus 4.6 ($5/$25/M) with extended thinking — 67% cheaper than before. DeepSeek R1 ($0.55/$2.19/M) for budget reasoning. OpenAI o3 for math/science benchmarks.
xAI Grok 4.1 Fast ($0.20/$0.50/M) — cheapest frontier model. DeepSeek V4 ($0.30/$0.50/M) with 90% cache-hit discounts. OpenAI/Anthropic batch APIs at 50% off for async workloads.
DeepSeek V4 (1M context, $0.30/M) or Google Gemini 2.5 Pro (1M context, $1.25/M). Claude (200K) for highest quality within context window.
Ollama (free, run locally) for development. Replicate or Baseten for hosted open-source models with dedicated infrastructure. Cloudflare Workers AI for edge inference.
Compare LLM API pricing, search free tiers, and track pricing changes — all from your AI coding assistant.
claude mcp add agentdeals -- npx -y agentdealsWorks with Claude Desktop, Cursor, Cline, Windsurf → Full setup guide