Which LLM APIs have a genuinely free tier or free credits in 2026, and what tokens cost once you exceed it. OpenAI, Anthropic, Google Gemini, Mistral, Groq, DeepSeek, Cerebras, OpenRouter, Cohere and xAI compared — free tier limits, rate limits, context windows and per-token pricing.
Published 2026-04-13 · Reviewed 2026-08-27, corrections outstanding · 19 providers compared · Figures compiled 2026-04-13, last checked 2026-08-27 · 30 pricing changes tracked
LLM API pricing, frontier rows read 2026-09-05: 19 providers across four categories — frontier labs, inference providers, open-source hosts, and specialized services. OpenAI and Anthropic price their top model identically: GPT-6 Astra and Claude Fable 5.1 are both $10/$50 per M tokens. Gemini 3.8 Flash is $0.75/$3.75 through December 31, 2026 and $1.50/$7.50 after it. Mistral Medium 3.5 is $1.50/$7.50. Google's Gemini free tier keeps Flash, Flash-Lite and 2.5 Pro, with 3.1 Pro paid-only. DeepSeek V4 delivers 1M context at $0.30/M input — the cheapest long-context option. Groq and Cerebras offer genuinely free inference at thousands of tokens/second.
Key trends: Inference providers (Groq, Cerebras, OpenRouter) are commoditizing open-source model access — free tiers with no credit card required. xAI Grok 4.1 is the cheapest frontier model at $0.20/M input. The gap between frontier and open-source quality is narrowing, making the price delta harder to justify for many use cases.
This guide covers: pricing tables, provider breakdowns, free tier analysis, cheapest-per-token rankings, pricing gotchas, recent changes, and best-for-use-case recommendations — compiled by hand from vendor pricing pages.
Frontier: xAI Grok 4.1 Fast ($0.20/M input, $0.50/M output) · Open-source: Groq Llama 4 Scout ($0.11/M input) · Long context (1M): DeepSeek V4 ($0.30/M input) · Reasoning: DeepSeek R1 ($0.55/M input) vs Claude Opus 5 ($5/M) vs OpenAI o3 ($2/M)
Most generous: Groq (30 RPM, 500K tok/day) · GitHub Models (100+ models free) · Mistral (1B tok/month) · Best for prototyping: OpenRouter (~30 free models) · Cerebras (1M tok/day) · Completely free: LLM7.io (donor-supported) · Ollama (self-hosted)
Frontier prices read from each vendor's own pricing page on 2026-09-05: OpenAI (developers.openai.com/api/docs/pricing), Anthropic (platform.claude.com/docs/en/about-claude/pricing), Google Gemini (ai.google.dev/gemini-api/docs/pricing), Mistral AI (docs.mistral.ai/inference/pricing). The other 15 rows carry no read date. Per-million-token pricing for flagship models. Hover rows to highlight. Click provider names for full vendor profiles.
| Provider | Free Tier | Flagship Model | Input /M | Output /M | Context |
|---|---|---|---|---|---|
| OpenAI | GPT-3.5 only | GPT-6 Astra | $10/M | $50/M | 1M |
| Anthropic | Console access | Claude Fable 5.1 | $10/M | $50/M | 1M |
| Google Gemini | Flash, Flash-Lite, 2.5 Pro | Gemini 3.8 Flash | $0.75/M | $3.75/M | 1M |
| Mistral AI | 2 RPM, 1B tok/mo | Mistral Medium 3.5 | $1.50/M | $7.50/M | 256K |
| Cohere | 1K calls/mo | Command R+ | $2.50/M | $10/M | 128K |
| xAI (Grok) | $25 signup credits | Grok 4.1 Fast | $0.20/M | $0.50/M | 128K |
| Groq | 30 RPM free | Llama 4 Scout 17B | $0.11/M | $0.18/M | 128K |
| Cerebras | 1M tok/day | Llama 3.1 70B | $0.60/M | $0.60/M | 128K |
| OpenRouter | ~30 free models | Multi-model gateway | Varies | Varies | Varies |
| NVIDIA NIM | 1K credits | Llama 3.1 70B | Credit-based | Credit-based | 128K |
| SiliconFlow | 100 req/day + $1 | DeepSeek-R1 | $0.14/M | $0.14/M | 64K |
| Hugging Face | $0.10/mo credits | 200+ models | Provider-dependent | Provider-dependent | Varies |
| Replicate | Free runs (curated) | Llama, Stable Diffusion | Per-second billing | Per-second billing | Varies |
| Baseten | $30 free credits | Custom deployments | Per-minute GPU | Per-minute GPU | Varies |
| Cloudflare Workers AI | 10K neurons/day | Llama, Mistral, SDXL | Neuron-based | Neuron-based | Varies |
| DeepSeek | 5M tokens free | DeepSeek V4 | $0.30/M | $0.50/M | 1M |
| GitHub Models | 50-150 req/day | GPT-4o, Llama, Mistral | Free (rate-limited) | Free (rate-limited) | Varies |
| LLM7.io | No published limits | 30+ models | Free | Free | Varies |
| Ollama | Free (light usage) | Llama, Mistral, Gemma | Free (self-host) | Free (self-host) | Varies |
LLM API providers fall into four categories. Each optimized for different use cases, budgets, and requirements.
The AI labs that train and serve their own flagship models. Higher prices, cutting-edge capabilities, proprietary architectures.
Limited to GPT-3.5 Turbo only, 3 requests/minute rate limit. No free trial credits for new accounts (discontinued mid-2025). Paid tiers unlock GPT-6 Astra at $10/$50 per MTok, GPT-5.6 Sol at $4/$20, GPT-5.6 Terra at $2/$12 and GPT-5.6 Luna at $0.20/$1.20. GPT-4o remains available at $2.50/$10. Batch and Flex at 50% of standard rates.
Key differentiator: Widest model selection; ecosystem leader with function calling, vision, and structured outputs
Limited access via console with rate limits. Fable 5.1: $10/$50 per MTok (input/output). Opus 5: $5/$25 per MTok. Sonnet 5: $2/$10 per MTok. Haiku 4.5: $1/$5 per MTok. Batch API at 50% discount. Adaptive thinking across the current lineup.
Key differentiator: 1M-token context on Fable 5.1, Opus 5 and Sonnet 5; adaptive thinking; Opus 5 holds frontier reasoning at $5/$25
Free tier covers Flash, Flash-Lite and Gemini 2.5 Pro. Gemini 3.8 Flash is $0.75/$3.75 per MTok through December 31, 2026 and $1.50/$7.50 from January 1, 2027. Gemini 2.5 Pro stays at $1.25/$10. Gemini 3.1 Pro is in preview at $2/$12 and requires paid billing. Mandatory spend caps enforced since April 1, 2026.
Key differentiator: 1M token context window; cheapest frontier input tokens; Flash models genuinely free
Experiment tier: 2 RPM, 1 billion tokens/month. No credit card required. Access to the current lineup: Mistral Medium 3.5 at $1.50/$7.50 per MTok, Mistral Large 3 at $0.50/$1.50, Mistral Small 4 at $0.15/$0.60 and Codestral at $0.30/$0.90. Le Chat consumer app included.
Key differentiator: European AI lab; open weights through Mistral Large 3 and the Ministral 3 family; Codestral for code generation
Trial key: 1,000 API calls/month across all endpoints (Chat, Embed, Rerank). Non-commercial use only. Access to Command R+, Rerank 3.5, Embed 4.
Key differentiator: Best-in-class RAG pipeline (Embed + Rerank + Chat); enterprise-focused; non-commercial free tier
$25 in free API credits on signup. Additional $150/month via data sharing program (opt-in, requires $5 minimum spend first). Grok 4.1 Fast: $0.20/M input, $0.50/M output — cheapest frontier model.
Key differentiator: Cheapest frontier pricing ($0.20/M input); $175/month possible in free credits via data sharing
Platforms that serve open-source models on optimized hardware. Cheaper than frontier labs, focused on speed and cost efficiency.
30 RPM, 100K-500K tokens/day depending on model. No credit card required. Custom LPU hardware delivers thousands of tokens/second. Models: Llama 4 Scout 17B, Llama 3.3 70B, Qwen3 32B, Whisper.
Key differentiator: Fastest inference speed (custom LPU hardware); generous free tier with no credit card
1M tokens/day, 10-30 requests/min. Custom wafer-scale chips for multi-thousand tokens/sec inference. Models: Llama 3.1 8B/70B, Qwen 3 235B, GPT-OSS 120B.
Key differentiator: Wafer-scale chip inference; competitive speeds with Groq; 1M tokens/day free
~30 free models including DeepSeek R1, Llama 3.3, Qwen3, Gemma 3. ~20 RPM per model. OpenAI-compatible API. Routes to cheapest provider automatically. One API key for 200+ models.
Key differentiator: Universal gateway to 200+ models; automatic provider routing; OpenAI-compatible API
~40 RPM, 1,000 free API credits. No credit card required for development. Models: Llama 3.1, Mistral, and NVIDIA models. Optimized with TensorRT-LLM for NVIDIA GPUs.
Key differentiator: NVIDIA-optimized inference; deploy on your own NVIDIA GPUs; enterprise support
100 requests/day and $1 free credits. Models: DeepSeek-R1, DeepSeek-V3, QwQ-32B, other open-source models. China-based provider.
Key differentiator: China-based with competitive pricing; strong DeepSeek model support
Platforms for deploying and serving any open-source model. Pay for compute, bring your own model.
$0.10/month free inference credits, 200+ models via Inference Providers, unlimited model hosting on Hub. Routes to multiple inference providers (AWS, GCP, etc.).
Key differentiator: Largest model hub (800K+ models); Inference Providers route to optimal backend; community ecosystem
Free runs on curated model collection without billing. No credit card required to start. Pay-per-second billing by hardware type (CPU/GPU) after free allowance.
Key differentiator: Run any open-source model; per-second GPU billing; one-click model deployment
$30 in free credits for new accounts. Basic plan: $0/month with pay-as-you-go billing after credits. Per-minute GPU/CPU billing for custom deployments, per-token for Model APIs.
Key differentiator: Deploy custom models with autoscaling; $30 free credits; optimized Truss framework
10,000 neurons/day free across text generation, image classification, translation, speech-to-text models. Runs on Cloudflare's 300+ edge locations. No cold starts.
Key differentiator: Edge inference on 300+ locations; zero cold starts; bundled with Workers ecosystem
Providers with unique positioning — free tiers, self-hosting, or niche model access.
5M free tokens for new accounts (30-day validity). DeepSeek V4: $0.30/M input, $0.50/M output (1M context). R1 reasoning: $0.55/M input, $2.19/M output. Cache hits 90% cheaper. Off-peak discounts up to 75% off. China-based.
Key differentiator: Cheapest 1M-context model; cache-hit discounts (90% off); R1 reasoning model competitive with frontier
10-15 RPM, 50-150 requests/day depending on model tier. OpenAI-compatible API endpoint. 100+ models including GPT-4o, Llama, Mistral, Phi. Free for GitHub users.
Key differentiator: Free access to 100+ models for GitHub users; OpenAI-compatible; great for prototyping
No published rate limits on free tier. Supported by donors. 30+ models including DeepSeek R1, Qwen2.5 Coder, text, image, and speech-to-text models. UK-based.
Key differentiator: Completely free donor-supported inference; 30+ models; UK-based
Free tier for light usage with 1 concurrent model. Ollama Cloud provides hosted API. Local installation runs models on your hardware at zero cost. Models: Llama, Mistral, Gemma, and hundreds more.
Key differentiator: Local-first with optional cloud; run models on your own hardware; largest model library
Free tiers range from genuinely production-viable (Groq, GitHub Models) to token giveaways that run out in hours. Here's the honest breakdown.
| Provider | Free Tier Type | Rate Limit | Context Window | When You Hit the Wall |
|---|---|---|---|---|
| OpenAI | Limited free tier | 3 RPM (free) | 1M | Days of moderate use |
| Anthropic | Pay-as-you-go | Rate-limited (free) | 1M | Immediately (pay per token) |
| Google Gemini | Limited free tier | 10-15 RPM (free) | 1M | Days of moderate use |
| Mistral AI | Generous free tier | 2 RPM (free) | 256K | Weeks to months of real usage |
| Cohere | Trial only | 1K calls/month | 128K | 1K calls/month limit |
| xAI (Grok) | Free credits | Standard | 128K | When credits run out |
| Groq | Generous free tier | 30 RPM, 100K-500K tok/day | 128K | Weeks to months of real usage |
| Cerebras | Generous free tier | 10-30 RPM, 1M tok/day | 128K | Weeks to months of real usage |
| OpenRouter | Generous free tier | ~20 RPM per model | Varies | Weeks to months of real usage |
| NVIDIA NIM | Free credits | ~40 RPM | 128K | When credits run out |
| SiliconFlow | Limited free tier | 100 req/day | 64K | Days of moderate use |
| Hugging Face | Limited free tier | Varies | Varies | Days of moderate use |
| Replicate | Generous free tier | Standard | Varies | Weeks to months of real usage |
| Baseten | Free credits | Standard | Varies | When credits run out |
| Cloudflare Workers AI | Generous free tier | 10K neurons/day | Varies | Weeks to months of real usage |
| DeepSeek | Free credits | Standard | 1M | When credits run out |
| GitHub Models | Generous free tier | 10-15 RPM, 50-150 req/day | Varies | Weeks to months of real usage |
| LLM7.io | Generous free tier | No published limits | Varies | Weeks to months of real usage |
| Ollama | Generous free tier | 1 concurrent model | Varies | Weeks to months of real usage |
Token pricing is rarely the full story. These are the costs and limits that surprise developers.
OpenAI o3 and o4-mini models generate internal "reasoning tokens" that count toward output pricing but aren't visible in the response. A simple query can generate 10x more reasoning tokens than output tokens. Monitor usage carefully — your bill reflects total tokens, not just visible output.
Claude's thinking tokens are billed at output rates ($25/M on Opus 5, $50/M on Fable 5.1). A complex reasoning task can produce 10K+ thinking tokens before the actual answer. Budget for 2–5x the output tokens you'd expect from a non-thinking request. Anthropic's current lineup decides its own thinking budget rather than taking one from the request, so the multiplier is harder to cap than it was on the 4.6 generation.
Google's Gemini free tier covers Flash, Flash-Lite and Gemini 2.5 Pro. The current flagship, Gemini 3.1 Pro, has no free tier access — it needs a paid billing account, and paid accounts carry mandatory spend caps.
A 128K context window doesn't mean the model performs well at 128K tokens. Quality degrades in the middle of long contexts ("lost in the middle" problem). For retrieval-heavy tasks, expect effective context of 30–50% of the advertised window. DeepSeek and Gemini's 1M windows are more susceptible to this.
OpenAI and Anthropic increase rate limits based on cumulative spend, not the plan you're on. A new account with $100 in credits still starts at Tier 1 limits. Groq's free tier limits are per-model, so switching models resets your quota.
LLM API pricing is the most volatile in the developer tools space. Here are the changes we've tracked. See full change timeline for all tracked changes.
| Date | Vendor | Change | Impact |
|---|---|---|---|
| discovered Sep 9, 2026 · effective date unknown | GitHub | GitHub Copilot access for students is now limited. Previously offering full access, it now offers a plan with 'limited chat and agent usage with models available through auto model selection only'. Additionally, new Copilot Student signups are paused until 2026-04-20 due to infrastructure strain. Source ↗ | MEDIUM |
| discovered Sep 7, 2026 · effective date unknown | OpenRouter | The free tier now has a rate limit of 50 requests/day, and offers access to 25+ free models and 4 free providers. Previously ~30 free models and ~20 RPM per model. There is also a $25,000/month inference limit with no fees, followed by a 5% fee. Source ↗ | MEDIUM |
| discovered Sep 3, 2026 · effective date unknown | Google Gemini API | The free tier and pricing have been significantly restructured. The current pricing page details Gemini 3.6 Flash, 3.5 Flash, and 3.5 Live Translate with different pricing tiers (Free and Paid) for Standard, Batch, Flexible, and Priority modes. The free tier now offers limited free usage of certain features like Google Search and Maps integration. Source ↗ | HIGH |
| discovered Sep 3, 2026 · effective date unknown | DeepSeek API | Pricing has significantly changed. The previous pricing of $0.28/$0.42/M tokens is no longer accurate. New models (deepseek-v4-flash, deepseek-v4-pro, deepseek-v4-flash-vision-exp) are available with different pricing tiers for peak and off-peak hours, and cache hits/misses. Source ↗ | HIGH |
| discovered Sep 3, 2026 · effective date unknown | Mistral AI | The free tier is no longer an API allowance of 2 RPM and 1B tokens a month across all models. It now carries limited messages, web searches and image generations in Vibe, plus access to Mistral Studio. $10 a month in API credits is sold separately. Source ↗ | HIGH |
| discovered Sep 2, 2026 · effective date unknown | xAI | Grok 4.6 now costs $2.00/M input tokens and $6.00/M output tokens. Source ↗ | HIGH |
| discovered Sep 2, 2026 · effective date unknown | Cloudflare | The program now offers up to $350k in credits instead of $250k. Tier 1 now offers $350k, Tier 2 offers $100k, and Tier 3 offers $10k. Workers AI credit caps are now $2,500 for Tier 3, $10,000 for Tier 2, and $50,000 for Tier 1. AI Gateway is temporarily not covered by credits. Source ↗ | HIGH |
| discovered Sep 2, 2026 · effective date unknown | Anthropic API | Anthropic API's own page now redirects to platform.claude.com, and the lineup behind it advanced. Claude Fable 5.1 is $10/$50 per MTok, Opus 5 is $5/$25, Sonnet 5 is $2/$10 and Haiku 4.5 is $1/$5. Anthropic also states that Sonnet 5's $2/$10 introductory pricing is now the standard price and the rise to $3/$15 scheduled for September 1, 2026 will not occur. Source ↗ | HIGH |
| discovered Sep 1, 2026 · effective date unknown | Cloudflare Workers AI | The pricing has been restructured. While a free tier of 10,000 Neurons per day still exists, the pricing is now more granular and based on per-model unit pricing, billed in Neurons at $0.011 / 1,000 Neurons. Some models now require a paid plan or AI Gateway credits. Source ↗ | HIGH |
| effective Aug 26, 2026 | OpenAI | Assistants API deprecated, full shutdown August 26, 2026. Developers must migrate to Responses API + Conversations API Source ↗ | HIGH |
| effective Jun 1, 2026 | Google Gemini API | Gemini 2.0 Flash and 2.0 Flash-Lite will be deprecated June 1, 2026. Developers must migrate to Gemini 2.5 Flash or 3.x Flash models. Google is consolidating the model lineup — 2.0 generation reaching end of life as 2.5 and 3.x become production-ready. Source ↗ | HIGH |
| effective May 12, 2026 | OpenAI | DALL-E 2 and DALL-E 3 API access discontinued. Developers must migrate to gpt-image-1 (different pricing model, quality tiers changed from standard/hd to low/medium/high) or switch to free alternatives like Pollinations.AI or Lumenfall.ai Source ↗ | HIGH |
| effective May 7, 2026 | OpenAI | Realtime API beta endpoints deprecated. Developers must remove OpenAI-Beta header, use new client_secrets endpoint, specify session_type, and update event names. GA Realtime API is the direct replacement. Source ↗ | HIGH |
| effective Apr 17, 2026 | DeepSeek | DeepSeek V3.2 replaces V4 branding. Pricing dropped: chat model $0.30→$0.28/M input, $0.50→$0.42/M output. Reasoner model pricing unified with chat at $0.28/$0.42 (was $0.55/$2.19). Free token welcome package appears removed. Source ↗ | MEDIUM |
| effective Apr 13, 2026 | xAI | No longer in force (2026-09-05). $25/month free API credits no longer offered — Grok API paid-only No longer in force as of 2026-09-05. This record cites nothing and xAI does not publish its sign-up credit terms on a page we can read, so it cannot be retracted on evidence; it is withdrawn because the claim it makes is contradicted by our own later record. Our xAI offer, verified 2026-08-14 and re-checked 2026-09-02, reads "Sign-up gives $25 in free API credits". Recorded separately: the URL that offer cites, docs.x.ai/developers/models, contains no occurrence of "free", "credit" or "$25" in its 5,452 characters of visible text, so that claim is also uncited and needs a source of its own. We hold no source for this record, so it does not set xAI's rating. | LOW |
| effective Apr 12, 2026 | Gem | Removed: source page no longer accessible or deal program discontinued We hold no source for this record, so it does not set Gem's rating. | LOW |
| effective Apr 12, 2026 | Mode | Removed: source page no longer accessible or deal program discontinued We hold no source for this record, so it does not set Mode's rating. | LOW |
| effective Apr 8, 2026 | Google Gemini API | No longer in force (2026-09-05). Gemini API free tier restricted to Flash and Flash-Lite models only (April 2026). Gemini 2.5 Pro and other Pro models now require a paid billing account. Previously free users could access Pro models with rate limits. Combined with April 1 spend cap enforcement, this significantly narrows what is available at $0. No longer in force as of 2026-09-05: this record's own source page gives Gemini 2.5 Pro a Free Tier of "Free of charge" for input and for output. The contrast is on the same table - Gemini 3.1 Pro Preview reads "Not available" in the Free Tier column, which is what a paid-only model looks like there. Gemini 3.1 Pro is paid-only; 2.5 Pro is not. We hold no evidence either way about the state on 2026-04-08, so this is recorded as no longer in force rather than retracted. Source ↗ | HIGH |
| effective Apr 3, 2026 | OpenAI Codex | Switched from per-seat subscription to pay-as-you-go token-based pricing. Teams can add Codex-only seats billed on token consumption with no rate limits. ChatGPT Business price cut from $25 to $20/month (annual). $100 credit per new Codex team member (up to $500/team, limited time) Source ↗ | MEDIUM |
| effective Apr 1, 2026 | Google Gemini API | Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. Source ↗ | MEDIUM |
Groq (free, fast, no credit card) or OpenRouter (~30 free models, try different providers). GitHub Models for accessing GPT-4o free with a GitHub account.
OpenAI GPT-5.6 Terra ($2/M in, $12/M out) — widest ecosystem, function calling, structured outputs; GPT-4o is still sold at $2.50/$10. Claude Sonnet 5 ($2/$10/M) for nuanced conversation and a 1M-token context.
Claude Opus 5 ($5/$25/M), or Claude Fable 5.1 ($10/$50/M) for long-horizon agentic work. DeepSeek R1 ($0.55/$2.19/M) for budget reasoning. OpenAI GPT-6 Astra ($10/$50/M) for the hardest end-to-end work.
xAI Grok 4.1 Fast ($0.20/$0.50/M) — cheapest frontier model. DeepSeek V4 ($0.30/$0.50/M) with 90% cache-hit discounts. OpenAI/Anthropic batch APIs at 50% off for async workloads.
DeepSeek V4 (1M context, $0.30/M) or Google Gemini 2.5 Pro (1M context, $1.25/M). Claude (200K) for highest quality within context window.
Ollama (free, run locally) for development. Replicate or Baseten for hosted open-source models with dedicated infrastructure. Cloudflare Workers AI for edge inference.
Compare LLM API pricing, search free tiers, and track pricing changes — all from your AI coding assistant.
claude mcp add agentdeals -- npx -y agentdealsWorks with Claude Desktop, Cursor, Cline, Windsurf → Full setup guide