Which LLM APIs have a genuinely free tier or free credits in 2026, and what tokens cost once you exceed it. OpenAI, Anthropic, Google Gemini, Mistral, Groq, DeepSeek, Cerebras, OpenRouter, Cohere and xAI compared — free tier limits, rate limits, context windows and per-token pricing.

LLM API Free Tiers and Free Credits — The 2026 Comparison

Published 2026-04-13 · Reviewed 2026-08-27, corrections outstanding · 19 providers compared · Figures compiled 2026-04-13, last checked 2026-08-27 · 30 pricing changes tracked

19
Providers Compared
9
Generous Free Tier
9
Credits / Limited / Trial
4
Provider Categories

LLM API pricing, frontier rows read 2026-09-05: 19 providers across four categories — frontier labs, inference providers, open-source hosts, and specialized services. OpenAI and Anthropic price their top model identically: GPT-6 Astra and Claude Fable 5.1 are both $10/$50 per M tokens. Gemini 3.8 Flash is $0.75/$3.75 through December 31, 2026 and $1.50/$7.50 after it. Mistral Medium 3.5 is $1.50/$7.50. Google's Gemini free tier keeps Flash, Flash-Lite and 2.5 Pro, with 3.1 Pro paid-only. DeepSeek V4 delivers 1M context at $0.30/M input — the cheapest long-context option. Groq and Cerebras offer genuinely free inference at thousands of tokens/second.

Key trends: Inference providers (Groq, Cerebras, OpenRouter) are commoditizing open-source model access — free tiers with no credit card required. xAI Grok 4.1 is the cheapest frontier model at $0.20/M input. The gap between frontier and open-source quality is narrowing, making the price delta harder to justify for many use cases.

This guide covers: pricing tables, provider breakdowns, free tier analysis, cheapest-per-token rankings, pricing gotchas, recent changes, and best-for-use-case recommendations — compiled by hand from vendor pricing pages.

Cheapest per Million Tokens

Frontier: xAI Grok 4.1 Fast ($0.20/M input, $0.50/M output) · Open-source: Groq Llama 4 Scout ($0.11/M input) · Long context (1M): DeepSeek V4 ($0.30/M input) · Reasoning: DeepSeek R1 ($0.55/M input) vs Claude Opus 5 ($5/M) vs OpenAI o3 ($2/M)

Best Free Tiers

Most generous: Groq (30 RPM, 500K tok/day) · GitHub Models (100+ models free) · Mistral (1B tok/month) · Best for prototyping: OpenRouter (~30 free models) · Cerebras (1M tok/day) · Completely free: LLM7.io (donor-supported) · Ollama (self-hosted)

Jump to section

  1. Pricing Comparison Table
  2. Provider Breakdown (Frontier, Inference, Open-Source, Specialized)
  3. What You Actually Get for Free
  4. Pricing Gotchas
  5. Recent Pricing Changes
  6. Best-for-Use-Case Recommendations
  7. FAQ

Pricing Comparison Table

Frontier prices read from each vendor's own pricing page on 2026-09-05: OpenAI (developers.openai.com/api/docs/pricing), Anthropic (platform.claude.com/docs/en/about-claude/pricing), Google Gemini (ai.google.dev/gemini-api/docs/pricing), Mistral AI (docs.mistral.ai/inference/pricing). The other 15 rows carry no read date. Per-million-token pricing for flagship models. Hover rows to highlight. Click provider names for full vendor profiles.

Provider Free Tier Flagship Model Input /M Output /M Context
OpenAIGPT-3.5 onlyGPT-6 Astra$10/M$50/M1M
AnthropicConsole accessClaude Fable 5.1$10/M$50/M1M
Google GeminiFlash, Flash-Lite, 2.5 ProGemini 3.8 Flash$0.75/M$3.75/M1M
Mistral AI2 RPM, 1B tok/moMistral Medium 3.5$1.50/M$7.50/M256K
Cohere1K calls/moCommand R+$2.50/M$10/M128K
xAI (Grok)$25 signup creditsGrok 4.1 Fast$0.20/M$0.50/M128K
Groq30 RPM freeLlama 4 Scout 17B$0.11/M$0.18/M128K
Cerebras1M tok/dayLlama 3.1 70B$0.60/M$0.60/M128K
OpenRouter~30 free modelsMulti-model gatewayVariesVariesVaries
NVIDIA NIM1K creditsLlama 3.1 70BCredit-basedCredit-based128K
SiliconFlow100 req/day + $1DeepSeek-R1$0.14/M$0.14/M64K
Hugging Face$0.10/mo credits200+ modelsProvider-dependentProvider-dependentVaries
ReplicateFree runs (curated)Llama, Stable DiffusionPer-second billingPer-second billingVaries
Baseten$30 free creditsCustom deploymentsPer-minute GPUPer-minute GPUVaries
Cloudflare Workers AI10K neurons/dayLlama, Mistral, SDXLNeuron-basedNeuron-basedVaries
DeepSeek5M tokens freeDeepSeek V4$0.30/M$0.50/M1M
GitHub Models50-150 req/dayGPT-4o, Llama, MistralFree (rate-limited)Free (rate-limited)Varies
LLM7.ioNo published limits30+ modelsFreeFreeVaries
OllamaFree (light usage)Llama, Mistral, GemmaFree (self-host)Free (self-host)Varies
The price floor: xAI Grok 4.1 Fast at $0.20/M input is the cheapest frontier model. For open-source, Groq and SiliconFlow serve Llama and DeepSeek models at $0.10–0.15/M. DeepSeek V4 offers 1M context at $0.30/M input — 4x cheaper than Gemini 2.5 Pro for long-context work. The batch APIs from OpenAI and Anthropic offer 50% discounts for non-real-time workloads.

Provider Breakdown

LLM API providers fall into four categories. Each optimized for different use cases, budgets, and requirements.

Frontier Labs

The AI labs that train and serve their own flagship models. Higher prices, cutting-edge capabilities, proprietary architectures.

OpenAI Limited free tier

Limited to GPT-3.5 Turbo only, 3 requests/minute rate limit. No free trial credits for new accounts (discontinued mid-2025). Paid tiers unlock GPT-6 Astra at $10/$50 per MTok, GPT-5.6 Sol at $4/$20, GPT-5.6 Terra at $2/$12 and GPT-5.6 Luna at $0.20/$1.20. GPT-4o remains available at $2.50/$10. Batch and Flex at 50% of standard rates.

Key differentiator: Widest model selection; ecosystem leader with function calling, vision, and structured outputs

Anthropic Pay-as-you-go

Limited access via console with rate limits. Fable 5.1: $10/$50 per MTok (input/output). Opus 5: $5/$25 per MTok. Sonnet 5: $2/$10 per MTok. Haiku 4.5: $1/$5 per MTok. Batch API at 50% discount. Adaptive thinking across the current lineup.

Key differentiator: 1M-token context on Fable 5.1, Opus 5 and Sonnet 5; adaptive thinking; Opus 5 holds frontier reasoning at $5/$25

Google Gemini Limited free tier

Free tier covers Flash, Flash-Lite and Gemini 2.5 Pro. Gemini 3.8 Flash is $0.75/$3.75 per MTok through December 31, 2026 and $1.50/$7.50 from January 1, 2027. Gemini 2.5 Pro stays at $1.25/$10. Gemini 3.1 Pro is in preview at $2/$12 and requires paid billing. Mandatory spend caps enforced since April 1, 2026.

Key differentiator: 1M token context window; cheapest frontier input tokens; Flash models genuinely free

Mistral AI Generous free tier

Experiment tier: 2 RPM, 1 billion tokens/month. No credit card required. Access to the current lineup: Mistral Medium 3.5 at $1.50/$7.50 per MTok, Mistral Large 3 at $0.50/$1.50, Mistral Small 4 at $0.15/$0.60 and Codestral at $0.30/$0.90. Le Chat consumer app included.

Key differentiator: European AI lab; open weights through Mistral Large 3 and the Ministral 3 family; Codestral for code generation

Cohere Trial only

Trial key: 1,000 API calls/month across all endpoints (Chat, Embed, Rerank). Non-commercial use only. Access to Command R+, Rerank 3.5, Embed 4.

Key differentiator: Best-in-class RAG pipeline (Embed + Rerank + Chat); enterprise-focused; non-commercial free tier

xAI (Grok) Free credits

$25 in free API credits on signup. Additional $150/month via data sharing program (opt-in, requires $5 minimum spend first). Grok 4.1 Fast: $0.20/M input, $0.50/M output — cheapest frontier model.

Key differentiator: Cheapest frontier pricing ($0.20/M input); $175/month possible in free credits via data sharing

Inference Providers

Platforms that serve open-source models on optimized hardware. Cheaper than frontier labs, focused on speed and cost efficiency.

Groq Generous free tier

30 RPM, 100K-500K tokens/day depending on model. No credit card required. Custom LPU hardware delivers thousands of tokens/second. Models: Llama 4 Scout 17B, Llama 3.3 70B, Qwen3 32B, Whisper.

Key differentiator: Fastest inference speed (custom LPU hardware); generous free tier with no credit card

Cerebras Generous free tier

1M tokens/day, 10-30 requests/min. Custom wafer-scale chips for multi-thousand tokens/sec inference. Models: Llama 3.1 8B/70B, Qwen 3 235B, GPT-OSS 120B.

Key differentiator: Wafer-scale chip inference; competitive speeds with Groq; 1M tokens/day free

OpenRouter Generous free tier

~30 free models including DeepSeek R1, Llama 3.3, Qwen3, Gemma 3. ~20 RPM per model. OpenAI-compatible API. Routes to cheapest provider automatically. One API key for 200+ models.

Key differentiator: Universal gateway to 200+ models; automatic provider routing; OpenAI-compatible API

NVIDIA NIM Free credits

~40 RPM, 1,000 free API credits. No credit card required for development. Models: Llama 3.1, Mistral, and NVIDIA models. Optimized with TensorRT-LLM for NVIDIA GPUs.

Key differentiator: NVIDIA-optimized inference; deploy on your own NVIDIA GPUs; enterprise support

SiliconFlow Limited free tier

100 requests/day and $1 free credits. Models: DeepSeek-R1, DeepSeek-V3, QwQ-32B, other open-source models. China-based provider.

Key differentiator: China-based with competitive pricing; strong DeepSeek model support

Open-Source Model Hosts

Platforms for deploying and serving any open-source model. Pay for compute, bring your own model.

Hugging Face Limited free tier

$0.10/month free inference credits, 200+ models via Inference Providers, unlimited model hosting on Hub. Routes to multiple inference providers (AWS, GCP, etc.).

Key differentiator: Largest model hub (800K+ models); Inference Providers route to optimal backend; community ecosystem

Replicate Generous free tier

Free runs on curated model collection without billing. No credit card required to start. Pay-per-second billing by hardware type (CPU/GPU) after free allowance.

Key differentiator: Run any open-source model; per-second GPU billing; one-click model deployment

Baseten Free credits

$30 in free credits for new accounts. Basic plan: $0/month with pay-as-you-go billing after credits. Per-minute GPU/CPU billing for custom deployments, per-token for Model APIs.

Key differentiator: Deploy custom models with autoscaling; $30 free credits; optimized Truss framework

Cloudflare Workers AI Generous free tier

10,000 neurons/day free across text generation, image classification, translation, speech-to-text models. Runs on Cloudflare's 300+ edge locations. No cold starts.

Key differentiator: Edge inference on 300+ locations; zero cold starts; bundled with Workers ecosystem

Specialized & Free Providers

Providers with unique positioning — free tiers, self-hosting, or niche model access.

DeepSeek Free credits

5M free tokens for new accounts (30-day validity). DeepSeek V4: $0.30/M input, $0.50/M output (1M context). R1 reasoning: $0.55/M input, $2.19/M output. Cache hits 90% cheaper. Off-peak discounts up to 75% off. China-based.

Key differentiator: Cheapest 1M-context model; cache-hit discounts (90% off); R1 reasoning model competitive with frontier

GitHub Models Generous free tier

10-15 RPM, 50-150 requests/day depending on model tier. OpenAI-compatible API endpoint. 100+ models including GPT-4o, Llama, Mistral, Phi. Free for GitHub users.

Key differentiator: Free access to 100+ models for GitHub users; OpenAI-compatible; great for prototyping

LLM7.io Generous free tier

No published rate limits on free tier. Supported by donors. 30+ models including DeepSeek R1, Qwen2.5 Coder, text, image, and speech-to-text models. UK-based.

Key differentiator: Completely free donor-supported inference; 30+ models; UK-based

Ollama Generous free tier

Free tier for light usage with 1 concurrent model. Ollama Cloud provides hosted API. Local installation runs models on your hardware at zero cost. Models: Llama, Mistral, Gemma, and hundreds more.

Key differentiator: Local-first with optional cloud; run models on your own hardware; largest model library

What You Actually Get for Free

Free tiers range from genuinely production-viable (Groq, GitHub Models) to token giveaways that run out in hours. Here's the honest breakdown.

Provider Free Tier Type Rate Limit Context Window When You Hit the Wall
OpenAILimited free tier3 RPM (free)1MDays of moderate use
AnthropicPay-as-you-goRate-limited (free)1MImmediately (pay per token)
Google GeminiLimited free tier10-15 RPM (free)1MDays of moderate use
Mistral AIGenerous free tier2 RPM (free)256KWeeks to months of real usage
CohereTrial only1K calls/month128K1K calls/month limit
xAI (Grok)Free creditsStandard128KWhen credits run out
GroqGenerous free tier30 RPM, 100K-500K tok/day128KWeeks to months of real usage
CerebrasGenerous free tier10-30 RPM, 1M tok/day128KWeeks to months of real usage
OpenRouterGenerous free tier~20 RPM per modelVariesWeeks to months of real usage
NVIDIA NIMFree credits~40 RPM128KWhen credits run out
SiliconFlowLimited free tier100 req/day64KDays of moderate use
Hugging FaceLimited free tierVariesVariesDays of moderate use
ReplicateGenerous free tierStandardVariesWeeks to months of real usage
BasetenFree creditsStandardVariesWhen credits run out
Cloudflare Workers AIGenerous free tier10K neurons/dayVariesWeeks to months of real usage
DeepSeekFree creditsStandard1MWhen credits run out
GitHub ModelsGenerous free tier10-15 RPM, 50-150 req/dayVariesWeeks to months of real usage
LLM7.ioGenerous free tierNo published limitsVariesWeeks to months of real usage
OllamaGenerous free tier1 concurrent modelVariesWeeks to months of real usage

Pricing Gotchas

Token pricing is rarely the full story. These are the costs and limits that surprise developers.

OpenAI Reasoning Tokens: 5–10x Cost Multiplier

OpenAI o3 and o4-mini models generate internal "reasoning tokens" that count toward output pricing but aren't visible in the response. A simple query can generate 10x more reasoning tokens than output tokens. Monitor usage carefully — your bill reflects total tokens, not just visible output.

Anthropic Extended Thinking: Output Tokens Add Up

Claude's thinking tokens are billed at output rates ($25/M on Opus 5, $50/M on Fable 5.1). A complex reasoning task can produce 10K+ thinking tokens before the actual answer. Budget for 2–5x the output tokens you'd expect from a non-thinking request. Anthropic's current lineup decides its own thinking budget rather than taking one from the request, so the multiplier is harder to cap than it was on the 4.6 generation.

Gemini Free Tier: Gemini 3.1 Pro Needs a Paid Account

Google's Gemini free tier covers Flash, Flash-Lite and Gemini 2.5 Pro. The current flagship, Gemini 3.1 Pro, has no free tier access — it needs a paid billing account, and paid accounts carry mandatory spend caps.

Context Window ≠ Effective Context

A 128K context window doesn't mean the model performs well at 128K tokens. Quality degrades in the middle of long contexts ("lost in the middle" problem). For retrieval-heavy tasks, expect effective context of 30–50% of the advertised window. DeepSeek and Gemini's 1M windows are more susceptible to this.

Rate Limits Scale with Spend, Not Plan

OpenAI and Anthropic increase rate limits based on cumulative spend, not the plan you're on. A new account with $100 in credits still starts at Tier 1 limits. Groq's free tier limits are per-model, so switching models resets your quota.

Recent Pricing Changes

LLM API pricing is the most volatile in the developer tools space. Here are the changes we've tracked. See full change timeline for all tracked changes.

Date Vendor Change Impact
discovered Sep 9, 2026 · effective date unknownGitHubGitHub Copilot access for students is now limited. Previously offering full access, it now offers a plan with 'limited chat and agent usage with models available through auto model selection only'. Additionally, new Copilot Student signups are paused until 2026-04-20 due to infrastructure strain. Source ↗MEDIUM
discovered Sep 7, 2026 · effective date unknownOpenRouterThe free tier now has a rate limit of 50 requests/day, and offers access to 25+ free models and 4 free providers. Previously ~30 free models and ~20 RPM per model. There is also a $25,000/month inference limit with no fees, followed by a 5% fee. Source ↗MEDIUM
discovered Sep 3, 2026 · effective date unknownGoogle Gemini APIThe free tier and pricing have been significantly restructured. The current pricing page details Gemini 3.6 Flash, 3.5 Flash, and 3.5 Live Translate with different pricing tiers (Free and Paid) for Standard, Batch, Flexible, and Priority modes. The free tier now offers limited free usage of certain features like Google Search and Maps integration. Source ↗HIGH
discovered Sep 3, 2026 · effective date unknownDeepSeek APIPricing has significantly changed. The previous pricing of $0.28/$0.42/M tokens is no longer accurate. New models (deepseek-v4-flash, deepseek-v4-pro, deepseek-v4-flash-vision-exp) are available with different pricing tiers for peak and off-peak hours, and cache hits/misses. Source ↗HIGH
discovered Sep 3, 2026 · effective date unknownMistral AIThe free tier is no longer an API allowance of 2 RPM and 1B tokens a month across all models. It now carries limited messages, web searches and image generations in Vibe, plus access to Mistral Studio. $10 a month in API credits is sold separately. Source ↗HIGH
discovered Sep 2, 2026 · effective date unknownxAIGrok 4.6 now costs $2.00/M input tokens and $6.00/M output tokens. Source ↗HIGH
discovered Sep 2, 2026 · effective date unknownCloudflareThe program now offers up to $350k in credits instead of $250k. Tier 1 now offers $350k, Tier 2 offers $100k, and Tier 3 offers $10k. Workers AI credit caps are now $2,500 for Tier 3, $10,000 for Tier 2, and $50,000 for Tier 1. AI Gateway is temporarily not covered by credits. Source ↗HIGH
discovered Sep 2, 2026 · effective date unknownAnthropic APIAnthropic API's own page now redirects to platform.claude.com, and the lineup behind it advanced. Claude Fable 5.1 is $10/$50 per MTok, Opus 5 is $5/$25, Sonnet 5 is $2/$10 and Haiku 4.5 is $1/$5. Anthropic also states that Sonnet 5's $2/$10 introductory pricing is now the standard price and the rise to $3/$15 scheduled for September 1, 2026 will not occur. Source ↗HIGH
discovered Sep 1, 2026 · effective date unknownCloudflare Workers AIThe pricing has been restructured. While a free tier of 10,000 Neurons per day still exists, the pricing is now more granular and based on per-model unit pricing, billed in Neurons at $0.011 / 1,000 Neurons. Some models now require a paid plan or AI Gateway credits. Source ↗HIGH
effective Aug 26, 2026OpenAIAssistants API deprecated, full shutdown August 26, 2026. Developers must migrate to Responses API + Conversations API Source ↗HIGH
effective Jun 1, 2026Google Gemini APIGemini 2.0 Flash and 2.0 Flash-Lite will be deprecated June 1, 2026. Developers must migrate to Gemini 2.5 Flash or 3.x Flash models. Google is consolidating the model lineup — 2.0 generation reaching end of life as 2.5 and 3.x become production-ready. Source ↗HIGH
effective May 12, 2026OpenAIDALL-E 2 and DALL-E 3 API access discontinued. Developers must migrate to gpt-image-1 (different pricing model, quality tiers changed from standard/hd to low/medium/high) or switch to free alternatives like Pollinations.AI or Lumenfall.ai Source ↗HIGH
effective May 7, 2026OpenAIRealtime API beta endpoints deprecated. Developers must remove OpenAI-Beta header, use new client_secrets endpoint, specify session_type, and update event names. GA Realtime API is the direct replacement. Source ↗HIGH
effective Apr 17, 2026DeepSeekDeepSeek V3.2 replaces V4 branding. Pricing dropped: chat model $0.30→$0.28/M input, $0.50→$0.42/M output. Reasoner model pricing unified with chat at $0.28/$0.42 (was $0.55/$2.19). Free token welcome package appears removed. Source ↗MEDIUM
effective Apr 13, 2026xAINo longer in force (2026-09-05). $25/month free API credits no longer offered — Grok API paid-only No longer in force as of 2026-09-05. This record cites nothing and xAI does not publish its sign-up credit terms on a page we can read, so it cannot be retracted on evidence; it is withdrawn because the claim it makes is contradicted by our own later record. Our xAI offer, verified 2026-08-14 and re-checked 2026-09-02, reads "Sign-up gives $25 in free API credits". Recorded separately: the URL that offer cites, docs.x.ai/developers/models, contains no occurrence of "free", "credit" or "$25" in its 5,452 characters of visible text, so that claim is also uncited and needs a source of its own. We hold no source for this record, so it does not set xAI's rating.LOW
effective Apr 12, 2026GemRemoved: source page no longer accessible or deal program discontinued We hold no source for this record, so it does not set Gem's rating.LOW
effective Apr 12, 2026ModeRemoved: source page no longer accessible or deal program discontinued We hold no source for this record, so it does not set Mode's rating.LOW
effective Apr 8, 2026Google Gemini APINo longer in force (2026-09-05). Gemini API free tier restricted to Flash and Flash-Lite models only (April 2026). Gemini 2.5 Pro and other Pro models now require a paid billing account. Previously free users could access Pro models with rate limits. Combined with April 1 spend cap enforcement, this significantly narrows what is available at $0. No longer in force as of 2026-09-05: this record's own source page gives Gemini 2.5 Pro a Free Tier of "Free of charge" for input and for output. The contrast is on the same table - Gemini 3.1 Pro Preview reads "Not available" in the Free Tier column, which is what a paid-only model looks like there. Gemini 3.1 Pro is paid-only; 2.5 Pro is not. We hold no evidence either way about the state on 2026-04-08, so this is recorded as no longer in force rather than retracted. Source ↗HIGH
effective Apr 3, 2026OpenAI CodexSwitched from per-seat subscription to pay-as-you-go token-based pricing. Teams can add Codex-only seats billed on token consumption with no rate limits. ChatGPT Business price cut from $25 to $20/month (annual). $100 credit per new Codex team member (up to $500/team, limited time) Source ↗MEDIUM
effective Apr 1, 2026Google Gemini APIBilling-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. Source ↗MEDIUM
The trend: Frontier model pricing is in freefall. Anthropic dropped Opus pricing 67% in 2026. Google is aggressively undercutting on input tokens ($1.25/M for Gemini Pro). Open-source inference is approaching zero — Groq and Cerebras give away millions of tokens daily. The implication: if you're paying more than $5/M input tokens, you should evaluate whether a cheaper model handles your use case.

Best-for-Use-Case Recommendations

Pick the Right LLM API

Best for prototyping

Groq (free, fast, no credit card) or OpenRouter (~30 free models, try different providers). GitHub Models for accessing GPT-4o free with a GitHub account.

Best for production chat / assistants

OpenAI GPT-5.6 Terra ($2/M in, $12/M out) — widest ecosystem, function calling, structured outputs; GPT-4o is still sold at $2.50/$10. Claude Sonnet 5 ($2/$10/M) for nuanced conversation and a 1M-token context.

Best for complex reasoning

Claude Opus 5 ($5/$25/M), or Claude Fable 5.1 ($10/$50/M) for long-horizon agentic work. DeepSeek R1 ($0.55/$2.19/M) for budget reasoning. OpenAI GPT-6 Astra ($10/$50/M) for the hardest end-to-end work.

Best for high-volume / cost-sensitive

xAI Grok 4.1 Fast ($0.20/$0.50/M) — cheapest frontier model. DeepSeek V4 ($0.30/$0.50/M) with 90% cache-hit discounts. OpenAI/Anthropic batch APIs at 50% off for async workloads.

Best for long-context (100K+ tokens)

DeepSeek V4 (1M context, $0.30/M) or Google Gemini 2.5 Pro (1M context, $1.25/M). Claude (200K) for highest quality within context window.

Best for self-hosting / privacy

Ollama (free, run locally) for development. Replicate or Baseten for hosted open-source models with dedicated infrastructure. Cloudflare Workers AI for edge inference.

Frequently Asked Questions

Which LLM API has the best free tier in 2026?
Groq's free tier is 30 RPM with 100K-500K tokens/day, no credit card required, with fast LPU-accelerated inference. GitHub Models gives free access to 100+ models (GPT-4o, Llama, Mistral) for GitHub users. OpenRouter provides ~30 free open-source models. For frontier models specifically, Mistral's Experiment tier gives 1B tokens/month at 2 RPM.
How much does GPT-4o cost per token?
GPT-4o costs $2.50 per million input tokens and $10 per million output tokens. For reference, 1 million tokens is roughly 750,000 words. The batch API offers 50% discount ($1.25/$5 per M tokens). GPT-4o-mini is significantly cheaper at $0.15/$0.60 per M tokens.
How much does Claude cost per token?
Claude Fable 5.1 costs $10/M input and $50/M output tokens. Opus 5 is $5/$25 per M tokens, Sonnet 5 is $2/$10, and Haiku 4.5 is the budget option at $1/$5. The Batch API offers 50% discount on all models.
What is the cheapest LLM API for production use?
For frontier-quality models: xAI Grok 4.1 Fast at $0.20/M input, $0.50/M output. For open-source models: DeepSeek V4 at $0.30/M input, $0.50/M output with cache-hit discounts up to 90%. Groq and Cerebras offer free tiers that can handle moderate production traffic. Google Gemini Flash models are free with rate limits.
Should I use a frontier lab API or an inference provider?
Use frontier lab APIs (OpenAI, Anthropic, Google) when you need their proprietary models (GPT-4o, Claude, Gemini Pro) or specific features (function calling, vision, extended thinking). Use inference providers (Groq, Cerebras, OpenRouter) when running open-source models — they're 5-10x cheaper and often faster. Many apps work well with Llama 3.3 70B or DeepSeek R1 at a fraction of frontier pricing.

Data Source & Methodology

Powered by AgentDeals. The tables on this page were compiled by hand from official vendor pricing pages and have not been re-checked since. Pricing changes are tracked via our deal changes timeline (554 total changes tracked). The pricing changes we track are updated continuously; the tables above are not.

Query this data programmatically via /api/llm-pricing (JSON), our MCP tools, or REST API — search for LLM providers, compare pricing, or track changes from your AI coding assistant.

Get this data in your AI editor

Compare LLM API pricing, search free tiers, and track pricing changes — all from your AI coding assistant.

claude mcp add agentdeals -- npx -y agentdeals

Related Guides

Explore all 1,547 developer tool deals → Browse the full index or connect via MCP