Compare 25+ LLM API providers and their free tiers — Groq, OpenRouter, Gemini, Mistral, NVIDIA NIM, and more, beside paid APIs from OpenAI, Anthropic and Cerebras. Exact rate limits and token quotas. Catalogue dates August to September 2026.

Best Free LLM APIs for Developers

Free LLM API access has never been better. Groq delivers gpt-oss-120b at 30 RPM on custom LPU hardware — the fastest free inference available. Mistral's Free plan includes $10 a month in API credits. OpenRouter aggregates 25+ free models through one OpenAI-compatible API. GitHub Models is recorded as Retired — GitHub ended it, so it is no longer a way to reach 100+ models for free.

This page compares 20 LLM API providers — from proprietary model APIs (OpenAI, Anthropic, Gemini) to open-model inference platforms (Groq, Cerebras, NVIDIA NIM) and AI gateways (OpenRouter, Portkey). The rate limit comparison table below has the data developers actually need when choosing a provider.

Recent LLM API Pricing Changes

View all 476 pricing changes →

Proprietary Model APIs

First-party APIs from the companies that train frontier models. Higher quality ceilings but typically lower free tier limits — these are the APIs behind GPT-4, Claude, Gemini, and Grok.

Google Gemini API Free (Reduced) stable

Free tier: Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash, 3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, Gemini 3 Flash Preview, Gemini Embedding 2 and Gemma 4 are free of charge. Gemini 3.1 Pro Preview is paid only. Google publishes no free-tier rate limits; each project's limits are shown in Google AI Studio. Google's pricing table marks free-tier use as "Used to improve our products". Since 2026-09-18 Google serves the Gemini 2.5 models only to users who have used them before, and points new projects to 3.5 Flash-Lite or 3.8 Flash. Paid, per million tokens (input/output): Gemini 3.8 Flash $0.75/$3.75 until 2026-12-31, then double; Gemini 3.5 Flash $1.50/$9; Gemini 3.5 Flash-Lite $0.30/$2.50; Gemini 3.1 Flash-Lite $0.25/$1.50; Gemini 3.1 Pro Preview $2/$12 (prompts up to 200K tokens). Accounts opened after 2026-03-02 cannot spend the $300 Google Cloud welcome credit on the Gemini API.

Mistral's Free plan includes $10 per month in API credits and lets you test Mistral models in Studio, alongside limited messages, web searches and coding sessions in Vibe. API keys work in Free mode with no credit card, within usage and rate limits. In Free mode, Mistral may use your inputs and outputs to train its models unless you opt out. API prices: Ministral 3 (3B) $0.1/$0.1 (per 1M tokens); Mistral Small 4 $0.15/$0.6 (per 1M tokens); Mistral Large 3 $0.5/$1.5 (per 1M tokens); Mistral Medium 3.5 $1.5/$7.5 (per 1M tokens). Batch processing is half price, and cached input tokens cost up to 90% less.

Cohere Free stable

AI model API for Command chat models, Embed, Rerank, Transcribe and Parse. Every account starts with a Trial API key: calls made with it are free, limited to 1,000 API calls a month, and may not be used for production or commercial purposes. Trial rate limits: 20 requests a minute per Chat model, 2,000 Embed inputs a minute and 10 Rerank requests a minute. Trial keys can use all of Cohere's models and APIs. Production keys are pay-as-you-go. Prices per 1M tokens: Command R7B $0.0375/$0.15; Command R $0.15/$0.60. The pricing page lists Command A+ (Apache 2.0) at $0 through an API key and as a model download.

OpenAI Pay-as-you-go

AI API platform. One model is priced Free in OpenAI's own table: the moderation model omni-moderation-latest. No GPT model is priced free. GPT prices per 1M tokens (Standard, short context): gpt-6-astra $10.00/$50.00; gpt-6-sol $2.00/$10.00; gpt-6-luna $0.10/$0.50; gpt-5.6-sol $4.00/$20.00 (a promotional price, available at least through November 21, 2026); gpt-5.6-terra $2.00/$12.00; gpt-5.6-luna $0.20/$1.20. Batch and Flex processing halve these prices. Embeddings from $0.02 per 1M tokens. Three further free amounts are sub-quotas inside paid tools: 1 GB of File search storage (then $0.10/GB per day), 1 GB per account per month of ChatKit upload storage, and search content tokens from the web search preview tool on non-reasoning models.

xAI Pay-as-you-go

Grok API, pay as you go: sign up at console.x.ai, then load it with credits. grok-4.7 $2.00/$6.00 (per 1M tokens, under 200k prompt tokens). grok-4.3 $1.25/$2.50 (under 200k prompt tokens). grok-build-0.1 $1.00/$2.00 (under 200k prompt tokens). Grok 4.1 Fast was retired on 2026-05-15; its model names now route to grok-4.3 at grok-4.3 rates.The page we cite for this offer does not name it when we last looked, on 2026-09-05, so we cannot confirm these terms today.

Anthropic API Pay-as-you-go

Claude API with usage-based pricing per million tokens (input/output). Claude Fable 5.1 $10/$50. Claude Opus 5.5 $4/$20. Claude Sonnet 5 $2/$10. Claude Haiku 4.5 $1/$5. The Batch API gives a 50% discount on input and output tokens. New users receive a small amount of free credits to test the API.

Open-Model Inference Platforms

Platforms that host open-weight models (Llama, Mistral, Qwen, Gemma) on optimized hardware. Often the most generous free tiers — you get fast inference on powerful models without paying or sharing data for training.

ML model hub. Free users get $0.10 a month of Inference Providers credits (subject to change); Inference Providers serves 200+ models, and extra usage requires a credits purchase. PRO ($9/month) gets $2.00 a month. Free accounts get 100GB of private storage and best-effort public storage; substantial storage needs PRO, Team or Enterprise.When we last read the page we cite for this offer, on 2026-09-13, we refused the change we considered recording because it named no figure that had moved, and refusing a change is not a confirmation of the terms above, so we cannot confirm these terms today.

Groq Free caution 2026-08-16 reduced

Fast LLM inference on Groq's LPU hardware. Free plan: gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B, each at 30 requests a minute, 1,000 requests and 200,000 tokens a day, plus Whisper speech-to-text at 2,000 requests a day. Llama 3.3 70B and Llama 3.1 8B left the free and developer plans on 2026-08-16 and are Enterprise-only. Developer plan prices per 1M tokens: gpt-oss-120b $0.15/$0.60; gpt-oss-20b $0.075/$0.30.

OpenRouter Free stable

AI model router. Free plan: 25+ free models, 4 free providers, 50 requests a day, no BYOK. Free models are capped at 20 requests a minute; accounts that have bought at least $10 of credits get 1,000 free-model requests a day. The Standard plan (pay-as-you-go) charges a 5.5% fee on credit purchases and includes BYOK up to $25,000 of list-price inference a month with no fees, 5% after.

Workers AI runs AI models on Cloudflare's global network. Every account gets 10,000 Neurons per day at no charge, reset daily at 00:00 UTC; on the Workers Free plan, requests beyond that fail until the reset. On the Workers Paid plan (from $5 per month), usage above 10,000 Neurons per day costs $0.011 per 1,000 Neurons. Seven models, including @cf/moonshotai/kimi-k2.6, @cf/zai-org/glm-5.3 and @cf/deepseek-ai/deepseek-v4-pro-0813, require the Workers Paid plan or prepaid AI Gateway credits.

Cerebras Trial

Ultra-fast LLM inference API — no permanent free tier. New accounts get $5 in free credits after adding a verified payment method, expiring 30 days after they are granted. Free Trial limits on gpt-oss-120b and qwen-3.8-27b: 5 requests/min and 1M tokens/day. Pay as you go after: GPT OSS 120B $0.35/$0.75 per M tokens, Qwen 3.8 27B $0.99/$1.49 per M tokens. Multi-thousand tokens/sec inference speedWhen we last read the page we cite for this offer, on 2026-09-17, we could read no amount, tier or rate on the page, so we cannot confirm these terms today.

Replicate Trial

ML model hosting and inference. A new account can run the models in Replicate's Try for Free collection without buying credit, for a limited number of runs that Replicate does not state; after that you add billing and buy credit. Replicate says these models are not free forever. Most models bill by the second for the hardware they run on, such as $0.000025/sec on CPU Small and $0.001525/sec on one H100. Official models bill per output instead, such as FLUX 1.1 Pro at $0.04 per image; per million input/output tokens: DeepSeek-R1 $3.75/$10.When we last read the page we cite for this offer, on 2026-09-17, we found the page did not mention this offer, which is not evidence it ended, so we cannot confirm these terms today.

NVIDIA NIM Free

Free inference endpoints on build.nvidia.com: models marked Free Endpoint, including Kimi K3, DeepSeek V4.1 Flash and NVIDIA Nemotron, can be called at no cost. Up to 40 requests per minute; limits may vary by model, and traffic from other users may cause throttling. NVIDIA's API Trial Terms allow free use for testing and evaluation only, not production.The page we cite for this offer states no amount, tier or rate we can read when we last looked, on 2026-09-02, so we cannot confirm these terms today.

Cloud-hosted Ollama for running open-source LLMs — The Free plan costs $0 and includes starter usage credits. Buying usage credits unlocks all models. It includes 1 concurrent request.

AI Gateways & Specialized Inference

Model routers, observability gateways, and specialized inference platforms. These add routing, monitoring, or domain-specific capabilities on top of LLM APIs.

Baseten Basic (Free Credits)

ML model deployment platform. New workspaces receive credits for testing and deployment; Baseten does not state the amount. Basic plan: $0 per month, pay as you go. Dedicated deployments bill per minute. Model APIs bill per 1M tokens: GLM-5.3 $1.40/$4.40. GLM-5.3-Flash $0.15/$0.50.When we last read the page we cite for this offer, on 2026-09-10, we found a change we could not reconcile with the terms we publish, so we cannot confirm these terms today.

Keywords AI Free stable

Rebranded to Respan; keywordsai.co redirects to respan.ai. Free plan: full platform, 100k logs, 1k scores, 5 datasets, 2 evaluators, 5 prompts. No credit card required.

Lumenfall.ai Free stable

AI media gateway: one OpenAI-compatible API to image and video generation models from several providers, billed at each provider's price with no markup, platform fee or subscription. One model, FLUX.1 [schnell] FP8, is listed free per image ("Free to try", served through Fireworks AI). New accounts also get a $1 credit with no card required; Lumenfall's terms say promotional credits typically expire after 90 days. After that you top up prepaid credits. Under Lumenfall's terms (2026-03-05), content sent through free features, including free API usage, may be used to train AI models.

Social media scheduling, publishing and analytics. Free plan at $0/month. The pricing page states no per-plan entitlements: the feature list shown is identical for all four plans.

easy-to-use, free image generation AI with free API available. No signups or API keys required, and several option for integrating into a website or workflow. [#opensource](https://github.com/pollinations/pollinations)The page we cite for this offer states no amount, tier or rate we can read when we last looked, on 2026-09-14, so we cannot confirm these terms today.

Portkey Free

Control panel for Gen AI apps featuring an observability suite & an AI gateway. Send & log up to 10,000 requests for free every month.When we last read the page we cite for this offer, on 2026-09-18, we could read no amount, tier or rate on the page, so we cannot confirm these terms today.

Free LLM API Rate Limit Comparison

The data developers actually need — exact rate limits, token quotas, and model availability for every free LLM API tier.

Provider Type Free Tier Paid Rate
Groq Inference Fast LLM inference on Groq's LPU hardware. Free plan: gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B, each at 30 requests a minute, 1,000 requests and 200,000 tokens a day, plus Whisper speech-to-text at 2,000 requests a day. Full profile $0.075/$0.30 (gpt-oss-20b) – $0.15/$0.60 (gpt-oss-120b) per MTok
Cerebras Inference Ultra-fast LLM inference API — no permanent free tier. New accounts get $5 in free credits after adding a verified payment method, expiring 30 days after they are granted. Full profileWhen we last read the page we cite for this offer, on 2026-09-17, we could read no amount, tier or rate on the page, so we cannot confirm these terms today. $0.35/$0.75 (GPT OSS 120B) – $0.99/$1.49 (Qwen 3.8 27B) per MTok
Mistral AI Provider Mistral's Free plan includes $10 per month in API credits and lets you test Mistral models in Studio, alongside limited messages, web searches and coding sessions in Vibe. Full profile $0.1/$0.1 (Ministral 3) – $1.5/$7.5 (Mistral Medium 3.5) per MTok
OpenRouter Gateway AI model router. Free plan: 25+ free models, 4 free providers, 50 requests a day, no BYOK. Free models are capped at 20 requests a minute; accounts that have bought at least $10 of credits get 1,000 free-model requests a day. Full profile — vendor pricing
GitHub Models Retired GitHub retired GitHub Models on 2026-07-30, so there is no free tier. Full profile —
Google Gemini API Provider Free tier: Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash, 3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, Gemini 3 Flash Preview, Gemini Embedding 2 and Gemma 4 are free of charge. Gemini 3.1 Pro Preview is paid only. Full profile $0.25/$1.50 (Gemini 3.1 Flash-Lite) – $2/$12 (Gemini 3.1 Pro Preview) per MTok
NVIDIA NIM Inference Free inference endpoints on build.nvidia.com: models marked Free Endpoint, including Kimi K3, DeepSeek V4.1 Flash and NVIDIA Nemotron, can be called at no cost. Full profileThe page we cite for this offer states no amount, tier or rate we can read when we last looked, on 2026-09-02, so we cannot confirm these terms today. — vendor pricing
Cloudflare Workers AI Inference Workers AI runs AI models on Cloudflare's global network. Every account gets 10,000 Neurons per day at no charge, reset daily at 00:00 UTC; on the Workers Free plan, requests beyond that fail until the reset. Full profile — vendor pricing
OpenAI Provider AI API platform. One model is priced Free in OpenAI's own table: the moderation model omni-moderation-latest. No GPT model is priced free. Full profile $0.10/$0.50 (gpt-6-luna) – $10.00/$50.00 (gpt-6-astra) per MTok
Anthropic API Provider Claude API with usage-based pricing per million tokens (input/output). Claude Fable 5.1 $10/$50. Claude Opus 5.5 $4/$20. Claude Sonnet 5 $2/$10. Claude Haiku 4.5 $1/$5. The Batch API gives a 50% discount on input and output tokens. Full profile $1/$5 (Claude Haiku 4.5) – $10/$50 (Claude Fable 5.1) per MTok
Hugging Face Platform ML model hub. Free users get $0.10 a month of Inference Providers credits (subject to change); Inference Providers serves 200+ models, and extra usage requires a credits purchase. PRO ($9/month) gets $2.00 a month. Full profileWhen we last read the page we cite for this offer, on 2026-09-13, we refused the change we considered recording because it named no figure that had moved, and refusing a change is not a confirmation of the terms above, so we cannot confirm these terms today. — vendor pricing
xAI Provider Grok API, pay as you go: sign up at console.x.ai, then load it with credits. grok-4.7 $2.00/$6.00 (per 1M tokens, under 200k prompt tokens). grok-4.3 $1.25/$2.50 (under 200k prompt tokens). Full profileThe page we cite for this offer does not name it when we last looked, on 2026-09-05, so we cannot confirm these terms today. $1.00/$2.00 (grok-build-0.1) – $2.00/$6.00 (grok-4.7) per MTok

Mistral's Free plan includes $10 a month in API credits. OpenRouter gives one API key for 25+ free models. GitHub Models is recorded as Retired, so the widest free selection it carried is no longer one of the options here. Of the proprietary frontier APIs, xAI and Anthropic are pay-as-you-go (Anthropic gives new users a small amount of free credits to test the API), and OpenAI prices no GPT model free. Catalogue dates August to September 2026.

Which Free LLM API Should I Use?

Need the fastest free LLM inference?
Groq — custom LPU hardware delivers the fastest token generation, 30 RPM free with gpt-oss-120b. No credit card required.
Need maximum free token volume?
Mistral AI — $10 a month in API credits on the Free plan.
Want one API key for many models?
OpenRouter — 25+ free models through one OpenAI-compatible API; free models are capped at 20 requests a minute and 50 a day, or 1,000 a day once you have bought at least $10 of credits. GitHub Models is recorded as Retired and is no longer a second route to many models.
Need a long context window?
Google Gemini API — 1M token context window on Flash models. Google publishes no free-tier limits; AI Studio shows each project's.
Building AI agents and need routing/observability?
Portkey — AI gateway with load balancing, fallbacks, and caching across providers. Keywords AI for LLM monitoring and optimization.
Want to run models at the edge?
Cloudflare Workers AI — 10,000 Neurons a day at no charge on every account; on the Workers Free plan, requests beyond that fail until the daily reset at 00:00 UTC. Some models, including Kimi K2.6 and GLM-5.3, need the Workers Paid plan or prepaid AI Gateway credits.
Want to test hosted models before paying?
NVIDIA NIM — free endpoints on build.nvidia.com for models marked Free Endpoint, such as Kimi K3, DeepSeek V4.1 Flash and NVIDIA Nemotron, at up to 40 requests per minute; NVIDIA's API Trial Terms allow testing and evaluation only, not production. Baseten — new workspaces receive credits for testing and deployment; Baseten does not state the amount.
Want completely free, self-hosted inference?
Ollama — open source (MIT) and free to run on your own machine; its cloud models are paid with usage credits, and the Free plan includes starter credits. Or download Hugging Face models and run them locally; Hugging Face's hosted Inference Providers API serves 200+ models, with $0.10 a month of credits for free users (subject to change).

Looking for more? Browse all AI / ML tools or see our broader AI & ML tools comparison covering 83+ tools including AI coding, observability, and specialized services. See also: our spend caps deep-dive.

Get this data in your AI editor

Get LLM API recommendations from your AI assistant. Compare rate limits, models, and pricing — directly in your editor.

claude mcp add agentdeals -- npx -y agentdeals