Gemini API billing changes in March and April 2026: spend caps by tier ($250 to $100K+ a month, enforced from April 1), prepay for some new users, and Gemini 3.1 Pro Preview paid only. The free tier covers the Gemini 3.x Flash and Flash-Lite models.
Published 2026-03-26 · Reviewed 2026-09-13, corrections outstanding · Spend caps enforced April 1 · 3.1 Pro paid-only · Prepaid billing live · Figures compiled 2026-03-26 · 47 of 47 figures in the “Free LLM API Comparison” table come from our records for 1,574 developer tools
$250
Tier 1 Spend Cap
Dec 2025
Free tier cut
0
Free Pro models for new projects
8
Providers Compared
Google began enforcing monthly spend caps on the Gemini API on April 1, 2026 (Tier 1: $250, Tier 2: $2,000, Tier 3: $20,000 to $100,000+). When a billing account reaches its cap, requests pause until the next billing month. Since March 23, 2026, AI Studio may ask new users to prepay to set up billing (minimum $5). Gemini 3.1 Pro Preview is paid only.
For developers who built on Gemini's generous early free tier, the API has fundamentally changed: On 2025-12-06 Google cut 2.5 Flash's free tier from 250 requests a day to about 20, and 2.5 Pro's to none. The free tier covers the Gemini 3.x Flash and Flash-Lite models. Google publishes no free-tier limits; AI Studio shows each project's. Since 2026-09-18 Google serves the Gemini 2.5 models only to users who used them before. New projects use 3.5 Flash-Lite or 3.8 Flash. Below we cover what changed, who's affected, and which alternatives offer better free access.
From our tracker: Google cut the Gemini API free tier's rate limits on 2025-12-06. Gemini 2.5 Flash went from 10 requests a minute and 250 a day to 5 and 20, and Gemini 2.5 Pro went to no free requests. Google's rate-limits page stopped listing free-tier limits the same day. Source ↗
Not published. About 5 RPM and 20 RPD on 2.5 Flash, none on 2.5 Pro (2025-12-06)
92% fewer daily requests on 2.5 Flash
Gemini 2.0 Flash
Available, standard limits
Shut down 2026-06-01
Google recommends 3.6 Flash or 3.1 Flash-Lite
Spend caps (Apr 1)
No hard caps — billed without pausing
$250/mo (Tier 1), $2K/mo (Tier 2), $20K+ (Tier 3)
Requests pause at cap
Billing model
Pay-as-you-go for all
Prepay may be required for new users (from March 23, 2026)
Minimum $5 prepayment
Gemini 3.1 Pro
N/A (new model)
Paid only
No free tier
Context window
1M tokens
1M tokens
Unchanged
Flash-Lite limits
Generous (unspecified)
Not published
Free tier preserved
Spend cap details: Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. Source ↗ Developers on pay-as-you-go plans should set project-level budget alerts in Google Cloud Console to avoid unexpected pausing.
2. Timeline of Gemini API Changes
Dated changes to the Gemini API's free tier and billing.
Dec 2025
Free Tier Cut
Google cut the Gemini API free tier's rate limits on 2025-12-06. Gemini 2.5 Flash went from 10 requests a minute and 250 a day to 5 and 20, and Gemini 2.5 Pro went to no free requests. Google's rate-limits page stopped listing free-tier limits the same day. Source ↗
Apr 2026
Spend Caps Enforced
Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. Source ↗
3. Who's Affected
The impact depends on how you use the Gemini API and when you started building on it.
Free tier developers (most affected)
If you built on the free tier before December 2025: on 2025-12-06, 2.5 Flash's daily requests fell from 250 to about 20 and 2.5 Pro's to none. Gemini 3.1 Pro Preview has no free tier. Consider switching to Groq (30 RPM free) or OpenRouter (25+ free models).
Pay-as-you-go developers
The April 1 spend caps affect you directly. If your billing account hits its tier's spend cap, requests pause until the next billing month — not just rate-limited, but fully paused. Set project-level budget alerts now. Teams sharing a billing account are especially vulnerable since all projects count toward the same cap.
Low-volume / hobby developers
If you make a few requests a minute on a 3.x Flash model, the free tier still works — just with less headroom. The 1M token context window remains Gemini's key differentiator. For occasional use, the impact is minimal.
4. New Tier Structure & Spend Caps
Google revamped the Gemini API billing structure with automatic tier upgrades based on spend level. Each tier has a monthly spend cap that pauses all API requests when reached.
Tier
Monthly Spend Cap
Free Rate Limits
Paid Rate Limits
Key Constraint
Free
$0
Not published; shown per project in AI Studio
N/A
No free Pro model for new projects.
Tier 1 (Pay-as-you-go)
$250/mo
Same as free
Higher RPM per model
Requests pause at $250 aggregate spend
Tier 2
$2,000/mo
Same as free
Even higher RPM
Auto-upgraded at spend threshold
Tier 3+
$20K–$100K+
Same as free
Highest RPM, priority
Enterprise scale, custom caps
What "spend cap" means in practice: Unlike rate limits (which reject individual requests), spend caps pause all requests for the remainder of the billing month once the tier's aggregate spend limit is reached. This is a billing-account-level control — all projects under the same billing account share the cap. Google added project-level spend caps on March 12, 2026.
5. Prepaid Billing & Paid-Only Models
Two additional changes that affect new and high-usage developers.
Prepaid billing for new users
Since March 23, 2026, AI Studio may ask a new user to prepay to set up billing (minimum $5); others choose between Prepay and Postpay.
Gemini 3.1 Pro is paid-only
Gemini 3.1 Pro Preview has no free tier. The free tier covers the 3.x Flash and Flash-Lite models, including 3.8 Flash.
Flash-Lite models remain free
The Gemini 3.x Flash and Flash-Lite models are free. Google publishes no free-tier limits; AI Studio shows each project's. For lightweight tasks such as classification, extraction and simple Q&A, 3.5 Flash-Lite or 3.1 Flash-Lite costs nothing.
Claude API with usage-based pricing per million tokens (input/output). Claude Fable 5.1 $10/$50. Claude Opus 5.5 $4/$20. Claude Sonnet 5 $2/$10. Claude Haiku 4.5 $1/$5. The Batch API gives a 50% discount on input and output tokens. Full profile
AI API platform. One model is priced Free in OpenAI's own table: the moderation model omni-moderation-latest. No GPT model is priced free. Full profile
$0.10/$0.50 (gpt-6-luna) – $10.00/$50.00 (gpt-6-astra) per MTok
Fast LLM inference on Groq's LPU hardware. Free plan: gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B, each at 30 requests a minute, 1,000 requests and 200,000 tokens a day, plus Whisper speech-to-text at 2,000 requests a day. Full profile
$0.075/$0.30 (gpt-oss-20b) – $0.15/$0.60 (gpt-oss-120b) per MTok
Mistral's Free plan includes $10 per month in API credits and lets you test Mistral models in Studio, alongside limited messages, web searches and coding sessions in Vibe. Full profile
$0.1/$0.1 (Ministral 3) – $1.5/$7.5 (Mistral Medium 3.5) per MTok
AI model router. Free plan: 25+ free models, 4 free providers, 50 requests a day, no BYOK. Free models are capped at 20 requests a minute; accounts that have bought at least $10 of credits get 1,000 free-model requests a day. Full profile
Ultra-fast LLM inference API — no permanent free tier. New accounts get $5 in free credits after adding a verified payment method, expiring 30 days after they are granted. Full profileWhen we last read the page we cite for this offer, on 2026-09-17, we could read no amount, tier or rate on the page, so we cannot confirm these terms today.
Pay-as-you-go LLM API with no free tier, billed from a prepaid balance. DeepSeek-V4.1-Flash (deepseek-flash): $0.30/M input, $1.20/M output at peak. DeepSeek-V4-Pro (deepseek-v4-pro): $1.32/M input, $3.96/M output at peak. Full profile
$0.30/$1.20 (DeepSeek-V4.1-Flash) – $1.32/$3.96 (DeepSeek-V4-Pro) per MTok
—
Key takeaway:Groq's free plan allows 30 requests a minute and 1,000 a day per model. Mistral AI includes $10 a month in API credits. OpenRouter serves 25+ free models through one API. For the full comparison, see our Free LLM APIs guide.
7. What to Do
Practical steps depending on your situation, from monitoring spend to migrating to alternatives.
If you're staying on Gemini
1. Set project-level spend caps in AI Studio (available since March 12, 2026; Google marks them experimental). This prevents one project from consuming the entire billing account's spend cap. 2. Monitor usage via the Gemini API usage dashboard. Set billing alerts before you hit the spend cap. 3. Move to the 3.x models — Gemini 2.0 Flash shut down on 2026-06-01, and new projects cannot use the 2.5 models. Google points new projects to gemini-3.5-flash-lite or gemini-3.8-flash. 4. Separate billing accounts if you run multiple projects — every project on a billing account shares its tier spend cap.
If you're evaluating alternatives
1. For maximum free requests:Groq — 30 RPM, no credit card, ultra-fast inference. 2. For model variety:OpenRouter — 25+ free models through one OpenAI-compatible API. 3. For long context: Gemini's 1M context window is still the largest free option. If context is your key requirement, stay on Gemini and manage the rate limits. 4. For production workloads:Anthropic and OpenAI also cap monthly spend by usage tier. Anthropic pauses API usage at its tier's cap ($500 a month on Start) until the next month, and OpenAI sets each organization a monthly usage limit ($100 on Tier 1).
Bottom Line
If you need a free LLM API today:
Groq and OpenRouter publish their free limits, which Google no longer does: Groq's free plan allows 30 requests a minute and 1,000 a day per model; OpenRouter's free models allow 20 a minute and 50 a day.
If you need the 1M context window:
Stay on Gemini Flash. Set budget alerts, and use 3.5 Flash-Lite or 3.8 Flash for new work.
If you're a paying Gemini customer:
Set project-level spend caps in AI Studio (available since March 12, 2026). Separate billing accounts for production vs. development. Monitor the billing documentation for tier details.
Related Guides
More resources for evaluating LLM APIs and tracking pricing changes.
Methodology: Gemini API pricing details sourced from Google's Gemini API pricing page and billing documentation. Rate limit reductions recorded in our deal change tracker (3 Gemini changes tracked since December 2025). Compiled 2026-03-26. Risk assessments are based on pricing history and free tier stability.
This guide covers Gemini API pricing changes through September 2026. For the full LLM API comparison, see Free LLM APIs. For Gemini alternatives, see /alternative-to/google-gemini-api. Browse all 1,574 developer tools at /search.
More Alternatives Guides
LocalStack CE Alternatives— LocalStack CE shuts down March 23, 2026 — compare 9 free open-source AWS emulators
Postman Alternatives— Postman killed free team collaboration March 1, 2026 — 5 free API testing alternatives
HCP Terraform Alternatives— HCP Terraform legacy plan ends March 31, 2026 — free IaC alternatives compared
Freshping Alternatives— Freshping shut down March 6, 2026 — 13 free uptime monitoring alternatives
Heroku Alternatives— Heroku removed free tier Nov 2022, entered sustaining mode Feb 2026 — 8 free PaaS options
Firebase Alternatives— Firebase Studio is closing (no new workspaces since June 22, 2026; shutdown March 22, 2027) + Cloud Storage for Firebase now requires Blaze — 7 BaaS alternatives
GitHub Actions Alternatives— GitHub postponed its self-hosted runner fee, so self-hosted runners stay free — 10 free CI/CD alternatives compared
Best Free AI & ML Tools for Developers in 2026— 65+ AI/ML tools and their free tiers compared — LLM APIs, AI coding assistants, ML platforms, observability, and specialized AI services
Best Free LLM APIs in 2026— 25+ LLM API providers and their free tiers compared — proprietary model APIs, open-model inference platforms, and AI gateways with exact rate limits
Best Free API Development Tools in 2026— 39+ free API development tools compared — REST/GraphQL clients, mocking, documentation, marketplaces, and integration platforms
The Complete Free Startup Stack for 2026— Complete free SaaS infrastructure stack — 10 categories with recommended picks, scaling guidance, and stability ratings
The Complete Free AI/ML Stack for 2026— Complete free AI/ML development stack — 10 categories with recommended picks, scaling guidance, and stability ratings
The Complete Free DevOps Stack for 2026— Complete free DevOps infrastructure stack — 10 categories with recommended picks, scaling guidance, and stability ratings
The Complete Free Frontend Stack for 2026— Complete free frontend/Jamstack development stack — 10 categories with recommended picks, scaling guidance, and stability ratings
The Complete Free Next.js Stack for 2026— Complete free Next.js full-stack infrastructure — 10 layers with recommended picks, growth cost analysis, and stability ratings
Google Developer Program 2026— Standalone Google Developer Program Premium no longer takes sign-ups — current plans, Cloud credits and free alternatives
Vercel vs Netlify Free Tier Comparison— Deep comparison of Vercel and Netlify free tiers — bandwidth, functions, builds, commercial use, and scaling costs
Railway vs Render Free Tier Comparison— Deep comparison of Railway and Render free tiers — usage-based vs fixed pricing, databases, sleep behavior, and scaling costs
Datadog vs New Relic Free Tier Comparison— Deep comparison of Datadog and New Relic free tiers — per-host vs per-GB pricing, APM, logs, synthetics, and scaling costs
Free Tier Risk Index— Predictive risk analysis for developer free tiers — grades dated and scored against what happened next, category heatmap, pattern analysis, counter-trends
Free Tier Tracker— Q1 2026 free tier erosion report — which developer free tiers were removed, reduced, or expanded
Startup Credits Comparison 2026— The definitive startup credits comparison — 15+ programs across cloud infrastructure, fintech, and developer tools with eligibility requirements, vesting schedules, and stacking strategies
AI Coding Tools Pricing Guide— AI coding tools pricing comparison — free tiers, pro plans, power tiers, and recent March 2026 pricing changes
AI Coding Tools Pricing Comparison 2026— The definitive AI coding tools comparison — 17 tools across IDE, CLI, cloud agent, and app builder categories with free tier analysis and cost breakdowns
CI/CD Tools Pricing Comparison 2026— The definitive CI/CD pricing comparison — 17+ tools across general, cloud-native, mobile, and self-hosted categories with free tier analysis and cost breakdowns
Database Pricing Comparison 2026— The definitive database pricing comparison — 25+ services across managed Postgres, serverless/edge, document/NoSQL, cloud provider, and specialized categories with free tier analysis and cost breakdowns
Vector Database Pricing Comparison 2026— The definitive vector database pricing comparison — 11 services across dedicated cloud, open-source, pgvector, embedded, and serverless categories with free tier analysis for RAG/AI
Cloud Hosting & PaaS Pricing Comparison 2026— The definitive cloud hosting pricing comparison — 15 platforms across PaaS, edge/serverless, full-featured, and static categories with free tier analysis, pricing gotchas, and Railway referral
LLM API Free Tiers & Free Credits 2026— Which LLM APIs have a genuinely free tier or free credits — frontier labs, inference providers, open-source hosts, and specialized services with free tier analysis and token cost breakdowns
AWS Free Tier Complete Guide 2026— Complete AWS free tier guide — every free service, real limits, hidden costs, and Aurora PostgreSQL on the Free Tier (March 2026)
GCP Free Tier Complete Guide 2026— Complete GCP free tier guide — 20+ always-free products, $300 trial, hidden costs, and comparison with AWS and Azure
Azure Free Tier Complete Guide 2026— Complete Azure free tier guide — 65+ always-free services, $200 trial, Cosmos DB lifetime free tier, and comparison with AWS and GCP
Testing & QA Tools Free Tier Comparison 2026— Side-by-side comparison of 15+ testing tool free tiers — E2E, visual regression, load testing, API testing, local dev, and the testing cost trap at scale
API Development Tools Free Tier Comparison 2026— Side-by-side comparison of 12+ API development tool free tiers — users, collections, requests, mock servers, local-first vs cloud, and the API tool migration trap
Hosting & PaaS Free Tier Comparison 2026— Side-by-side comparison of 12+ hosting free tiers — bandwidth, compute, build minutes, cold starts, commercial use restrictions, and the hosting cost trap at scale
Developer Security Tools Free Tier Comparison 2026— Side-by-side comparison of 20+ developer security tool free tiers — SAST, SCA, DAST, secrets detection, container security, and the DevSecOps cost trap at scale
Firebase Studio Shutdown Guide— Firebase Studio shuts down March 22, 2027 — free cloud IDE alternatives with compute, storage, and collaboration limits compared
Developer Tool Shutdown Tracker 2026— Living tracker of developer tool shutdowns, API sunsets, and deprecation deadlines in 2026 — with migration paths and alternatives
Track Gemini API pricing changes and compare free LLM APIs from your AI assistant. Get alerts on rate limit changes and find alternatives — directly in your editor.
claude mcp add agentdeals -- npx -y agentdeals
Works with Claude Desktop, Cursor, Cline, Windsurf → Full setup guide