Gemini API billing changes in March and April 2026: spend caps by tier ($250 to $100K+ a month, enforced from April 1), prepay for some new users, and Gemini 3.1 Pro Preview paid only. The free tier covers the Gemini 3.x Flash and Flash-Lite models.

Gemini API Pricing 2026 — Free Tier Changes, Spend Caps & Alternatives

Published 2026-03-26 · Reviewed 2026-09-13, corrections outstanding · Spend caps enforced April 1 · 3.1 Pro paid-only · Prepaid billing live · Figures compiled 2026-03-26 · 47 of 47 figures in the “Free LLM API Comparison” table come from our records for 1,574 developer tools

$250
Tier 1 Spend Cap
Dec 2025
Free tier cut
0
Free Pro models for new projects
8
Providers Compared

Google began enforcing monthly spend caps on the Gemini API on April 1, 2026 (Tier 1: $250, Tier 2: $2,000, Tier 3: $20,000 to $100,000+). When a billing account reaches its cap, requests pause until the next billing month. Since March 23, 2026, AI Studio may ask new users to prepay to set up billing (minimum $5). Gemini 3.1 Pro Preview is paid only.

For developers who built on Gemini's generous early free tier, the API has fundamentally changed: On 2025-12-06 Google cut 2.5 Flash's free tier from 250 requests a day to about 20, and 2.5 Pro's to none. The free tier covers the Gemini 3.x Flash and Flash-Lite models. Google publishes no free-tier limits; AI Studio shows each project's. Since 2026-09-18 Google serves the Gemini 2.5 models only to users who used them before. New projects use 3.5 Flash-Lite or 3.8 Flash. Below we cover what changed, who's affected, and which alternatives offer better free access.

From our tracker: Google cut the Gemini API free tier's rate limits on 2025-12-06. Gemini 2.5 Flash went from 10 requests a minute and 250 a day to 5 and 20, and Gemini 2.5 Pro went to no free requests. Google's rate-limits page stopped listing free-tier limits the same day. Source ↗

In This Guide

  1. What Changed on April 1
  2. Timeline of Gemini API Changes
  3. Who's Affected
  4. New Tier Structure & Spend Caps
  5. Prepaid Billing & Paid-Only Models
  6. Free LLM API Comparison
  7. What to Do

1. What Changed on April 1

Since April 1, 2026, billing-account spend caps pause API requests when an account reaches its tier's cap.

What ChangedBeforeAfterImpact
Free tier rate limits2.5 Flash: 10 RPM, 250 RPD. 2.5 Pro: 2 RPM, 50 RPDNot published. About 5 RPM and 20 RPD on 2.5 Flash, none on 2.5 Pro (2025-12-06)92% fewer daily requests on 2.5 Flash
Gemini 2.0 FlashAvailable, standard limitsShut down 2026-06-01Google recommends 3.6 Flash or 3.1 Flash-Lite
Spend caps (Apr 1)No hard caps — billed without pausing$250/mo (Tier 1), $2K/mo (Tier 2), $20K+ (Tier 3)Requests pause at cap
Billing modelPay-as-you-go for allPrepay may be required for new users (from March 23, 2026)Minimum $5 prepayment
Gemini 3.1 ProN/A (new model)Paid onlyNo free tier
Context window1M tokens1M tokensUnchanged
Flash-Lite limitsGenerous (unspecified)Not publishedFree tier preserved
Spend cap details: Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. Source ↗ Developers on pay-as-you-go plans should set project-level budget alerts in Google Cloud Console to avoid unexpected pausing.

2. Timeline of Gemini API Changes

Dated changes to the Gemini API's free tier and billing.

Dec 2025

Free Tier Cut

Google cut the Gemini API free tier's rate limits on 2025-12-06. Gemini 2.5 Flash went from 10 requests a minute and 250 a day to 5 and 20, and Gemini 2.5 Pro went to no free requests. Google's rate-limits page stopped listing free-tier limits the same day. Source ↗

Apr 2026

Spend Caps Enforced

Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. Source ↗

3. Who's Affected

The impact depends on how you use the Gemini API and when you started building on it.

Free tier developers (most affected)

If you built on the free tier before December 2025: on 2025-12-06, 2.5 Flash's daily requests fell from 250 to about 20 and 2.5 Pro's to none. Gemini 3.1 Pro Preview has no free tier. Consider switching to Groq (30 RPM free) or OpenRouter (25+ free models).

Pay-as-you-go developers

The April 1 spend caps affect you directly. If your billing account hits its tier's spend cap, requests pause until the next billing month — not just rate-limited, but fully paused. Set project-level budget alerts now. Teams sharing a billing account are especially vulnerable since all projects count toward the same cap.

Low-volume / hobby developers

If you make a few requests a minute on a 3.x Flash model, the free tier still works — just with less headroom. The 1M token context window remains Gemini's key differentiator. For occasional use, the impact is minimal.

4. New Tier Structure & Spend Caps

Google revamped the Gemini API billing structure with automatic tier upgrades based on spend level. Each tier has a monthly spend cap that pauses all API requests when reached.

TierMonthly Spend CapFree Rate LimitsPaid Rate LimitsKey Constraint
Free$0Not published; shown per project in AI StudioN/ANo free Pro model for new projects.
Tier 1 (Pay-as-you-go)$250/moSame as freeHigher RPM per modelRequests pause at $250 aggregate spend
Tier 2$2,000/moSame as freeEven higher RPMAuto-upgraded at spend threshold
Tier 3+$20K–$100K+Same as freeHighest RPM, priorityEnterprise scale, custom caps
What "spend cap" means in practice: Unlike rate limits (which reject individual requests), spend caps pause all requests for the remainder of the billing month once the tier's aggregate spend limit is reached. This is a billing-account-level control — all projects under the same billing account share the cap. Google added project-level spend caps on March 12, 2026.

5. Prepaid Billing & Paid-Only Models

Two additional changes that affect new and high-usage developers.

Prepaid billing for new users

Since March 23, 2026, AI Studio may ask a new user to prepay to set up billing (minimum $5); others choose between Prepay and Postpay.

Gemini 3.1 Pro is paid-only

Gemini 3.1 Pro Preview has no free tier. The free tier covers the 3.x Flash and Flash-Lite models, including 3.8 Flash.

Flash-Lite models remain free

The Gemini 3.x Flash and Flash-Lite models are free. Google publishes no free-tier limits; AI Studio shows each project's. For lightweight tasks such as classification, extraction and simple Q&A, 3.5 Flash-Lite or 3.1 Flash-Lite costs nothing.

6. Free LLM API Comparison

How Gemini's free tier compares to alternatives.

ProviderFree TierPaid RateRisk
Google Gemini API Free tier: Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash, 3.5 Flash, 3.5 Flash-Lite, 3.1 Flash-Lite, Gemini 3 Flash Preview, Gemini Embedding 2 and Gemma 4 are free of charge. Gemini 3.1 Pro Preview is paid only. Full profile $0.25/$1.50 (Gemini 3.1 Flash-Lite) – $2/$12 (Gemini 3.1 Pro Preview) per MTok stable
Anthropic Claude API Claude API with usage-based pricing per million tokens (input/output). Claude Fable 5.1 $10/$50. Claude Opus 5.5 $4/$20. Claude Sonnet 5 $2/$10. Claude Haiku 4.5 $1/$5. The Batch API gives a 50% discount on input and output tokens. Full profile $1/$5 (Claude Haiku 4.5) – $10/$50 (Claude Fable 5.1) per MTok —
OpenAI API AI API platform. One model is priced Free in OpenAI's own table: the moderation model omni-moderation-latest. No GPT model is priced free. Full profile $0.10/$0.50 (gpt-6-luna) – $10.00/$50.00 (gpt-6-astra) per MTok —
Groq Fast LLM inference on Groq's LPU hardware. Free plan: gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B, each at 30 requests a minute, 1,000 requests and 200,000 tokens a day, plus Whisper speech-to-text at 2,000 requests a day. Full profile $0.075/$0.30 (gpt-oss-20b) – $0.15/$0.60 (gpt-oss-120b) per MTok caution
Mistral AI Mistral's Free plan includes $10 per month in API credits and lets you test Mistral models in Studio, alongside limited messages, web searches and coding sessions in Vibe. Full profile $0.1/$0.1 (Ministral 3) – $1.5/$7.5 (Mistral Medium 3.5) per MTok caution
OpenRouter AI model router. Free plan: 25+ free models, 4 free providers, 50 requests a day, no BYOK. Free models are capped at 20 requests a minute; accounts that have bought at least $10 of credits get 1,000 free-model requests a day. Full profile — vendor pricing stable
Cerebras Ultra-fast LLM inference API — no permanent free tier. New accounts get $5 in free credits after adding a verified payment method, expiring 30 days after they are granted. Full profileWhen we last read the page we cite for this offer, on 2026-09-17, we could read no amount, tier or rate on the page, so we cannot confirm these terms today. $0.35/$0.75 (GPT OSS 120B) – $0.99/$1.49 (Qwen 3.8 27B) per MTok —
DeepSeek Pay-as-you-go LLM API with no free tier, billed from a prepaid balance. DeepSeek-V4.1-Flash (deepseek-flash): $0.30/M input, $1.20/M output at peak. DeepSeek-V4-Pro (deepseek-v4-pro): $1.32/M input, $3.96/M output at peak. Full profile $0.30/$1.20 (DeepSeek-V4.1-Flash) – $1.32/$3.96 (DeepSeek-V4-Pro) per MTok —
Key takeaway: Groq's free plan allows 30 requests a minute and 1,000 a day per model. Mistral AI includes $10 a month in API credits. OpenRouter serves 25+ free models through one API. For the full comparison, see our Free LLM APIs guide.

7. What to Do

Practical steps depending on your situation, from monitoring spend to migrating to alternatives.

If you're staying on Gemini

1. Set project-level spend caps in AI Studio (available since March 12, 2026; Google marks them experimental). This prevents one project from consuming the entire billing account's spend cap.
2. Monitor usage via the Gemini API usage dashboard. Set billing alerts before you hit the spend cap.
3. Move to the 3.x models — Gemini 2.0 Flash shut down on 2026-06-01, and new projects cannot use the 2.5 models. Google points new projects to gemini-3.5-flash-lite or gemini-3.8-flash.
4. Separate billing accounts if you run multiple projects — every project on a billing account shares its tier spend cap.

If you're evaluating alternatives

1. For maximum free requests: Groq — 30 RPM, no credit card, ultra-fast inference.
2. For model variety: OpenRouter — 25+ free models through one OpenAI-compatible API.
3. For long context: Gemini's 1M context window is still the largest free option. If context is your key requirement, stay on Gemini and manage the rate limits.
4. For production workloads: Anthropic and OpenAI also cap monthly spend by usage tier. Anthropic pauses API usage at its tier's cap ($500 a month on Start) until the next month, and OpenAI sets each organization a monthly usage limit ($100 on Tier 1).

Bottom Line

If you need a free LLM API today:

Groq and OpenRouter publish their free limits, which Google no longer does: Groq's free plan allows 30 requests a minute and 1,000 a day per model; OpenRouter's free models allow 20 a minute and 50 a day.

If you need the 1M context window:

Stay on Gemini Flash. Set budget alerts, and use 3.5 Flash-Lite or 3.8 Flash for new work.

If you're a paying Gemini customer:

Set project-level spend caps in AI Studio (available since March 12, 2026). Separate billing accounts for production vs. development. Monitor the billing documentation for tier details.

Related Guides

More resources for evaluating LLM APIs and tracking pricing changes.

Methodology: Gemini API pricing details sourced from Google's Gemini API pricing page and billing documentation. Rate limit reductions recorded in our deal change tracker (3 Gemini changes tracked since December 2025). Compiled 2026-03-26. Risk assessments are based on pricing history and free tier stability.

This guide covers Gemini API pricing changes through September 2026. For the full LLM API comparison, see Free LLM APIs. For Gemini alternatives, see /alternative-to/google-gemini-api. Browse all 1,574 developer tools at /search.

Get this data in your AI editor

Track Gemini API pricing changes and compare free LLM APIs from your AI assistant. Get alerts on rate limit changes and find alternatives — directly in your editor.

claude mcp add agentdeals -- npx -y agentdeals