Gemini API billing overhaul April 2026: enforced spend caps ($250-$100K+/mo by tier), prepaid billing for new users, Gemini 3.1 Pro paid-only. Free tier covers Flash, Flash-Lite and Gemini 2.5 Pro. Compare 8 free LLM API alternatives.

Gemini API Pricing 2026 — Free Tier Changes, Spend Caps & Alternatives

Published 2026-03-26 · Spend caps enforced April 1 · 3.1 Pro paid-only · Prepaid billing live · Figures compiled 2026-03-26, last checked 2026-04-02

$250
Tier 1 Spend Cap
50-80%
Rate Limit Cuts
1
Free Pro Model (2.5 Pro)
8
Providers Compared

Google overhauled Gemini API billing effective April 1, 2026. Enforced monthly spend caps by tier (Tier 1: $250/mo, Tier 2: $2,000/mo, Tier 3: $20K-$100K+) automatically pause API requests when reached. New users face prepaid billing — buy credits before using the API. The latest flagship Gemini 3.1 Pro is paid-only with no free tier access.

For developers who built on Gemini's generous early free tier, the API has fundamentally changed: Flash went from ~250 to 20-50 requests/day, spend caps add another constraint for paid users, and the newest model requires payment. Free tier access is preserved for Flash, Flash-Lite and Gemini 2.5 Pro, with restructured rate limits. Below we cover what changed, who's affected, and which alternatives offer better free access. See also our complete pricing overhaul guide with cost analysis by usage tier and migration recommendations.

From our tracker: Google quietly slashed Gemini API free tier rate limits by 50-80% in late 2025. Gemini 2.5 Flash went from ~250 requests/day to 20-50. Pro model free tier removed entirely. A Google PM admitted generous limits were only supposed to be available for a single weekend Source ↗

In This Guide

  1. What's Changing April 1
  2. Timeline of Gemini API Changes
  3. Who's Affected
  4. New Tier Structure & Spend Caps
  5. Prepaid Billing & Paid-Only Models
  6. Free LLM API Comparison
  7. What to Do

1. What's Changing April 1

Google is introducing billing-account-level spend caps that automatically pause API requests when your tier's cap is reached. Combined with rate limit reductions already in effect, this is the third major Gemini API pricing change in six months.

What ChangedBeforeAfterImpact
Free tier rate limitsFlash: ~250 RPD, Pro: availableFlash: 10 RPM (~20-50 RPD). 2.5 Pro: free, 3.1 Pro: paid50-80% reduction
Gemini 2.0 FlashAvailable, standard limitsDeprecated, scheduled for retirementMust migrate to 2.5
Spend caps (Apr 1)No hard caps — billed without pausing$250/mo (Tier 1), $2K/mo (Tier 2), $20K+ (Tier 3)Requests pause at cap
Billing modelPay-as-you-go for allPrepaid billing for new usersBuy credits in advance
Gemini 3.1 ProN/A (new model)Paid-only — no free tier accessLatest flagship paywalled
Context window1M tokens1M tokensUnchanged
Flash-Lite limitsGenerous (unspecified)15 RPM (preserved)Free tier preserved
Spend cap details: Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. Source ↗ Developers on pay-as-you-go plans should set project-level budget alerts in Google Cloud Console to avoid unexpected pausing.

2. Timeline of Gemini API Changes

Three significant changes in six months — a pattern of tightening that started with rate limits and is now extending to spend controls.

Dec 2025

Rate Limits Slashed 50-80%

Google quietly slashed Gemini API free tier rate limits by 50-80% in late 2025. Gemini 2.5 Flash went from ~250 requests/day to 20-50. Pro model free tier removed entirely. A Google PM admitted generous limits were only supposed to be available for a single weekend Source ↗

Mar 2026

Gemini 2.0 Flash Deprecated

Gemini 2.0 Flash and 2.0 Flash-Lite models deprecated, scheduled for retirement. Free tier continues with Gemini 2.5 models. Developers on 2.0 must migrate Source ↗

Apr 2026

Spend Caps Enforced

Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next billing month. Affects pay-as-you-go developers who may see unexpected request pausing. Source ↗

3. Who's Affected

The impact depends on how you use the Gemini API and when you started building on it.

Free tier developers (most affected)

If you built on Gemini's generous early free tier (250+ RPD Flash, free Pro access), your application likely broke when limits dropped 50-80% in late 2025. Flash is now 10 RPM, and the flagship Gemini 3.1 Pro is paid-only. Consider switching to Groq (30 RPM free) or OpenRouter (~30 free models).

Pay-as-you-go developers

The April 1 spend caps affect you directly. If your billing account hits its tier's spend cap, requests pause until the next billing month — not just rate-limited, but fully paused. Set project-level budget alerts now. Teams sharing a billing account are especially vulnerable since all projects count toward the same cap.

Low-volume / hobby developers

If you're making fewer than 10 requests per minute on Flash, the free tier still works — just with less headroom. The 1M token context window remains Gemini's key differentiator. For occasional use, the impact is minimal.

4. New Tier Structure & Spend Caps

Google revamped the Gemini API billing structure with automatic tier upgrades based on spend level. Each tier has a monthly spend cap that pauses all API requests when reached.

TierMonthly Spend CapFree Rate LimitsPaid Rate LimitsKey Constraint
Free$0Flash: 10 RPM, Flash-Lite: 15 RPMN/ANo Pro/3.1 Pro access. Heavily limited.
Tier 1 (Pay-as-you-go)$250/moSame as freeHigher RPM per modelRequests pause at $250 aggregate spend
Tier 2$2,000/moSame as freeEven higher RPMAuto-upgraded at spend threshold
Tier 3+$20K–$100K+Same as freeHighest RPM, priorityEnterprise scale, custom caps
What "spend cap" means in practice: Unlike rate limits (which reject individual requests), spend caps pause all requests for the remainder of the billing month once the tier's aggregate spend limit is reached. This is a billing-account-level control — all projects under the same billing account share the cap. Google added project-level caps on March 16, 2026 as a mitigation tool.

5. Prepaid Billing & Paid-Only Models

Two additional changes that affect new and high-usage developers.

Prepaid billing for new users

New Gemini API users are now enrolled in prepaid billing — you buy credits in advance and consume them with API calls. This replaces the previous pay-as-you-go model for new accounts. Existing users retain pay-as-you-go billing for now, but the transition signals Google's move toward pre-commitment pricing.

Gemini 3.1 Pro is paid-only

The latest flagship model (Gemini 3.1 Pro) has no free tier access. Google's newest model is entirely behind the paywall. Free tier developers are limited to Flash, Flash-Lite and Gemini 2.5 Pro. This widens the gap between free and paid Gemini API capabilities.

Flash-Lite models remain free

Flash-Lite is still free at 15 RPM, with restructured rate limits. For lightweight tasks (classification, extraction, simple Q&A), Flash-Lite remains a viable zero-cost option. Flash is also free at 10 RPM — just significantly reduced from pre-2025 levels.

6. Free LLM API Comparison

How Gemini's free tier compares to alternatives. Sorted by free tier generosity — several providers offer significantly more free access than Gemini post-cuts.

ProviderFree Tier LimitsContextKey ModelsRisk
Google Gemini API 10 RPM (Flash), 15 RPM (Flash-Lite) 1M tokens 2.5 Flash, Flash-Lite, 2.5 Pro (free); 3.1 Pro (paid-only) high
Anthropic Claude API Pay-as-you-go only 1M tokens Fable 5.1, Opus 5, Sonnet 5, Haiku 4.5 none
OpenAI API GPT-3.5 only, 3 RPM 128K tokens GPT-3.5 Turbo (free), GPT-4o (paid) medium
Groq 30 RPM, 100K-500K tokens/day 128K tokens Llama 4, Qwen3, Whisper low
Mistral AI 2 RPM, 1B tokens/month 128K tokens Large, Codestral, Pixtral low
OpenRouter ~20 RPM per model, ~30 free models Varies by model DeepSeek R1, Llama 3.3, Qwen3, Gemma 3 low
Cerebras 10-30 RPM, 1M tokens/day 128K tokens Llama 3.1 8B, Qwen 3 235B, GPT-OSS 120B low
DeepSeek Pay-as-you-go, very low pricing 128K tokens DeepSeek-V3, DeepSeek-R1 low
Key takeaway: Groq offers 30 RPM free (3x Gemini Flash) with ultra-fast inference. Cerebras gives 1M tokens/day free. Mistral AI offers 1B tokens/month free. OpenRouter provides ~30 free models through one API. All are more generous than Gemini's post-cut free tier. For the full comparison, see our Free LLM APIs guide.

7. What to Do

Practical steps depending on your situation, from monitoring spend to migrating to alternatives.

If you're staying on Gemini

1. Set project-level budget caps in Google Cloud Console (available since March 16). This prevents one project from consuming the entire billing account's spend cap.
2. Monitor usage via the Gemini API usage dashboard. Set billing alerts before you hit the spend cap.
3. Migrate to 2.5 models — Gemini 2.0 Flash is deprecated and will be retired. Update your API calls to use gemini-2.5-flash or gemini-2.5-pro.
4. Separate billing accounts if you run multiple projects — spend caps are per billing account, not per project.

If you're evaluating alternatives

1. For maximum free requests: Groq — 30 RPM, no credit card, ultra-fast inference.
2. For maximum free tokens: Cerebras (1M tokens/day) or Mistral AI (1B tokens/month).
3. For model variety: OpenRouter — ~30 free models through one OpenAI-compatible API.
4. For long context: Gemini's 1M context window is still the largest free option. If context is your key requirement, stay on Gemini and manage the rate limits.
5. For production workloads: Anthropic Claude or OpenAI offer more predictable pricing without surprise pausing.

Bottom Line

If you need a free LLM API today:

Groq and OpenRouter offer more generous free tiers than Gemini post-cuts. Switch for better rate limits and no spend cap worries.

If you need the 1M context window:

Stay on Gemini Flash — no other free tier offers this context length. Set budget alerts and migrate to 2.5 models before the 2.0 retirement.

If you're a paying Gemini customer:

Set project-level caps immediately (available since March 16). Separate billing accounts for production vs. development. Monitor the billing documentation for tier details.

Related Guides

More resources for evaluating LLM APIs and tracking pricing changes.

Methodology: Gemini API pricing details sourced from Google's Gemini API pricing page and billing documentation. Rate limit reductions verified via our deal change tracker (3 Gemini changes tracked since December 2025). Compiled 2026-03-26, last checked 2026-04-02. Risk assessments are based on pricing history and free tier stability.

This guide covers Gemini API pricing changes through April 2026. For the full LLM API comparison, see Free LLM APIs. For Gemini alternatives, see /alternative-to/google-gemini-api. Browse all 1,547 developer tools at /search.

Get this data in your AI editor

Track Gemini API pricing changes and compare free LLM APIs from your AI assistant. Get alerts on rate limit changes and find alternatives — directly in your editor.

claude mcp add agentdeals -- npx -y agentdeals