Build AI apps on free tiers. 10 AI/ML infrastructure categories — LLM APIs, vector databases, experiment tracking, observability, and more. Exact limits, scaling guidance.

The Complete Free AI/ML Stack

Everything you need to build, train, and deploy AI applications — without spending a dollar. This guide recommends the best free tier for each layer of an AI/ML development stack — 10 categories from LLM APIs to speech AI — with exact limits pulled from our index of 1,547+ verified developer tools.

Designed for solo AI developers, indie hackers, and startup teams prototyping AI features. Each recommendation includes alternatives, a "when you'll outgrow it" guide, and stability notes based on our tracking of 554 real pricing changes. The limits on this page were read from vendor pricing pages between 2026-07-10 and 2026-09-09.

$0/month
Total estimated monthly cost for prototyping and experimentation
$0 covers the 7 of 10 picks whose free tier our own badge still reads as active. It does not cover GitHub Copilot (at risk), Weights & Biases (at risk), Roboflow (at risk).

Stack Overview

Category Recommended Key Limit Stability
🧠 LLM API Access Groq Ultra-fast LLM inference on LPU hardware — free tier: 30 RPM, 100K-500K tokens/day depending on model… active
💻 AI Coding Assistant GitHub Copilot AI coding assistant by GitHub/Microsoft. Works in VS Code/JetBrains/GitHub.com. Copilot Free: 2,000… at risk
📐 Vector Database Pinecone Vector database — 2 GB storage, 2M write units/month, 1M read units/month, 5 indexes, 5M embedding… active
📊 ML Experiment Tracking Weights & Biases The free tier is $0/mo and includes 1 user seat, experiment tracking, registry & lineage tracking, and is for… read 2026-09-03 at risk
🚀 Model Hosting & Inference Hugging Face ML model hub — $0.10/month free inference credits, 200+ models via Inference Providers, unlimited model… active
⚡ Compute & Training Kaggle Data science platform (Google) — entirely free: 30 hrs/week GPU compute (NVIDIA Tesla P100), 20 hrs/week TPU… active
🏷️ Data Labeling & Annotation Roboflow Public plan is free, requires no credit card, and includes 15 credits / month, 2 users, and Community… read 2026-08-28 at risk
🔍 AI Observability & Evaluation Langfuse Open-source LLM engineering platform that helps teams collaboratively debug, analyze, and iterate on their… active
🔀 AI Gateway & Routing Portkey Control panel for Gen AI apps featuring an observability suite & an AI gateway. Send & log up to 10,000… active
🎙️ Speech & Audio AI Deepgram Speech-to-text and text-to-speech API — $200 free credits on signup (no credit card required), ~43K minutes… active

🧠 LLM API Access

Recommended Groq Free active

Ultra-fast inference on LPU hardware — 30 RPM with 100K-500K tokens/day free. Supports Llama 3.3 70B, Mixtral, Gemma 2. Best balance of speed, limits, and model quality for prototyping.

Ultra-fast LLM inference on LPU hardware — free tier: 30 RPM, 100K-500K tokens/day depending on model. Supports Llama 4 Scout 17B, Llama 3.3 70B, Qwen3 32B, Whisper, and more. No credit card required

When you'll outgrow it: When you exceed 30 RPM or 500K tokens/day. At that point, Cerebras (1M tokens/day) or OpenRouter (~30 free models) extend the free runway. Production apps typically need paid tiers for reliability SLAs.
Full comparison guide →

💻 AI Coding Assistant

Recommended GitHub Copilot Free at risk

2,000 completions/month free. VS Code native with deep GitHub integration. The simplest on-ramp to AI-assisted development.

AI coding assistant by GitHub/Microsoft. Works in VS Code/JetBrains/GitHub.com. Copilot Free: 2,000 completions/month, an allowance of GitHub AI Credits, limited agents, auto model selection only. Copilot Student is…

When you'll outgrow it: When you exceed 2,000 completions/month or need unlimited chat. Cline and Aider are free with your own API keys — no completion limits, just API costs.
Full comparison guide →

📐 Vector Database

Recommended Pinecone Starter active

2 GB storage with 5 indexes and 2M write units/month on the Starter plan. Serverless architecture means zero ops. The most popular vector DB with excellent SDK support.

Vector database — 2 GB storage, 2M write units/month, 1M read units/month, 5 indexes, 5M embedding tokens/month. Pinecone Assistant: 100 docs / 1 GB

When you'll outgrow it: When you exceed 2 GB storage or need more than 5 indexes. Qdrant's 1 GB free forever cluster is a solid alternative with unlimited requests.
Full comparison guide →

📊 ML Experiment Tracking

Recommended Weights & Biases Free at risk

Unlimited experiments with 100 GB storage free. Industry-standard experiment tracking with hyperparameter sweeps, model registry, and team dashboards.

The free tier is $0/mo and includes 1 user seat, experiment tracking, registry & lineage tracking, and is for personal projects only. Corporate use is not allowed. read 2026-09-03

When you'll outgrow it: When you need more than 100 GB storage or team collaboration features beyond 1 user. Neptune.ai offers 200 GB metadata storage for individuals.
Full comparison guide →

🚀 Model Hosting & Inference

Recommended Hugging Face Free active

Free inference API with $0.10/month credits and access to 200+ models. The largest open-source model hub — deploy models, share datasets, and collaborate on ML projects.

ML model hub — $0.10/month free inference credits, 200+ models via Inference Providers, unlimited model hosting on Hub

Also consider:

Replicate Free active
When you'll outgrow it: When you need dedicated endpoints or higher throughput. Free inference API has rate limits and cold starts — production apps need Inference Endpoints ($0.06/hr+).
Full comparison guide →

Compute & Training

Recommended Kaggle Free active

30 hrs/week GPU (Tesla T4) and 20 hrs/week TPU — entirely free. No credit card needed. Integrated datasets and community notebooks.

Data science platform (Google) — entirely free: 30 hrs/week GPU compute (NVIDIA Tesla P100), 20 hrs/week TPU with a 9-hour cap per session, 20 GB working disk per session. Unlimited public notebooks, datasets, and…

When you'll outgrow it: When you need longer sessions (Kaggle limits to 9 hrs), more VRAM (T4 = 16 GB), or persistent storage. Google Colab offers T4 free but with session time limits.
Full comparison guide →

🏷️ Data Labeling & Annotation

Recommended Roboflow Public (Free) at risk

250,000 images free with 10 projects, 2 users, and $60/month compute credits. Best for computer vision projects with built-in model training and deployment.

Public plan is free, requires no credit card, and includes 15 credits / month, 2 users, and Community Support. Data and models are open source on Roboflow Universe. read 2026-08-28

When you'll outgrow it: When you exceed 250K images or need more than 2 users. Labelbox offers 500 LBUs/month. Scale AI gives 1,000 annotation units free.
Full comparison guide →

🔍 AI Observability & Evaluation

Recommended Langfuse Free active

50K observations/month with all features, open-source. Trace LLM calls, evaluate outputs, manage prompts, and debug chains — the emerging standard for LLM ops.

Open-source LLM engineering platform that helps teams collaboratively debug, analyze, and iterate on their LLM applications. Free forever plan includes 50k observations per month and all platform features…

When you'll outgrow it: When you exceed 50K observations/month. Langtrace (50K traces/month) and LangWatch (1K traces/month) are alternatives. For production, self-host Langfuse (open-source) for unlimited traces.
Full comparison guide →

🔀 AI Gateway & Routing

Recommended Portkey Free active

10,000 requests/month free. AI gateway with load balancing, fallbacks, semantic caching, and 200+ LLM providers. Route between models for cost optimization.

Control panel for Gen AI apps featuring an observability suite & an AI gateway. Send & log up to 10,000 requests for free every month.

When you'll outgrow it: When you exceed 10K requests/month. Keywords AI also offers 10K requests/month free. For self-hosted, LiteLLM (open-source) has no limits.
Full comparison guide →

🎙️ Speech & Audio AI

Recommended Deepgram Free Credits active

$200 free credits on signup (~43K minutes transcription). Industry-leading speech-to-text accuracy with real-time streaming, multiple languages, and speaker diarization.

Speech-to-text and text-to-speech API — $200 free credits on signup (no credit card required), ~43K minutes transcription with Nova model, credits never expire. Access to all model endpoints

Also consider:

AssemblyAI Free active
When you'll outgrow it: When you exhaust $200 credits. AssemblyAI offers $50 free credits (~185 hrs). For ongoing free usage, Whisper (open-source) can run locally via Hugging Face.
Full comparison guide →

Stability Notes

Recent pricing changes affecting vendors in this stack. Based on our tracking of 554 deal changes across 1,547+ developer tools.

pricing restructured GitHub Copilot: Premium request model introduced across all tiers. Free: 50/mo, Pro: 300, Pro+: 1,500, Business: 300/user, Enterprise: 1,000/user. Overag... Source ↗
pricing restructured Cursor: Moved from flat monthly subscription to credit-based pricing model with new $200/month Ultra tier Source ↗
new free tier GitHub Copilot: GitHub launched Copilot Free tier — AI code completion available to all GitHub users at no cost. 2,000 completions and 50 chat messages p... Source ↗
restriction Google Gemini API: Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next b... Source ↗
new tier Cursor: Cursor introduced Pro+ tier at $60/month between Pro ($20) and Ultra ($200). Pro+ includes 10x premium model requests vs Pro, priority ac... Source ↗
restriction Google Gemini API: No longer in force (2026-09-05). Gemini API free tier restricted to Flash and Flash-Lite models only (April 2026). Gemini 2.5 Pro and oth... Source ↗
free tier removed Clarifai: Retracted — this record was our error (2026-09-05). Free tier replaced by 14-day trial for Reasoning Engine Retracted 2026-09-05: Clarifa... We hold no source for this record, so it does not set Clarifai's rating.
rebranded Keywords AI: Rebranded to Respan, free tier preserved We hold no source for this record, so it does not set Keywords AI's rating.
restriction GitHub Copilot: Student plan now managed under dedicated 'GitHub Copilot Student' plan. Premium models (GPT-5.4, Claude Opus, Claude Sonnet) removed from... Source ↗
pricing restructured Cursor: Retracted — this record was our error (2026-09-02). Cursor now offers 6 plans: Free, Hobby ($10/mo), Pro ($20/mo), Pro+ ($60/mo), Busines... Source ↗
product deprecated Google Gemini API: Gemini 2.0 Flash and 2.0 Flash-Lite will be deprecated June 1, 2026. Developers must migrate to Gemini 2.5 Flash or 3.x Flash models. Goo... Source ↗
restriction GitHub Copilot: No longer in force (2026-09-02). GitHub paused new Copilot Pro trials due to abuse. Existing trials continued through their term. New use... Source ↗

View all 554 pricing changes →

Open-Source Self-Hosted Alternatives

For each category, there's an open-source option you can run on your own hardware — no API limits, no vendor lock-in, complete data privacy.

Category Tool Details
LLM Inference Ollama Run Llama, Mistral, Gemma locally — no API limits, complete privacy. GPU recommended.
Experiment Tracking MLflow Apache-licensed ML lifecycle platform. Self-host for unlimited experiments and model registry.
Vector Database Chroma Open-source embedding database. Run locally in Python with zero configuration.
AI Observability Langfuse (self-hosted) Same Langfuse, no observation limits. Deploy via Docker or Kubernetes.
Data Labeling Label Studio Open-source annotation tool. Unlimited projects, all data types (text, image, audio, video).
AI Gateway LiteLLM Open-source proxy to 100+ LLMs. Unified API, load balancing, spend tracking. No request limits.
Compute Ollama + llama.cpp Run quantized models on consumer hardware. M-series Macs run 7B-13B models at usable speeds.

When You'll Outgrow Free Tiers

Most AI projects can prototype entirely on free tiers. Here's what typically hits limits first:

  1. GPU compute time (Kaggle 30 hrs/week) — fine-tuning and training sessions eat GPU hours fast
  2. LLM API rate limits (Groq 30 RPM) — production apps with concurrent users need higher throughput
  3. Vector storage (Pinecone 2 GB) — RAG applications with large document corpora
  4. Speech credits (Deepgram $200 one-time) — real-time transcription burns through credits quickly
  5. Observability volume (Langfuse 50K obs/month) — high-traffic AI apps with detailed tracing

The good news: the open-source alternatives above have no limits. Ollama + Chroma + MLflow + Langfuse (self-hosted) gives you a complete local AI stack at zero ongoing cost.

Need general SaaS infrastructure too? See our Free Startup Stack for hosting, databases, auth, and more. Or search our full index of 1,547+ developer deals.

Get this data in your AI editor

Get personalized AI stack recommendations from your AI assistant. Compare free tiers, check model limits, and plan your AI infrastructure — directly in your editor.

claude mcp add agentdeals -- npx -y agentdeals