Build AI apps on free tiers. 10 AI/ML infrastructure categories — LLM APIs, vector databases, experiment tracking, observability, and more. Exact limits, scaling guidance.

The Complete Free AI/ML Stack

Published 2026-03-25 · Not yet reviewed

Everything you need to build, train, and deploy AI applications — without spending a dollar. This guide recommends the best free tier for each layer of an AI/ML development stack — 10 categories from LLM APIs to speech AI — with exact limits for each. Figures in the tables below come from our records for 1,574 developer tools.

Designed for solo AI developers, indie hackers, and startup teams prototyping AI features. Each recommendation includes alternatives, a "when you'll outgrow it" guide, and stability notes based on our tracking of 476 real pricing changes. The limits on this page were read from vendor pricing pages between 2026-07-10 and 2026-09-26.

$0/month
Total estimated monthly cost for prototyping and experimentation
$0 covers the 3 of 10 picks whose free tier our own badge still reads as active. It does not cover Groq (at risk), GitHub Copilot (at risk), Hugging Face (unrated — change refused), Roboflow (at risk), Langfuse (unrated — change not reconciled), Portkey (unrated — change not established), Deepgram (credits only).

Stack Overview

Category Recommended Key Limit Stability
🧠 LLM API Access Groq Fast LLM inference on Groq's LPU hardware. Free plan: gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B, each at 30… at risk
💻 AI Coding Assistant GitHub Copilot AI coding assistant by GitHub. Copilot Free is $0 with no credit card: 2,000 code completions a month, a… at risk
📐 Vector Database Pinecone Vector database. The Starter plan is free: up to 2 GB storage, 2M write units and 1M read units a month, 1 GB… active
📊 ML Experiment Tracking Weights & Biases ML experiment tracking, model registry, and LLM app tracing and evaluation (Weave). The cloud Free plan is… active
🚀 Model Hosting & Inference Hugging Face ML model hub. Free users get $0.10 a month of Inference Providers credits (subject to change); Inference… unrated — change refused
⚡ Compute & Training Kaggle Google's data science platform with free notebooks, datasets and competitions. Notebooks get a weekly GPU… active
🏷️ Data Labeling & Annotation Roboflow Computer vision platform for labeling, training and deploying models. On September 18, 2026 Roboflow… at risk
🔍 AI Observability & Evaluation Langfuse Open-source LLM engineering platform for tracing, evaluating and debugging AI applications. The Langfuse… unrated — change not reconciled
🔀 AI Gateway & Routing Portkey Control panel for Gen AI apps featuring an observability suite & an AI gateway. Send & log up to 10,000… unrated — change not established
🎙️ Speech & Audio AI Deepgram Free $200 Credit then pay-as-you-go. Flux TTS is free until 9/12/2026. Voice Agent API's Flux TTS is free… read 2026-08-28 credits only

🧠 LLM API Access

Recommended Groq Free at risk

Ultra-fast inference on LPU hardware — 30 RPM, 1,000 requests and 200K tokens a day per model, free. Serves gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B. Best balance of speed, limits, and model quality for prototyping.

Fast LLM inference on Groq's LPU hardware. Free plan: gpt-oss-120b, gpt-oss-20b and Qwen3.8 27B, each at 30 requests a minute, 1,000 requests and 200,000 tokens a day, plus Whisper speech-to-text at 2,000 requests a…

When you'll outgrow it: When you exceed 30 RPM or 200K tokens a day on a model. At that point, OpenRouter (25+ free models) extends the free runway. Production apps typically need paid tiers for reliability SLAs.
Full comparison guide →

💻 AI Coding Assistant

Recommended GitHub Copilot Free at risk

2,000 completions/month free. VS Code native with deep GitHub integration. The simplest on-ramp to AI-assisted development.

AI coding assistant by GitHub. Copilot Free is $0 with no credit card: 2,000 code completions a month, a limited monthly allowance of GitHub AI Credits for chat and agent features, and models through auto model…

When you'll outgrow it: When you exceed 2,000 completions/month or need unlimited chat. Cline and Aider are free with your own API keys — no completion limits, just API costs.
Full comparison guide →

📐 Vector Database

Recommended Pinecone Free active

2 GB storage with 5 indexes and 2M write units/month on the Starter plan. Serverless architecture means zero ops. The most popular vector DB with excellent SDK support.

Vector database. The Starter plan is free: up to 2 GB storage, 2M write units and 1M read units a month, 1 GB of egress a month (reads that return data stop at the cap until the next billing period), 5 indexes, 1…

Also consider:

Qdrant Free active
When you'll outgrow it: When you exceed 2 GB storage or need more than 5 indexes. Qdrant Cloud's free cluster (1 GB RAM, 4 GB disk) is an alternative; Qdrant suspends it after a week without use.
Full comparison guide →

📊 ML Experiment Tracking

Recommended Weights & Biases Free active

Free cloud plan for personal development: up to 5 model seats, 5 GB/month of storage and 1 GB/month of Weave ingestion. Industry-standard experiment tracking.

ML experiment tracking, model registry, and LLM app tracing and evaluation (Weave). The cloud Free plan is $0/mo and designed for personal development, with up to 5 model seats, 5 GB/mo of storage and 1 GB/mo of Weave…

Also consider:

Comet ML Free at risk
When you'll outgrow it: When you need more than 5 GB/month of storage or more than 5 model seats. W&B Pro starts at $60/month with a 30-day free trial.
Full comparison guide →

🚀 Model Hosting & Inference

Recommended Hugging Face Free unrated — change refused

Free inference API with $0.10/month credits and access to 200+ models. The largest open-source model hub — deploy models, share datasets, and collaborate on ML projects.

ML model hub. Free users get $0.10 a month of Inference Providers credits (subject to change); Inference Providers serves 200+ models, and extra usage requires a credits purchase. PRO ($9/month) gets $2.00 a month. Free…

When you'll outgrow it: When you need dedicated endpoints or higher throughput. Free users get $0.10 a month of Inference Providers credits — production apps can use Inference Endpoints (dedicated, from $0.033/hour).
Full comparison guide →

⚡ Compute & Training

Recommended Kaggle Free active

A weekly quota of 30 GPU hours (one P100 or two T4s) and up to 20 TPU hours, at no charge. Integrated datasets and community notebooks.

Google's data science platform with free notebooks, datasets and competitions. Notebooks get a weekly GPU quota of 30 hours, sometimes more depending on demand (one NVIDIA Tesla P100 or two Tesla T4s), and a TPU v3-8…

When you'll outgrow it: When you need longer sessions (Kaggle caps them at 12 hours on CPU or GPU and 9 on TPU), more VRAM (T4 = 16 GB), or persistent storage. Google Colab offers T4 free but with session time limits.
Full comparison guide →

🏷️ Data Labeling & Annotation

Recommended Roboflow Free at risk

Free Core plan: 10 credits that refresh every month, private projects and models, and model weight download. Best for computer vision projects with built-in model training and deployment.

Computer vision platform for labeling, training and deploying models. On September 18, 2026 Roboflow announced a free tier for its Core plan: 10 credits that refresh every month, private projects and models, and model…

When you'll outgrow it: When you need more than the 10 monthly credits. Labelbox offers 500 LBUs/month.
Full comparison guide →

🔍 AI Observability & Evaluation

Recommended Langfuse Free unrated — change not reconciled

50K units/month free on Cloud Hobby (traces, observations and scores each count), all platform features with limits, 30 days of data access, open-source. Trace LLM calls, evaluate outputs, manage prompts, and debug chains — the emerging standard for LLM ops.

Open-source LLM engineering platform for tracing, evaluating and debugging AI applications. The Langfuse Cloud Hobby plan is free with no credit card: 50k units a month (each trace, observation and score counts as one…

When you'll outgrow it: When you exceed 50K units/month; Langfuse Core is $29/month with 100K units. Langtrace (50K traces/month) and LangWatch (1K traces/month) are alternatives. For production, self-host Langfuse (open-source) for unlimited traces.
Full comparison guide →

🔀 AI Gateway & Routing

Recommended Portkey Free unrated — change not established

10,000 requests/month free. AI gateway with load balancing, fallbacks, semantic caching, and 200+ LLM providers. Route between models for cost optimization.

Control panel for Gen AI apps featuring an observability suite & an AI gateway. Send & log up to 10,000 requests for free every month.

When you'll outgrow it: When you exceed 10K requests/month. Keywords AI also offers 10K requests/month free. For self-hosted, LiteLLM (open-source) has no limits.
Full comparison guide →

🎙️ Speech & Audio AI

Recommended Deepgram Free Credits credits only

$200 free credits on signup (~43K minutes transcription). Industry-leading speech-to-text accuracy with real-time streaming, multiple languages, and speaker diarization.

Free $200 Credit then pay-as-you-go. Flux TTS is free until 9/12/2026. Voice Agent API's Flux TTS is free through September 12, 2026. read 2026-08-28

When you'll outgrow it: When you exhaust $200 credits. AssemblyAI offers $50 free credits (~185 hrs). For ongoing free usage, Whisper (open-source) can run locally via Hugging Face.
Full comparison guide →

Stability Notes

Recent pricing changes affecting vendors in this stack. Based on our tracking of 476 deal changes across 1,574+ developer tools.

pricing restructured GitHub Copilot: Premium request model introduced across all tiers. Free: 50/mo, Pro: 300, Pro+: 1,500, Business: 300/user, Enterprise: 1,000/user. Overag... Source ↗
pricing restructured Cursor: Cursor moved Pro from 500 requests a month to usage-based pricing and introduced a $200/month Ultra tier. Source ↗
new free tier GitHub Copilot: GitHub launched Copilot Free tier — AI code completion available to all GitHub users at no cost. 2,000 completions and 50 chat messages p... Source ↗
restriction Google Gemini API: Billing-account-level spend caps enforced starting April 1, 2026. When a tier's spend cap is reached, API requests pause until the next b... Source ↗
rebranded Keywords AI: Rebranded to Respan, free tier preserved We hold no source for this record, so it does not set Keywords AI's rating.
restriction GitHub Copilot: Student plan now managed under dedicated 'GitHub Copilot Student' plan. Premium models (GPT-5.4, Claude Opus, Claude Sonnet) removed from... Source ↗
product deprecated Google Gemini API: Gemini 2.0 Flash and 2.0 Flash-Lite were shut down on 2026-06-01, as Google announced on 2026-02-18. Google names gemini-3.6-flash as the... Source ↗
restriction GitHub Copilot: No longer in force (2026-09-02). GitHub paused new Copilot Pro trials due to abuse. Existing trials continued through their term. New use... Source ↗
restriction GitHub Copilot: No longer in force (2026-06-17). GitHub paused new signups for Copilot Pro, Pro+, and Copilot Student plans citing infrastructure strain ... Source ↗
restriction GitHub Copilot: No longer in force (2026-09-01). GitHub paused new self-serve signups for Copilot Business by organizations on GitHub Free and GitHub Tea... Source ↗
pricing restructured GitHub Copilot: GitHub retired premium requests and now bills every Copilot plan in GitHub AI Credits. Copilot Free kept 2,000 code completions a month, ... Source ↗
limits reduced Roboflow: Roboflow's free Public plan went from "$60/mo free credits" to "15 credits / month" between the Internet Archive's 2026-06-23 and 2026-07... Source ↗

View all 476 pricing changes →

Open-Source Self-Hosted Alternatives

For each category, there's an open-source option you can run on your own hardware — no API limits, no vendor lock-in, complete data privacy.

Category Tool Details
LLM Inference Ollama Run Llama, Mistral, Gemma locally — no API limits, complete privacy. GPU recommended.
Experiment Tracking MLflow Apache-licensed ML lifecycle platform. Self-host for unlimited experiments and model registry.
Vector Database Chroma Open-source embedding database. Run locally in Python with zero configuration.
AI Observability Langfuse (self-hosted) Same Langfuse, no observation limits. Deploy via Docker or Kubernetes.
Data Labeling Label Studio Open-source annotation tool. Unlimited projects, all data types (text, image, audio, video).
AI Gateway LiteLLM Open-source proxy to 100+ LLMs. Unified API, load balancing, spend tracking. No request limits.
Compute Ollama + llama.cpp Run quantized models on consumer hardware. M-series Macs run 7B-13B models at usable speeds.

When You'll Outgrow Free Tiers

Most AI projects can prototype entirely on free tiers. Here's what typically hits limits first:

  1. GPU compute time (Kaggle 30 hrs/week) — fine-tuning and training sessions eat GPU hours fast
  2. LLM API rate limits (Groq 30 RPM) — production apps with concurrent users need higher throughput
  3. Vector storage (Pinecone 2 GB) — RAG applications with large document corpora
  4. Speech credits (Deepgram $200 one-time) — real-time transcription burns through credits quickly
  5. Observability volume (Langfuse 50K units/month) — high-traffic AI apps with detailed tracing

The good news: the open-source alternatives above have no limits. Ollama + Chroma + MLflow + Langfuse (self-hosted) gives you a complete local AI stack at zero ongoing cost.

Need general SaaS infrastructure too? See our Free Startup Stack for hosting, databases, auth, and more. Or search our full index of 1,574+ developer deals.

Get this data in your AI editor

Get personalized AI stack recommendations from your AI assistant. Compare free tiers, check model limits, and plan your AI infrastructure — directly in your editor.

claude mcp add agentdeals -- npx -y agentdeals