Updated 2026-10-02

AI API Pricing Comparison 2026: GPT vs Claude vs Gemini

Full Pricing Table: All Models Compared

All major AI API models ranked by output price, from cheapest to most expensive. Prices are per million tokens.

ModelProviderInput / 1MOutput / 1MContext WindowTier
GPT-6 LunaOpenAI$0.10$0.501.05Mbudget
Gemini 3.5 Flash-LiteGoogle$0.30$2.501Mbudget
Gemini 3.8 FlashGoogle$0.75$3.751Mbalanced
Claude Haiku 4.5Anthropic$1$5200Kbudget
GPT-6 SolOpenAI$2$101.05Mflagship
Claude Sonnet 5Anthropic$2$101Mbalanced
Gemini 3.1 ProGoogle$2$121Mflagship
Claude Opus 5.5Anthropic$4$201Mflagship
Claude Opus 5Anthropic$5$251Mflagship
GPT-6 AstraOpenAI$10$501.05Mfrontier
Claude Fable 5.1Anthropic$10$501Mfrontier

Output prices range from $0.50 (GPT-6 Luna) to $50 (Claude Fable 5.1) — a 100x difference.

Cheapest AI API for Chat Applications

Chat applications are input-heavy — each user message, system prompt, and conversation history counts as input tokens. For high-volume chat, input price matters most.

The three cheapest models by input price:

  1. GPT-6 Luna — $0.10 / 1M input tokens. Processing 75,000 chat messages costs $0.20.
  2. Gemini 3.5 Flash-Lite — $0.30 / 1M input tokens. Processing 75,000 chat messages costs $0.61.
  3. Gemini 3.8 Flash — $0.75 / 1M input tokens. Processing 75,000 chat messages costs $1.52.

Based on 75,000 messages at ~27 tokens per message (2,025,000 total input tokens).

Best Value for Code Generation

Code generation is output-heavy — the model writes more tokens than it reads. Output price is the dominant cost factor here.

The three cheapest models by output price:

  1. GPT-6 Luna — $0.50 / 1M output tokens. Generating 10,000 lines of code costs $0.05.
  2. Gemini 3.5 Flash-Lite — $2.50 / 1M output tokens. Generating 10,000 lines of code costs $0.25.
  3. Gemini 3.8 Flash — $3.75 / 1M output tokens. Generating 10,000 lines of code costs $0.38.

Based on 10,000 lines at ~10 tokens per line (100,000 total output tokens).

Most Affordable for Document Analysis

Document analysis is input-heavy — you feed in large amounts of text and get back relatively short summaries or answers. Input price dominates.

Cost of processing 1,000 pages across the cheapest models:

  1. GPT-6 Luna — $0.10 / 1M input tokens. Processing 1,000 pages costs $0.07.
  2. Gemini 3.5 Flash-Lite — $0.30 / 1M input tokens. Processing 1,000 pages costs $0.20.
  3. Gemini 3.8 Flash — $0.75 / 1M input tokens. Processing 1,000 pages costs $0.50.

Based on 1,000 pages at ~667 tokens per page (667,000 total input tokens).

Real-World Cost Examples

Three concrete scenarios showing the range between the cheapest and most expensive models.

Processing 100 PDFs

100 documents at 20 pages each = 2,000 pages (1,334,000 input tokens)

Cheapest: GPT-6 Luna — $0.13

Most expensive: Claude Fable 5.1 — $13.34

Building a Chatbot

1,000 messages/day for 30 days (810,000 input tokens)

Cheapest: GPT-6 Luna — $0.08

Most expensive: Claude Fable 5.1 — $8.10

Analyzing 1,000 Images

1,000 images at ~1,000 tokens each (1,000,000 input tokens)

Cheapest: GPT-6 Luna — $0.10

Most expensive: Claude Fable 5.1 — $10

How to Choose the Right Model

Use frontier models (like GPT-6 Astra) when accuracy matters most — complex reasoning, nuanced writing, or tasks where mistakes are costly. The higher price buys measurably better performance on hard problems.

Use budget models (like GPT-6 Luna) for high-volume, routine tasks — classification, summarization, simple chat, or data extraction. They're often 10–60x cheaper and perfectly adequate for straightforward work.

Start cheap, upgrade selectively. Run your workload on a budget model first. If quality isn't sufficient, move up one tier at a time. Most teams find that 80% of their API calls work fine on balanced or budget models, with only the hardest tasks needing frontier.

Want to see exactly what your budget buys?

Try our AI token calculator — pick a model, set a budget, and see how many words, pages, or lines of code you get.