Quick Answer
As of August 20, 2026, the floor for mainstream LLM APIs now sits near $0.20 per million input tokens, held by GPT-5.6 Luna at 0.20/1.20 after OpenAI's July 30 price cut. Open-weight and offshore providers go lower still on output. Flagships run $5 to $10 input: GPT-5.6 Sol ($5/$30), Claude Opus 5 ($5/$25), Claude Fable 5 ($10/$50). Mid-tier production models cluster at $2 input.
On July 30, 2026, OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%, and Reddit did what Reddit does. Within hours, r/ClaudeCode had a thread titled “OpenAI just cut prices by up to 80% and Anthropic is crickets,” r/Anthropic was arguing there’s no way the new models cost that little to run, and one r/LLMDevs thread rounded up cheap options “for the token price hike era.” Everyone had an opinion. Almost nobody had a current price table.
That’s the problem with LLM API pricing: it’s a price war now, and the ladder moves under you mid-quarter. The rates you budgeted in August are wrong in September. This page is the fix: every major model’s current rate, ranked by cost, verified against vendor pages, with the multipliers and free tiers that decide what you actually pay.
Cheap tokens and cheap AI are not the same thing, and the gap between them is the part everyone skips. Cheap tokens are not the same as cheap AI. In CloudZero’s 2026 AI ROI survey of 260 finance leaders, 64% said that being able to tie AI spend to outcomes would change how they invest. A price per token is not a measure of value received. Nothing on this page can tell you whether the tokens earned their keep.
How does LLM API pricing work?
Every major provider bills the same way: per token, split into input (what you send) and output (what the model writes back), quoted per million tokens (MTok). A token is roughly three-quarters of an English word.
Two ratios do most of the damage. Output costs 5x input at Anthropic and 6x input at OpenAI, so verbose responses, not long prompts, dominate most bills. And the spread inside a single provider’s lineup runs from 10x at Anthropic ($1 Haiku to $10 Fable) to 150x at OpenAI ($0.20 Luna to $30 GPT-5.5 Pro), which makes model choice the single biggest decision on any AI budget.
Beyond the headline rates, every provider offers the same three discounts with different fine print: cached input at roughly 10% of the standard rate, batch processing at 50% off for asynchronous jobs, and cheaper tiers for simpler work. The fine print is where comparisons get won and lost, and it gets its own section below.
Report
Finance needs to prove AI’s return: CloudZero report
260 senior finance leaders (more than half CFOs) told us why the speed of seeing AI spend, not the size of it, separates who pulls ahead on AI from who gets burned.
Which LLM API costs the least, model by model?
The LLM API pricing comparison table, cheapest to priciest by input rate. All prices per million tokens, standard tier.
| Model | Provider | Input | Output | Context window |
|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | 1.05M |
| GPT-5.4 Nano | OpenAI | $0.20 | $1.25 | 400K |
| Gemini 2.5 Flash | $0.25 | $1.50 | 1M | |
| DeepSeek V3.2 | DeepSeek | $0.28 | $0.42 | 128K |
| GPT-5.4 Mini | OpenAI | $0.75 | $4.50 | 400K |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | 200K |
| Claude Sonnet 5 * | Anthropic | $2.00 | $10.00 | 1M |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | 1.05M |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1M | |
| GPT-5.4 | OpenAI | $2.50 | $15.00 | 1M |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | 1.05M |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | 1M |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | 1M |
| GPT-5.5 Pro | OpenAI | $30.00 | $180.00 | 1M |
* Claude Sonnet 5 launched at 2.00/10.00 as introductory pricing. Anthropic made that rate permanent on August 11, 2026; the previously scheduled September 1 increase to 3.00/15.00 will not happen.
Rates verified against OpenAI’s API pricing, Anthropic’s pricing docs, Google’s Gemini API pricing, and DeepSeek’s platform pricing as of August 20, 2026. Prices move; the change log below tracks what shifted and when.
Three reads from the table that matter more than any single row:
- The $2 tier is the war zone. Claude Sonnet 5, GPT-5.6 Terra, and Gemini 3.1 Pro all sit at exactly $2 input. That’s not a coincidence; it’s the most contested price point in AI, and Anthropic just dug in, making Sonnet 5’s introductory $2 permanent rather than letting it lapse to $3. If your production workload lives here, you’re the customer everyone is fighting over. Act like it.
- Output rates break ties. At the flagship level, GPT-5.6 Sol and Claude Opus 5 both charge $5 input, but Opus is $25 output against Sol’s $30. For agentic work that generates a lot of text, that 17% gap compounds on the most expensive tokens you buy.
- The floor moved on July 30. Luna at $0.20/$1.20 sits in a current flagship family at prices that used to mean “budget legacy model.” For deeper dives on each provider, see CloudZero’s guides to Claude pricing, OpenAI API pricing, Gemini pricing, and DeepSeek pricing.
What is the cheapest LLM API?
Straight answer, no hedging, because that’s the actual question. The cheapest LLM API depends on which cheap you mean:
Cheapest output: the sub-$0.50 output tier belongs to open-weight and offshore providers, not the US frontier labs. Nothing from OpenAI or Anthropic comes close on generation-heavy workloads. The tradeoffs are smaller context windows and the fact that most enterprise teams route around these providers for data governance reasons, which is its own cost conversation.
Cheapest current-generation model: GPT-5.6 Luna at $0.20/$1.20. This is the July 30 story: frontier-family quality at prices that undercut everything except DeepSeek’s output rate. OpenAI framed the cut as advancing the price-performance frontier, which is corporate for “we came for the bottom of the market.”
Cheapest Anthropic option: Claude Haiku 4.5 at 1/5. It stops looking cheap next to Luna, and Haiku does not get the no-surcharge 1M window either: that starts at Claude 4.6, and Haiku 4.5 caps at 200K. The long-context advantage belongs to Sonnet and Opus, covered in the multipliers section.
Cheapest is a moving target. This exact answer changed on July 30 and parts of it change again September 1. Bookmark the price change log below.
One warning the cheap-API listicles skip: cheapest per token is not cheapest per task. A budget model that needs three attempts, longer prompts, or human cleanup costs more than a mid-tier model that nails it once. Benchmark leaderboards rank intelligence; this page ranks price; your evaluations on your own workload are the only LLM comparison that decides anything.
Are free LLM APIs actually free?
The most-searched question in this space, and the honest answer is: free exists, but it’s a trailhead, not a destination.
What’s genuinely available: most major providers issue trial credits to new API accounts (OpenAI and Anthropic both do, in small amounts that expire). Google offers a free tier through AI Studio for development use. Aggregators and open-weight routes offer free access to smaller models with rate limits that make them fine for learning and useless for production. The GitHub awesome-lists that rank on this SERP catalog dozens of these, and they go stale fast because free tiers are the first thing providers adjust.
What the free LLM API hunt actually costs you: every free tier is a funnel into metered usage, and the transition is silent. No invoice announces “your free credits ran out on Tuesday.” The spend just starts. That pattern, small unmonitored starts that compound, is exactly how AI becomes a major cloud line item before anyone assigns it an owner.
One CloudZero customer found OpenAI had reached 25% of their total cloud spend before they could see it, a story that started, like they all do, with somebody’s API key and good intentions.
So use free tiers for what they’re for: evaluation. Run your test suite, compare outputs, pick a tier, and then move to metered usage deliberately, with a budget and an alert threshold set on day one. The teams that get burned aren’t the ones who used free credits; they’re the ones who never noticed the credits end. If you want a rule: the day an API key touches production traffic is the day its spend gets an owner, a forecast, and a dashboard, even if the current bill is $4.
What do the rate cards not tell you?
This is where identical-looking prices diverge, and where AI model pricing comparisons go wrong. Four multipliers, compared across providers:
| Multiplier | OpenAI (GPT-5.6) | Anthropic (Claude 4.6+) |
|---|---|---|
| Long context | 2x input, 1.5x output above ~272K tokens | None. Full 1M window at standard rates |
| Cache reads | 10% of input rate | 10% of input rate |
| Cache writes | 1.25x input | 1.25x (5-min) or 2x (1-hour) |
| Batch discount | 50% off | 50% off |
Long context is the sleeper. Push a 400K-token request through GPT-5.6 Terra and you’re paying 4/18 on the entire request: double on input, half again on output. The same request on Claude Sonnet bills at standard rates. For retrieval-heavy agents that pack the window, this one line flips the OpenAI-vs-Anthropic math, and it never appears in a headline.
Reasoning tokens bill as output. OpenAI’s o-series (o4-mini at $1.10/$4.40 through o3-pro at $20/$80) and extended thinking on Claude both charge internal reasoning at output rates, whether or not you see the tokens. Effective costs on reasoning-heavy tasks run 3x to 10x the base rate. Budget for the thinking, not just the answer.
Caching is the same discount with different failure modes. Both providers sell 90%-off cache reads. Anthropic’s docs put the break-even precisely: caching “pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write).” Both also fail silently: reorder a prompt so the stable prefix changes, and every cached token quietly bills at 10x. Nobody sends a notification. Instrumentation does.
The tokenizer is a hidden repricing. Anthropic notes that Claude 4.7 and later models use a newer tokenizer producing roughly 30% more tokens for the same text. Sonnet 5, Opus 5, and Fable 5 all use it. A per-token comparison against OpenAI therefore understates Claude’s effective cost on identical input, and no rate card shows it.
Batch processing is the discount nobody claims. Both providers take 50% off input and output for asynchronous work, and after each price move the savings are worth rebanking; teams that skip it keep paying last quarter’s prices by default.
How do you compare LLM costs for your actual workload?
The universal formula: (input tokens / 1,000,000 x input rate) + (output tokens / 1,000,000 x output rate), then multiply by monthly call volume and apply cache and batch discounts to the eligible share.
Worked example, one workload across three tiers. A support assistant handling 10,000 conversations daily at 500 input and 300 output tokens each (150M input, 90M output per month):
| Model | Monthly cost |
|---|---|
| GPT-5.6 Sol | $3,450 |
| Claude Sonnet 5 | $1,200 |
| GPT-5.6 Terra | $1,380 |
| Claude Haiku 4.5 | $600 |
| GPT-5.6 Luna | $138 |
| DeepSeek V3.2 | $80 |
Same traffic, a 43x spread, decided entirely by routing. Which is the entire argument for multi-model architectures: route classification and extraction to the cheap end, reserve flagships for tasks where evaluations prove the lift, and re-run the math after every price change.
CloudZero’s interactive LLM cost calculator runs these numbers across providers side by side.
And the part the formula can’t do: telling you whether any of it was worth buying. That’s not a token math problem, it’s an API pricing attribution problem, and it’s where most teams have no answer at all.
Which pricing tier should you actually pick?
The table is fourteen rows; the decision is four rules.
Start one tier lower than your instinct. Almost everyone over-buys intelligence. If Luna, Haiku, or Nano handles the task in your evaluations, the flagship is a 25x tax on insecurity. Escalate on evidence, not instinct.
Pay for output discipline before you pay for a better model. Output costs 5x to 6x input everywhere, so capping response length and killing verbosity often saves more than switching providers. It’s the only optimization that works identically across every row of the table.
Let the multipliers pick between ties. When two models match on sticker price, the decision lives in the fine print: context length patterns (Anthropic wins long-context economics), caching fit (stable prompts cache, dynamic ones don’t), and whether the workload can wait 24 hours for the 50% batch discount.
Re-decide quarterly, minimum. The July 30 cut made GPT-5.4 more expensive than its own successor. Any routing decision older than a quarter is running on expired prices, and in a price war, loyalty to last quarter’s winner is just a donation to your vendor.
How CloudZero connects LLM pricing to AI ROI
Here’s the uncomfortable math this whole page builds toward. You can pick the perfect model at the perfect price and still have no idea whether your AI spend is profitable, because rate optimization and ROI are different questions. The first is “did we pay a fair price per token.” The second is “what did the tokens produce, for which customer, in which feature, at what margin.” Most companies can answer the first with a spreadsheet. Almost none can answer the second, which is exactly the gap those finance leaders in the survey wanted closed.
CloudZero was built for the second question. Native Anthropic and OpenAI integrations ingest token-level usage across every model in the table above, the AI Hub normalizes it alongside AWS, Azure, and GCP spend, and the CostFormation allocation engine maps every dollar to cost per customer, per feature, and per team without tags.
When a provider cuts prices and your usage triples a week later, anomaly detection comparing the last 36 hours against 12 months of history tells the engineer who owns the feature, before the invoice does. Multi-model routing makes this harder, not easier: one AI feature can draw from OpenAI, Anthropic, and a GPU cluster in the same request, three bills in three formats with no shared label. Unifying that into per-feature unit economics is the difference between a price comparison and an investment decision, and it’s the framework in CloudZero’s AI spend management guide.
Schedule a demo to see how leading global organizations connect model-level AI spend to business outcomes or take a self-guided tour to explore cost per customer and AI allocation in action.
Footnote: July 30, 2026, OpenAI cuts Luna 80% and Terra 20%. August 11, 2026, Anthropic makes Sonnet 5’s $2/$10 permanent and cancels the September 1 increase. July 24, 2026, DeepSeek retires V3.2 and the legacy API aliases.