Contents
How does Azure OpenAI pricing work? What models are available on Azure OpenAI in 2026? What changed in Azure OpenAI pricing this summer? PTUs vs pay-as-you-go: when does provisioned capacity make sense? Azure OpenAI vs OpenAI direct: which should you use? What do real Azure OpenAI workloads cost? How do you reduce Azure OpenAI costs? How does CloudZero track and optimize Azure OpenAI spend? Frequently asked questions about Azure OpenAI pricing

Quick Answer

Azure OpenAI pricing has two billing modes: Standard pay-as-you-go per token, and Provisioned Throughput Units (PTUs) at an hourly rate for reserved capacity, with monthly and annual reservations. The GPT-5.6 family (Sol, Terra, Luna) is now live on Azure alongside GPT-5.5, GPT-5.4, and the o-series. Global Standard deployments generally track OpenAI's direct list rates ($0.20 to $5 per million input tokens across the mainstream current models); Data Zone and Regional deployments cost more.

Your Azure OpenAI bill lands at $40,000 this month. Is that a problem? Nobody in the room knows. If that $40,000 powered a feature that drove $200,000 in revenue, the right move is scaling up. If it fed a chatbot nobody uses, the right move is a funeral. Azure Cost Management shows you the cost. It cannot show you which of those two companies you are.

That’s the frame for everything below, because Azure pricing for OpenAI models is genuinely more complicated than the direct API: same models, plus deployment geographies, provisioned capacity math, and reservation terms stacked on top. In CloudZero’s 2026 AI ROI survey of 260 finance leaders, 66% said their boards tie AI funding to returns. Boards don’t accept “it’s complicated” as a line item. So here’s the whole structure.

How does Azure OpenAI pricing work?

Microsoft’s Azure pricing model for OpenAI gives you three ways to pay:

  • Standard (pay-as-you-go). Per-token billing, input and output metered separately, just like the direct API. Flexible, no commitment, and the default for variable workloads.
  • Provisioned Throughput Units (PTUs). You reserve model processing capacity and pay an hourly rate per PTU whether you use it or not. Minimums are 15 PTUs for Global and Data Zone deployments and 50 for Regional (25 for mini models). Monthly and annual reservations cut the hourly rate further. PTUs buy predictable latency and predictable spend, at the price of paying for idle capacity.
  • Batch API. Asynchronous jobs returned within 24 hours at a 50% discount on Global Standard rates. Same discount logic as the direct API, and the first lever to pull for pipelines that can wait.

Then multiply by geography, because every Standard and Provisioned deployment comes in three flavors. Same model, three prices:

Deployment typeRelative costData boundaryUse when
GlobalBaseline (tracks OpenAI list rates)Routed anywhereDefault for most workloads
Data ZoneLists about 10% higherBounded to EU or USRegulatory residency at zone level
RegionalHighestPinned to one of up to 27 regionsStrict local-region requirements

Data residency is a real requirement for plenty of enterprises; it’s also a real multiplier, and it belongs in the forecast, not the postmortem.

One naming note: Microsoft’s product is Azure OpenAI, sometimes written Azure Open AI. The broader Azure AI pricing picture, AI Search, Foundry, and Speech, sits on top of the model rates covered here.

What models are available on Azure OpenAI in 2026?

The catalog moved a lot this summer, and this is where most Azure OpenAI service pricing guides are quietly out of date. As of 2026, Azure offers:

  • The GPT-5.6 family, now live. Sol, Terra, and Luna are all deployable on Azure in Global and Data Zone flavors, with separate short-context and long-context rates and Priority Processing tiers for Sol and Terra. Sol and Terra are also available as PTU deployments. If your Azure catalog knowledge predates July, this is the update that matters.
  • The recent generations. GPT-5.5 (Global, Data Zone, long-context tiers), the GPT-5.4 family including Pro, mini, and nano, GPT-5.3 Codex and Chat, GPT-5.2, GPT-5.1, and the original GPT-5 series, all still deployable, which matters for teams with model-version pins in production.
  • Reasoning and specialty models. o3 and o4-mini (Global, Data Zone, and Regional, with batch discounts), o3 deep research, Sora 2 for video, the GPT-Image series, realtime and audio models, embeddings, fine-tuning for GPT-4.1 and GPT-4o families, and the open-weight gpt-oss-120b.
  • The legacy shelf. GPT-4o, GPT-4.1, and the older Azure OpenAI GPT-4 and GPT-3.5 lines remain listed for existing deployments. If you’re still running these, you’re paying previous-generation rates for previous-generation output, and the migration math almost certainly favors moving.

On rates: Microsoft renders exact dollar figures dynamically by region and agreement, so the reliable reference is OpenAI’s direct list rates, which Global Standard Azure GPT pricing generally tracks for the same models:

Model (Global Standard reference)Input (per MTok)Output (per MTok)Notes
GPT-5.6 Sol$5.00$30.00PTU-eligible; long-context tier above ~272K
GPT-5.6 Terra$2.00$12.00PTU-eligible; reflects OpenAI’s July 30 cut (20%)
GPT-5.6 Luna$0.20$1.20Reflects OpenAI’s July 30 cut (80%)
GPT-5.5$5.00$30.00Long-context tier available
GPT-5.4$2.50$15.00Batch API eligible
GPT-5.4 mini$0.75$4.50Batch API eligible
GPT-5.4 nano$0.20$1.25Cheapest legacy tier
o4-mini$1.10$4.40Reasoning tokens bill as output
o3$2.00$8.00Reasoning tokens bill as output

Azure regional rates vary. Whether the July 30 cut is fully reflected in your Azure region is exactly the kind of thing to confirm in the Azure pricing calculator before you approve a budget, because new models and pricing changes reach the direct API first, and any lag on an 80% cut is real money.

What changed in Azure OpenAI pricing this summer?

Three dated changes worth building into any forecast:

  • Cache write billing starts on or after August 21, 2026. Microsoft’s pricing page currently carries a notice that cache write charges are not yet active and billing is expected to begin on or after August 21. Translation: prompt caching on Azure has been effectively write-free, and that subsidy is ending. Teams that built caching-heavy architectures on Azure should expect a new line item this month (on the direct API, cache writes bill at 1.25x input rates for current models; confirm Azure’s exact multiplier in the calculator). If your August and September Azure bills diverge and nothing else changed, start here.
  • GPT-5.6 arrived with split context pricing. Short-context and long-context tiers are priced separately on Azure, mirroring the direct API’s ~272K threshold where rates roughly double on input. Long-context workloads need the higher meter in the forecast.
  • The direct API got cheaper on July 30. OpenAI cut Luna 80% and Terra 20% on the direct API, leaving Sol unchanged at 5/30. For Azure customers this creates a live arbitrage question: if your region’s rates haven’t matched yet, workloads without residency requirements are temporarily cheaper off Azure, and your enterprise agreement discussion just got a new data point.

PTUs vs pay-as-you-go: when does provisioned capacity make sense?

The PTU decision is a utilization bet. Pay-as-you-go charges for tokens; PTUs charge for time. The break-even question is whether your throughput is consistent enough that reserved capacity beats metered usage. The structure, per Microsoft’s provisioned throughput billing documentation:

DeploymentMinimum PTUsBilling options
Global15Hourly, monthly reservation, annual reservation
Data Zone15Hourly, monthly reservation, annual reservation
Regional50 (25 for mini models)Hourly, monthly reservation, annual reservation

The honest math: at 15 PTU minimums for Global deployments, PTUs only pencil for sustained, predictable, high-volume inference, the kind where you can chart hourly token throughput and it looks like a plateau, not a mountain range. Bursty workloads pay for idle PTU hours. And the newest cheap models complicate the old logic: at Luna-class token rates, some workloads that used to justify provisioned capacity may never reach break-even on PTUs at all, because the metered price fell out from under the reservation.

Where PTUs win: latency-sensitive production at scale, workloads with contractual throughput requirements, and organizations that can commit to monthly or annual reservations, where the reservation discount does the heavy lifting. Microsoft’s own framing is that provisioned throughput delivers predictable performance and cost, which is true precisely when your usage is stable. If it isn’t, you’re paying for control you’re not using.

The uncomfortable part: most teams choose between PTUs and pay-as-you-go on a forecast built from a pilot, and pilots systematically underestimate production usage. Build three cases (pilot-level, planned rollout, and adoption spike), price all three in both modes, and commit to reservations only on the volume you’d bet the budget on.

Azure OpenAI vs OpenAI direct: which should you use?

The Azure OpenAI vs OpenAI question comes down to what you’re buying beyond tokens.

FactorAzure OpenAIOpenAI direct
Token ratesGlobal Standard tracks OpenAI list; Data Zone and Regional cost moreList rates
New models and price cutsLag by regionImmediate
Data residencyData Zone and Regional deploymentsNot available
Reserved capacityPTUs with hourly, monthly, and annual termsNot available
BillingConsolidated on the Microsoft invoiceSeparate OpenAI invoice
Stack integrationFabric, Cosmos DB, AI Search inside the Azure boundaryAPI only

Choose Azure when you need data residency (Data Zone and Regional deployments exist for exactly this), enterprise compliance and networking inside the Azure boundary, integration with the Azure stack (Fabric, Cosmos DB, AI Search), consolidated Microsoft billing, or PTU-style reserved capacity with SLAs. One budgeting note for RAG architectures: Azure AI Search pricing is its own meter, billed separately from model tokens, so a retrieval-augmented feature carries two Azure line items before it answers a single question.

Choose the direct API when you need the newest models and prices first (the July 30 cut reached the direct API immediately; Azure regions lag), simpler billing, or you’re pre-enterprise and paying list either way. The full direct-side breakdown lives in CloudZero’s OpenAI API pricing guide.

In practice, most enterprises run both, which is where Azure OpenAI costs get genuinely hard to manage: the same model, GPT-5.6 Terra, can appear on a Microsoft invoice (as Azure OpenAI, priced by deployment type) and an OpenAI invoice (as API usage) in the same month, for the same product feature. Two bills, two formats, one workload, zero shared labels. Add Claude on AWS Bedrock and you have the standard 2026 enterprise AI stack: three invoices describing one feature, none of them saying so.

This is the shift CloudZero founder and CTO Erik Peterson keeps pushing: stop treating these as mere expenses and start treating them as investments that need to show a return, which requires knowing what each one returned. Full model-side economics live in CloudZero’s Claude pricing guide for the Anthropic half of that stack.

What do real Azure OpenAI workloads cost?

Worked example at Global Standard rates tracking OpenAI list prices. A support assistant handling 10,000 conversations daily, 500 input and 300 output tokens each (150M input, 90M output monthly):

Model (Global Standard)Monthly cost
GPT-5.6 Sol~$3,450
GPT-5.6 Terra~$1,380
GPT-5.4 mini~$518
GPT-5.6 Luna~$138

Now apply the Azure adjustments the sticker math skips: Data Zone adds roughly 10% if you have residency requirements. Production deployments typically land 20% to 40% above listed token rates once networking, monitoring, support plans, and fine-tuned model hosting join the party. And from August 21, caching-heavy architectures pick up cache-write charges that were previously absent. The listed token rate is the floor of your Azure OpenAI price, not the estimate.

One more forecast trap specific to Azure: the metered-vs-provisioned crossover moves when token prices move. A workload that justified a PTU commitment against last quarter’s Terra rates may no longer justify it against the post-cut rates, because the metered alternative got 20% cheaper while the reservation didn’t. Reservations are a bet on price stability as much as usage stability, and 2026 is not offering price stability. Re-run the crossover math before every renewal, not just the first commitment.

For the metered-vs-provisioned crossover and your own token mix, run scenarios in CloudZero’s Azure cost calculator guide alongside Microsoft’s calculator.

How do you reduce Azure OpenAI costs?

Six levers, roughly in order of how much they move the bill:

  1. Re-tier for the new catalog. If production still runs GPT-5.4 or GPT-4o-era models, the 5.6 family offers better economics at most tiers. Legacy pins are the most common silent overspend on Azure.
  2. Batch everything that can wait. 50% off Global Standard for 24-hour turnaround. Nightly pipelines have no excuse.
  3. Question every Data Zone and Regional deployment. Residency requirements are real; reflexive region-pinning is not. Every deployment above Global should have a compliance reason attached, because it carries a permanent surcharge.
  4. Audit caching before August 21. Confirm your cache hit rates justify the architecture once writes start billing. High-hit caches still win comfortably; low-hit caches are about to become a fee.
  5. Right-size the PTU bet. Review provisioned utilization monthly. Idle PTU hours are the Azure OpenAI equivalent of unattached storage volumes: invisible, recurring, and nobody’s fault, which is the problem.
  6. Give every deployment an owner. Azure makes it one click to spin up a new OpenAI resource, which is how experimentation becomes 40 untagged deployments. Ownership plus visibility is the prerequisite for the other five levers, and it’s the one native tooling doesn’t provide.

How does CloudZero track and optimize Azure OpenAI spend?

Azure Cost Management can show you what Azure OpenAI cost last month, by resource and subscription. It cannot tell you which product feature or customer drove the tokens, whether the PTU reservation is earning its commitment, or how the Azure share compares with the direct API share of the same workload. Cost, yes. Meaning, no.

CloudZero closes that gap in four moves:

Allocate 100% of AI spend, no tags required. CloudZero’s Azure integration ingests all Azure spend including Azure OpenAI, and the allocation engine attributes every dollar to who’s responsible for it. That matters on Azure specifically, because the deployments nobody tagged are usually the experimental ones, and the experimental ones are where AI spend hides.

Analyze costs per AI service. CloudZero breaks AI spending down by type of service, SDLC stage, and AI model development stage, using custom Dimensions like cost per project, cost per model, and cost per user, trended over time. So “what does Azure OpenAI cost” becomes “what does the support copilot cost per resolved ticket, and is that trending the right way.”

Connect spend to ROI. Allocated costs map to unit economics, cost per customer, per feature, per team, which is how you calculate the return on every dollar of AI investment instead of just the size of it. Pete Rubio, SVP of Platform and Engineering at Rapid7, puts it this way: “It’s not just about saving money; it’s about enabling innovation while maintaining financial accountability and control,” which is the whole assignment in one sentence.

Budget, forecast, and catch the surprises. Budgets and forecasting turn allocated AI spend into forward numbers a board can hold you to, and anomaly detection comparing the last 36 hours against 12 months of history flags the moment a new deployment, or a billing change like cache writes going live on August 21, pushes spend outside normal, with hour-level data routed to the engineer who owns it.

The multi-invoice problem gets the same treatment: Azure OpenAI lands in the AI Hub next to direct OpenAI usage, Anthropic, AWS, and GCP, one normalized view of what each AI feature costs regardless of which invoice it arrived on.

That’s the difference between knowing your Azure OpenAI bill and knowing your AI unit economics: one is a number, the other is a decision. It’s also the answer to the $40,000 question this article opened with, which was never “is the bill too big” but “what did it buy.”

The receipts are public: Upstart saved $20 million and Drift cut $2.4 million from their AWS bill using CloudZero, and Rapid7 scaled its AI initiatives with the accountability quoted above. More on the customer stories page. Schedule a demo to see your Azure OpenAI spend the way your board wants it explained.

Frequently asked questions about Azure OpenAI pricing