Claude Code token limits explained: Managing AI coding costs

Understand Claude Code's context window and usage limits, what really drives token costs, and how to manage AI coding spend by tying usage to engineering ROI.

Chart of Claude Code's average estimated cost per commit based on used tokens

Claude Code token limits explained: Managing AI coding costs

Understand Claude Code's context window and usage limits, what really drives token costs, and how to manage AI coding spend by tying usage to engineering ROI.

Chart of Claude Code's average estimated cost per commit based on used tokens
Chapters

Published December 04, 2025 · Updated September 3, 2026

What are Claude Code’s token limits?

Earlier in 2026, Anthropic began describing Claude Code’s token limits in more relative terms rather than as fixed token counts. Claude has two types of limits: length limits and usage limits. The key distinction is that length limits determine how long a single conversation can become, whereas usage limits determine how much you can use Claude overall across your conversations. In other words, length limits measure the size and complexity of one conversation, while usage limits measure total activity over time. 

Claude Code length limits

Claude’s context window size is 200K tokens across all models and paid plans, except for Enterprise plans, which have a 500K context window on some models. Once a conversation or codebase exceeds that window, Claude may lose access to earlier details, which can make long debugging sessions, large refactors, and multi-file projects harder to manage.

Claude Code usage limits

Claude Code usage is governed by more than one limit. For Pro and Max users, the first is a 5-hour session limit. According to Anthropic’s Max plan documentation, Max 5x provides 5x Pro usage per session, while Max 20x provides 20x Pro usage per session. The session limit resets every five hours.

Paid plans also have weekly usage limits layered on top of the 5-hour session limit. Anthropic currently tracks weekly usage separately for Opus and all other models, with each resetting at a fixed time assigned to the account. Usage across Claude, Claude Code, and other Claude surfaces draws from the same plan limits, as outlined in Anthropic’s usage limit guidance.

Anthropic’s 5x and 20x plan labels refer to usage within the 5-hour session window. Weekly capacity is separate and does not appear to scale at the same rate. Based on Anthropic’s previously published usage estimates and real-world findings, the approximate relationship looks like this:

Claude Plan 5-Hour Session Capacity Approx. Weekly Capacity vs. Pro
Pro 1x baseline 1x baseline
Max 5x 5x Pro ~3.5x Pro
Max 20x 20x Pro ~6x Pro
Claude Pro and Max 5x and 20x capacity across 5-hour sessions and weekly usage limits

Anthropic does not currently publish fixed weekly multipliers for Pro, Max 5x, and Max 20x. Recently, the discrepancy between the 5-hour multipliers and weekly capacity has drawn attention in the Claude community; see this discussion on X. Actual usage also varies based on factors including model choice, conversation length and complexity, features used, and effort level.

Enterprise clients have a different usage model. On Anthropic’s current usage-based Enterprise plan, the seat fee covers access to Claude, Claude Code, and Cowork, while usage is billed separately based on actual token consumption at standard API rates. Unlike Pro, Max, Team, and legacy seat-based Enterprise plans, usage-based Enterprise has no included token allowance or per-seat usage limits. Admins can instead control consumption by setting spend limits at the organization and individual user levels.

How Claude Pro, Max 5x, and Max 20x usage limits compare

The difference between session limits and weekly capacity limits becomes even clearer when you look at each upgrade step directly. Moving up a tier increases 5-hour session capacity much faster than it increases weekly capacity:

Upgrade Increase in 5-Hour Capacity Approx. Increase in Weekly Capacity
Pro → Max 5x 5x ~3.5x
Max 5x → Max 20x 4x ~1.7x
Pro → Max 20x 20x ~6x
How upgrading between Claude plans changes 5-hour and weekly usage capacity limits

The biggest gap is between Max 5x and Max 20x: the upgrade provides 4x more capacity within a 5-hour window, but only about 1.7x more weekly capacity based on Anthropic’s historical estimates.

Anthropic can also impose additional weekly, monthly, model, or feature-specific limits to manage capacity. Because these limits and the way usage is metered can change, the figures above should be treated as approximate rather than fixed allowances.

What happens when you hit your Claude Code usage limit?

Hitting a Claude Code usage limit doesn’t necessarily mean you have to stop working. What happens next depends on your plan and how Claude Code is configured.

For Pro and Max users, you can wait for your usage limit to reset, upgrade to a higher-usage plan, or enable usage credits to continue working beyond your plan’s included allowance. Once enabled, additional usage is charged at consumption-based rates. Users can also switch to pay-as-you-go usage through a Claude Console account for more intensive coding workloads.

For Team and seat-based Enterprise plans, organizations can enable usage credits so developers can continue working after reaching their included limits. Usage-based Enterprise plans work differently: there are no per-seat usage limits, and consumption is billed at API rates.

For engineering leaders, this means hitting a Claude Code limit is increasingly a cost-management issue rather than simply an access issue. Teams may be able to keep coding past their included allowances, but doing so can introduce variable spend that needs to be tracked and governed.

How different models affect Claude Code token limits

Claude Code usage depends on several factors, including the length and complexity of your conversations, the features you use, and your selected model and effort settings. Model choice directly affects how quickly Claude Code usage is consumed. Claude Code model pricing is based on input and output tokens, as summarized in the following table:

Claude Code Model Current Model Tier Input Token Price Output Token Price Total Cost for 1M Input + 1M Output Relative Cost Across Model Tiers Best For
Claude Opus Opus 5 $5 / 1M tokens $25 / 1M tokens $30 5x Haiku Complex reasoning, large codebase work, high-autonomy agentic coding
Claude Sonnet Sonnet 5 $2 / 1M tokens $10 / 1M tokens $12 2x Haiku Everyday Claude Code use, refactoring, debugging, balanced speed and quality
Claude Haiku Haiku 4.5 $1 / 1M tokens $5 / 1M tokens $6 1x baseline Lower-cost tasks, fast iterations, simpler coding assistance
Claude Code model tiers, token pricing, relative costs, and recommended use cases.

Across all three models, output tokens are the bigger cost driver, with each model’s output tokens costing 5x more than its input tokens. And, for the same number of input and output tokens, Sonnet costs 2x more than Haiku, while Opus costs 5x more than Haiku. Practically speaking, that means heavy use of Opus will exhaust your Pro/Max allocation much faster than Sonnet or Haiku usage. If you’re running complex, multi-file agentic workflows with Opus, you'll hit your limits much sooner than you might expect.

A note on comparing token usage across models: Sonnet 5 introduced an updated tokenizer, which means the same input can translate into more tokens than it did with Sonnet 4.6. Anthropic estimates that the same input can map to roughly 1.0–1.35x as many tokens, depending on the content type. That means raw token counts aren't necessarily an apples-to-apples comparison across model generations, even when the underlying workload stays the same.

How effort levels affect Claude Code token usage

Model choice isn’t the only factor that determines how quickly you consume Claude Code usage. Reasoning effort also matters. Claude Code lets users adjust how much computational effort Claude applies to a task, trading off deeper reasoning against latency and token consumption.

Higher effort levels can improve performance on complex coding and agentic tasks, but they also consume more tokens and can cause users to hit usage limits faster. Lower effort levels can be more efficient for simpler tasks where extended reasoning isn’t necessary. Anthropic describes this as a tradeoff between more thinking and lower latency and fewer usage-limit hits.

For engineering teams, that means understanding Claude Code consumption increasingly requires looking at both the model and the effort level being used. Two developers using the same model for similar workloads can consume meaningfully different amounts of their usage allowance depending on how much reasoning effort they apply.

How advanced features affect Claude Code token limits

Claude has numerous types of advanced features that can greatly increase token usage. There are two worth noting: 

Agent Teams: In February 2026, Anthropic released Agent Teams. This multi-agent capability is now a built-in part of Claude Code, and it can significantly increase the number of tokens software engineers use during a session. Agent teams run multiple Claude Code instances at once, with each instance maintaining its own context window. As a result, token consumption grows based on how many teammates are active and how long they continue running. Anthropic notes that agent teams can consume about 7x more tokens than standard sessions when teammates operate in plan mode.

Dynamic Workflows: In May 2026, Anthropic released dynamic workflows (for those on Claude Enterprise plans), and they became available and turned on by default on June 8, 2026. Dynamic workflows can further expand token consumption by turning a single request into a scripted, multi-agent execution. Instead of Claude handling the task turn by turn in one conversation, a workflow can fan work out across dozens or even hundreds of subagents, each performing its own model calls and tool use. Anthropic notes that workflow runs can use meaningfully more tokens than completing the same task through a standard conversation, and those runs count against the organization’s usage and rate limits. 

How to reduce Claude Code token usage

Because Claude Code’s token usage scales with the amount of context it processes, keeping that context focused can help developers get more from their usage limits. Anthropic recommends several ways to reduce unnecessary token consumption:

  • Clear context between unrelated tasks. Use /clear when moving to a new task so Claude doesn’t continue processing irrelevant conversation history with every subsequent message.
  • Compact long-running conversations. Use /compact to summarize the conversation while preserving the information needed to continue working. Claude Code also automatically compacts conversations as they approach the context limit.
  • Use the right model for the task. Sonnet is suitable for most coding tasks, while Opus can be reserved for work that requires more complex reasoning. Simpler tasks can also be delegated to Haiku-powered subagents.
  • Watch what’s consuming context. The /context command shows what is taking up space in the current context window, making it easier to identify oversized instructions, tools, or other sources of unnecessary context.
  • Limit unnecessary tool context. Unused MCP servers can add to context consumption. Anthropic recommends disabling servers you aren’t actively using and using CLI tools where appropriate.

These practices can help developers stretch their Claude Code usage further, but optimizing for fewer tokens shouldn’t be the goal in isolation. For engineering organizations, the more useful question is whether the tokens being consumed are producing valuable outcomes, such as completed work, merged PRs, and faster delivery.

Claude Code token limits: What engineering leaders should know about AI coding costs

AI coding tools like Claude Code are more widely used in software development than ever—and costs have climbed just as fast. Yet, that spend remains hard to manage: consumption-based pricing is unpredictable, actual limits are opaque, and the link between AI usage and engineering outcomes is murky.

Anthropic also provides native analytics for tracking Claude Code usage, contribution, and cost, including sessions, token consumption by model, commits, pull requests, and estimated cost per user. These metrics are useful for understanding adoption and spend, but they don't show what happens to AI-assisted work after it leaves the tool. For a deeper look at the available data, ingestion options, and where tool-level telemetry stops, read our article on Claude Code analytics.

To see what your organization's AI spend is actually producing, start with The Field Guide to Measuring Token Efficiency in AI Engineering, which lays out the metrics worth tracking so you can make decisions grounded in your own data. From there, see how Token Intelligence traces AI token consumption to what it delivers across your people, teams, and outcomes—so you know what's productive, what's wasteful, and what to fix.

Frequently asked questions about Claude Code token limits

How much does Claude Code cost per developer?

Claude Code pricing depends on how it is accessed. Claude Pro costs $20/month, Max 5x costs $100/month, and Max 20x costs $200/month, with each subscription including usage subject to 5-hour and weekly limits. API and usage-based Enterprise customers are instead charged based on token consumption at Anthropic’s applicable model rates.

How many tokens do you get with Claude Pro vs. Max 5x?

Claude Pro costs $20/month, while Claude Max 5x costs $100/month. Anthropic does not publish a fixed token allowance for either plan. Instead, Max 5x provides 5x the usage capacity of Pro within each 5-hour session. Both plans also have separate weekly limits, and Anthropic does not currently publish a fixed weekly multiplier between Pro and Max 5x.

How many tokens do you get with Claude Pro vs. Max 20x?

Claude Pro costs $20/month, while Claude Max 20x costs $200/month. Anthropic does not publish a fixed token allowance. Max 20x provides 20x the usage capacity of Pro within each 5-hour session. The 20x figure applies to the 5-hour session limit, not total weekly usage, which is governed by separate weekly limits.

How many tokens do you get with Claude Max 5x vs. Max 20x?

Claude Max 5x costs $100/month and Claude Max 20x costs $200/month. Max 20x provides 4x the 5-hour session capacity of Max 5x because the tiers provide 20x and 5x Pro capacity, respectively. Weekly capacity also increases on Max 20x, but Anthropic does not currently publish the exact weekly multiplier between the two plans.

What is the Claude Code context window size?

Claude Code supports a 200K-token context window on standard paid models, with larger context windows available in some Enterprise configurations. The context window determines how much information Claude can process at once and is separate from the 5-hour and weekly usage limits that determine how much Claude you can use over time.

Does Claude Code have weekly usage limits?

Yes. Claude Code has weekly usage limits in addition to its 5-hour session limits. Anthropic’s usage limit guidance currently tracks weekly usage separately for Opus and all other models. These weekly limits reset independently of the 5-hour session window and can prevent additional usage even if a new 5-hour session has begun.

What happens when you hit your Claude Code usage limit?

When you hit a Claude Code usage limit, you generally need to wait for the relevant 5-hour or weekly limit to reset, use another model if capacity remains available, or enable additional paid usage where your plan supports it. The exact options depend on which Claude plan you use and which limit you reached.

What is Claude Code's pricing per million tokens?

Claude Opus 5 costs $5 per 1M input tokens and $25 per 1M output tokens. Claude Sonnet 5 costs $2 per 1M input tokens and $10 per 1M output tokens. Claude Haiku 4.5 costs $1 per 1M input tokens and $5 per 1M output tokens. Opus is best suited to complex reasoning and large-codebase work, Sonnet to everyday coding, refactoring, and debugging, and Haiku to faster, lower-cost tasks.

Why does Opus use Claude Code limits faster than Sonnet?

Opus can consume Claude Code usage limits faster because it is more computationally intensive than Sonnet. Anthropic also tracks Opus against a separate weekly usage limit, so heavy Opus usage can reach that model-specific cap even when capacity remains available for other models.

What metrics should you track to manage Claude Code spend?

To manage Claude Code spend, track cost and usage by developer, team, model, and workflow, then connect that consumption to engineering outcomes such as merged pull requests, cycle time, deployments, and incidents. Token usage shows how much AI is being consumed; outcome-level metrics show whether that spend is translating into useful engineering work.

Do Claude Code Agent Teams use more tokens?

Yes, significantly. Agent Teams (released February 2026) run multiple Claude Code instances at once, each with its own context window, so consumption scales with how many teammates are active. Anthropic notes Agent Teams can use about 7x more tokens than standard sessions when teammates run in plan mode.

What are Claude Code dynamic workflows and how do they affect token usage?

Dynamic workflows (Enterprise plans, on by default since June 8, 2026) turn a single request into a scripted, multi-agent execution that can fan work across dozens or hundreds of subagents, each making its own model calls. They use meaningfully more tokens than the same task in a standard conversation, and runs count against the org's usage and rate limits.

Did Anthropic remove Claude Code’s peak-hour limits?

Yes, for Pro and Max users. In March 2026, Anthropic temporarily reduced effective Claude Code session limits during peak weekday hours. On May 6, Anthropic removed that peak-hours reduction for Pro and Max accounts. It also doubled Claude Code’s 5-hour rate limits for Pro, Max, Team, and seat-based Enterprise plans. Weekly and other usage limits can still apply.

What is the current Claude Code plan limit controversy?

The Claude Code plan limit controversy centers on the difference between the Max 5x and Max 20x plan names and Claude’s separate weekly usage limits. The 5x and 20x multipliers apply to capacity within a 5-hour session, not total weekly usage. Because Anthropic does not publish equivalent weekly multipliers, some users argue that the $100 Max 5x and $200 Max 20x plans can appear to scale overall usage more than they actually do.

Thierry Donneau-Golencer

Thierry Donneau-Golencer

Thierry is Head of Product at Faros, where he builds solutions to empower teams and drive engineering excellence. His previous roles include AI research (Stanford Research Institute), an AI startup (Tempo AI, acquired by Salesforce), and large-scale business AI (Salesforce Einstein AI).

Graduation cap with a tassel over a dark gradient background.
AI ENGINEERING REPORT 2026
The Acceleration â€ĻWhiplash
The definitive data on AI's engineering impact. What's working, what's breaking, and what leaders need to do next.
  • Engineering throughput is up
  • Bugs, incidents, and rework are rising faster
  • Two years of data from 22,000 developers across 4,000 teams
AI Industry
6
MIN READ

An open source team barred AI code. Our data shows a better fix.

Restrictions on AI-generated code are spreading across open source projects as review queues overflow. Our Speed Trap report shows what that costs teams and what to do instead.

AI Industry
7
MIN READ

What is an AI-native engineering organization?

AI-native engineering organizations build software delivery around AI agents. See the six characteristics that separate AI-native from AI-assisted software development.

AI Industry
4
MIN READ

Your AI bill doesn't tell you what you think it does

Six webinar takeaways from Faros CEO Vitaly Gordon on measuring AI spend by cost per outcome, setting smarter quotas, and choosing models using your own code.