Skip to content

How to Stop Running Out of Codex Tokens in 2 Days

Problem

I have a Codex Pro subscription. For a while, I kept hitting my usage limit by Tuesday.

That left me with a frustrating choice for the rest of the week: slow down, switch tools, or wait for the allowance to reset.

My first assumption was simple: I needed more quota.

But after looking at how I actually used Codex, I realized that a large part of the problem was my workflow.

I was giving Codex too much context, asking vague questions, running work in parallel before validating the previous step, and sometimes asking it to generate code that already existed in mature open-source projects.

The biggest lesson was this:

The biggest Codex usage killer isn’t code generation itself. It’s unnecessary context, unnecessary exploration, and unnecessary rework.

Once I changed those habits, the same subscription started lasting much longer.

What Other Codex Users Were Seeing

A Reddit comment first made me look more closely at my workflow:

“It boggles my mind how people can use up the Codex Plus plan in 1 day… having 50 repos open, multiple worktrees, dozens of agents running continuously” — Reddit user (Score: 5)

That wasn’t exactly what I was doing, but the pattern was familiar: too much work happening at once, too much context, and too little control over what the agent actually needed to inspect.

Another user described something much closer to my experience:

“Token usage has been insanely bad for me the last two weeks. I’m hitting the weekly limit on pro in two days of fairly gentle usage” — Reddit user (Score: 3)

The important point is not whether those exact workloads match yours. The useful question is:

What work is Codex doing that never needed to happen?

That is where I found most of my savings.

How Codex Usage Works in 2026

One important correction to my original understanding: Codex usage is not simply a fixed bucket of “weekly tokens.”

Depending on your plan, Codex can be constrained by both a five-hour usage window and a weekly usage window. You need remaining allowance in the applicable windows to continue working.

The amount of useful work you get from that allowance also varies with:

  • the model you choose
  • reasoning settings
  • task complexity
  • how much repository context the agent needs
  • how long the task runs
  • how many tools and files it needs to inspect

OpenAI describes the published per-window figures as estimates rather than fixed message limits. In practice, a tiny mechanical edit and a cross-repository debugging session can consume very different amounts of allowance.

So I no longer try to model my plan as:

Oversimplified mental model
weekly_limit = 500000_tokens

Instead, I use the product’s Settings → Usage view and Codex’s /status command when available to check the current session, context, and limits.

That also means the token numbers later in this article are my workload estimates, not OpenAI’s billing formula.

Where My Usage Was Going

Codex token usage breakdown showing that context reading, reasoning, testing, and debugging can consume more tokens than the final code change

Here is how I originally thought about a typical “simple” task:

Conceptual token consumption breakdown
Task: "Fix the bug in auth.py"
├─ Read auth.py ~5,000 tokens
├─ Read related imports ~9,000 tokens
├─ Analyze codebase structure ~2,000 tokens
├─ Generate fix ~1,000 tokens
├─ Write edited file ~500 tokens
└─ Context / reasoning overhead ~1,000 tokens
Estimated total ~18,500 tokens

The exact number varies by model, tokenizer, tool behavior, caching, and how much context Codex actually reads. The useful part of this example is the shape of the cost.

The final code change might be tiny.

The expensive part can be everything Codex does before and around that change:

discover files
→ read files
→ understand dependencies
→ form a hypothesis
→ edit
→ test
→ reread
→ debug

In agentic coding, input context and repeated reasoning can cost much more than the few lines of code you actually wanted changed.

That changed how I optimize Codex.

I now focus on three levels:

  1. Reduce context
  2. Reduce unnecessary generation
  3. Reduce rework

Here are the changes that mattered most.

1. Reduce Context Before Optimizing Anything Else

My first rule used to be:

One project at a time.

I still follow it, but I would phrase the reason more carefully now.

Simply having a repository visible in an editor does not necessarily mean every file is automatically sent to the model. The real problem is giving Codex a workspace or instruction that encourages it to inspect far more than the task requires.

The expensive version of a prompt looks like this:

Too broad
Find why authentication sometimes fails.
Look through the project and fix it.

That gives the agent permission to explore widely.

I now give it an explicit starting scope:

Scoped investigation
Investigate the login failure.
Start with:
- src/auth/login.ts
- src/auth/session.ts
Do not inspect unrelated directories unless these files point to them.
First explain the likely cause.
Do not edit code yet.

This does two useful things:

  • limits unnecessary repository exploration
  • separates diagnosis from implementation

For large monorepos, I also avoid casually asking Codex to “understand the whole project” unless that is actually required.

Generated files, build output, vendored dependencies, huge logs, and unrelated packages should not become part of the investigation by default.

My rule now is:

Scope first, code second.

2. Search GitHub Before Writing Thousands of Lines From Scratch

Comparison of building a project from scratch with Codex versus reusing and customizing an existing open-source GitHub project

This became one of my biggest workflow changes.

When I started a new project, I used to ask something like:

Build me an admin dashboard.

or:

Create an AI knowledge base with authentication, RAG, file upload,
chat history, an admin panel, and Docker deployment.

Codex could do a lot of that work, but it also meant generating a large amount of infrastructure that already existed elsewhere.

And the real cost was not just the first generation.

Every generated module could later need to be:

generated
→ read again
→ modified
→ tested
→ debugged
→ reread

So I now search for a mature open-source base before asking Codex to implement common infrastructure.

For example, an AI knowledge-base project may require:

  • file upload and parsing
  • embeddings
  • vector search
  • RAG
  • chat UI
  • authentication
  • user management
  • admin tools
  • deployment configuration

If an existing project already covers 70–80% of that, my Codex workflow can become:

Reuse-first workflow
search
→ evaluate
→ fork
→ remove unnecessary features
→ adapt business logic
→ connect my services
→ deploy

instead of:

Build-everything workflow
design everything
→ generate everything
→ debug everything
→ maintain everything

The Prompt I Use Before Starting a New Project

Open-source-first prompt
I want to build XXX.
Do not write code yet.
First, search for existing open-source projects that could be used
directly or adapted for this requirement.
Evaluate:
1. Is the project actively maintained?
2. How recent are its commits and releases?
3. Are Issues and PRs active and healthy?
4. How difficult is deployment?
5. Does the technology stack fit my project?
6. Which features can be reused directly?
7. Which modules would require modification?
8. Is the license suitable for my intended use?
Then recommend one of:
- use directly
- fork and customize
- build from scratch
Finally, propose the smallest MVP.
Wait for confirmation before writing code.

I do not choose a project purely because it has the most stars.

I check:

  • recent commits
  • releases
  • open issues
  • PR activity
  • dependency freshness
  • license
  • deployment quality
  • stack fit
  • how much code I can realistically keep

The key idea is simple:

Less code generated today means less code Codex has to read, modify, and debug tomorrow.

3. Use the Cheapest Model That Can Reliably Finish the Task

Another mistake I made was treating “strongest model” as the default for every task.

That is convenient, but it is not always efficient.

Current Codex model availability and limits change over time, so I avoid hard-coding a permanent model hierarchy into my workflow. But the principle is stable:

Use the least expensive model that can reliably complete the task.

For example, lighter/faster models such as Luna-class options are often a better fit for:

  • formatting
  • extraction
  • repetitive edits
  • simple transformations
  • documentation cleanup
  • straightforward code changes

A mid-tier coding model such as Terra is a more natural default for:

  • normal bug fixing
  • tests
  • common feature work
  • day-to-day repository changes

I reserve stronger models such as Sol or Astra-class options for tasks such as:

  • architecture decisions
  • difficult cross-module debugging
  • ambiguous failures
  • large migrations
  • race conditions
  • complex refactoring

The same applies to reasoning effort.

This task:

Rename a field across 12 files and update the affected tests.

usually does not need the highest reasoning setting.

This one might:

Find the root cause of an intermittent race condition
across three services and propose the smallest safe fix.

Model routing is one of the easiest ways to make a limited allowance go further.

4. Plan First, Then Let Codex Touch the Code

One of the most wasteful workflows is letting the agent start editing before you know whether its interpretation is correct.

I now use a plan-first prompt for anything non-trivial:

Plan-first prompt
Do not modify files yet.
First:
1. Identify the likely root cause.
2. List the files that need to change.
3. Propose a solution in no more than 5 steps.
4. Identify possible side effects.
Wait for approval before implementation.

If the plan is wrong, I have spent a small amount of allowance discovering that.

That is much cheaper than:

scan repository
→ change 10 files
→ run tests
→ discover bad assumption
→ undo
→ try again

When available in my Codex host, /plan makes this workflow even easier.

My preferred sequence is:

Controlled workflow
Understand
↓
Plan
↓
Approve
↓
Implement
↓
Test

not:

Expensive workflow
Prompt
↓
Agent starts changing everything
↓
Discover wrong assumption
↓
Undo
↓
Retry

Planning is not overhead when it prevents an entire wrong implementation.

5. Keep AGENTS.md Small and Local

AGENTS.md is useful because Codex can automatically load project instructions from the repository hierarchy.

That also means it should contain high-value instructions—not a history of every decision ever made.

My root AGENTS.md now focuses on five things:

Project
Commands
Constraints
Definition of Done
Permissions

For example:

AGENTS.md
# Project
Next.js + TypeScript + PostgreSQL
# Commands
npm test
npm run lint
# Constraints
- Do not modify migrations without approval
- Do not add dependencies unless necessary
- Preserve public APIs
# Definition of Done
- tests pass
- lint passes
- no unrelated changes
# Permissions
Ask before deleting files or changing deployment config.

I avoid putting temporary task details into the root instructions, such as:

  • today’s bug description
  • meeting notes
  • long architecture history
  • temporary migration plans
  • thousands of lines of coding policy

For a monorepo, local instruction files are much cleaner:

Layered AGENTS.md
/
├── AGENTS.md
├── frontend/
│ └── AGENTS.md
└── backend/
└── AGENTS.md

Codex reads instructions from the repository root toward the current working directory, so more local rules can specialize the broader rules.

That means a CSS task does not need to carry every backend database rule with it.

6. Split Reasoning-Heavy Work, Batch Repetitive Work

At first, these two recommendations sounded contradictory:

  • break large tasks into smaller tasks
  • batch similar tasks together

They are both correct.

The distinction is reasoning vs repetition.

Split work when each stage needs validation

For example:

Reasoning-heavy workflow
architecture
→ validate
implementation
→ validate
tests
→ validate
documentation

I do not want Codex writing tests and documentation for an implementation that I have not reviewed yet.

Batch work when the operations are mechanically similar

For example:

Good batching candidates
- rename the same API field in 12 files
- update 15 imports
- add the same validation rule to several endpoints
- convert a set of similar functions

Previously I might have sent 10 separate requests:

Inefficient
Request 1 → load context → edit function A
Request 2 → load context → edit function B
Request 3 → load context → edit function C
...

Now I batch the related mechanical work when the same context applies.

In my own testing, a group of changes that might have cost roughly 50,000 tokens as separate tasks sometimes fell to around 30,000 when batched. That is an observation from my workflow, not a guaranteed Codex savings rate.

My rule is:

Batch repetitive work. Separate reasoning-heavy work.

7. Compact or Reset Stale Context

Long-running Codex conversations are useful because the agent remembers what has already happened.

But old context eventually becomes baggage.

I watch for a few signs:

  • Codex starts referring to an old requirement that no longer applies
  • the current task has drifted far from the original task
  • a small change is carrying dozens of turns of history
  • /status shows that context has grown significantly

When that happens, I decide whether I need to keep the history.

Useful options in supported Codex environments include:

  • /status — inspect session/context/limits
  • /compact — compress a long conversation
  • /fork — start a new main chat from the current history
  • a clean new thread — when the old history no longer matters

I do not reset the conversation after every small edit, because then Codex may have to rediscover important project state.

The goal is not “minimum context at all times.”

It is:

Keep useful state. Remove obsolete state.

8. Stop Repeating Failed Attempts

One of the easiest ways for an agent to burn allowance is to get stuck in a loop:

attempt
→ fail
→ small variation
→ fail
→ another small variation
→ fail

I now use a one-retry rule.

The first failure is useful: it gives us new evidence.

A second attempt is reasonable if that evidence changes the hypothesis.

But if the same type of failure happens again, I stop execution and switch back to diagnosis.

I use a prompt like this:

Stop repeated retries
Stop changing code.
Explain:
1. What we know
2. Which assumptions may be wrong
3. What evidence is missing
4. What diagnostic command should be run next
5. What alternative approach exists
Do not retry the same solution.

My shorthand is:

No new evidence, no new retry.

When I Don’t Use Codex

Another optimization is simply recognizing tasks that do not need an agent.

I usually handle these with my IDE, CLI, or a lighter tool when possible:

  • simple search-and-replace
  • obvious renames
  • JSON formatting
  • deterministic formatting
  • generated boilerplate
  • trivial documentation cleanup
  • grep, sed, or IDE refactors

Codex CLI can also run with local open-source providers through --oss; OpenAI currently documents local providers such as Ollama and LM Studio.

That makes local models useful for some high-volume, low-risk work where I do not need the strongest reasoning model.

For example:

codex --oss --local-provider ollama

I would rather save my included Codex allowance for work where repository-aware reasoning is actually valuable.

My Low-Usage Codex Prompt Template

This is the structure I use for many debugging and maintenance tasks:

Low-usage task template
Task:
Fix [specific problem].
Scope:
Only inspect:
- file A
- file B
Do not inspect unrelated directories unless required.
Constraints:
- preserve existing API
- do not add dependencies
- do not refactor unrelated code
Process:
1. Diagnose the problem.
2. Explain the root cause.
3. Propose the smallest fix.
4. Wait for confirmation before editing.
Validation:
Run [specific test command].
Stop if:
The same approach fails twice.

The important fields are:

Task
Scope
Constraints
Process
Validation
Stop condition

The more clearly I define those, the less room Codex has to spend time exploring something I did not ask for.

My 30-Second Codex Checklist

Efficient Codex workflow showing scoped tasks, model selection, planning, implementation, validation, and context cleanup

Before I send a non-trivial task, I now ask:

  • Can I do this faster manually?
  • Does GitHub already have a mature implementation?
  • Am I using an appropriate model for the task?
  • Is the task scope specific?
  • Does Codex really need to inspect the whole repository?
  • Should I ask for a plan before editing?
  • Can repetitive edits be batched?
  • Is this conversation carrying obsolete context?
  • Am I repeating a failed approach without new evidence?

It takes less than a minute and prevents a surprising amount of wasted work.

Results After One Month

After changing my workflow, these were my observed results:

MetricBeforeAfter
Days until I typically hit the limit26-7
Estimated tokens per typical task~20,000~8,000
Parallel agents3-41
Weekly overrunFrequentRare / none in the measured period

These are not official Codex benchmarks, and they are not a promise that everyone will save the same amount.

They reflect my workload and my usage habits.

For my workflow, the estimated usage per typical task fell by roughly 60%.

The important part was not a single magic setting. It was eliminating unnecessary work before the model started doing it.

8 Ways I Accidentally Burned Codex Usage

1. Asking Codex to inspect the entire repository

Broad instructions encouraged broad exploration.

2. Using vague prompts

“Make authentication better” requires far more interpretation than a precise change request.

3. Using the strongest model for trivial work

A mechanical edit does not always need the most capable model or highest reasoning setting.

4. Running too many agents in parallel

Parallelism is useful when tasks are truly independent. I was using it before validating whether the first direction was correct.

5. Building common functionality from scratch

Authentication, dashboards, RAG shells, admin panels, and SaaS foundations often have mature open-source starting points.

6. Carrying stale context across unrelated tasks

Useful memory slowly turned into irrelevant baggage.

7. Letting an agent retry the same failure

Without new evidence, repeated retries were mostly repeated spending.

8. Using Codex for trivial work

Some changes are faster and cheaper with an IDE command or a few lines of shell.

Final Takeaway

I originally thought the answer was to make Codex “use fewer tokens.”

That framing was too narrow.

The better way to think about it is:

Level 1: Reduce context

Only make Codex inspect what it actually needs.

Level 2: Reduce generation

Do not generate infrastructure that can be safely reused from a mature project.

Level 3: Reduce rework

Plan before large edits, validate between stages, and stop repeated failures.

The goal is not to make Codex do less useful work.

The goal is to stop spending your allowance on work that never needed to happen.

If you want the highest-impact starting point, I would make three changes first:

  1. Scope every task
  2. Plan before large edits
  3. Search GitHub before rebuilding common features

Those three changes did more for my Codex usage than obsessing over a fixed token number ever did.

Final Words + More Resources

My intention with this article was to help others share my knowledge and experience. If you want to contact me, you can contact by email: Email me

Here are also the most important links from this article along with some further resources that will help you in this scope:

Oh, and if you found these resources useful, don’t forget to support me by starring the repo on GitHub!

Comments