How to Stop Running Out of Codex Tokens in 2 Days
Problem
I have a Codex Pro subscription. For a while, I kept hitting my usage limit by Tuesday.
That left me with a frustrating choice for the rest of the week: slow down, switch tools, or wait for the allowance to reset.
My first assumption was simple: I needed more quota.
But after looking at how I actually used Codex, I realized that a large part of the problem was my workflow.
I was giving Codex too much context, asking vague questions, running work in parallel before validating the previous step, and sometimes asking it to generate code that already existed in mature open-source projects.
The biggest lesson was this:
The biggest Codex usage killer isn’t code generation itself. It’s unnecessary context, unnecessary exploration, and unnecessary rework.
Once I changed those habits, the same subscription started lasting much longer.
What Other Codex Users Were Seeing
A Reddit comment first made me look more closely at my workflow:
“It boggles my mind how people can use up the Codex Plus plan in 1 day… having 50 repos open, multiple worktrees, dozens of agents running continuously” — Reddit user (Score: 5)
That wasn’t exactly what I was doing, but the pattern was familiar: too much work happening at once, too much context, and too little control over what the agent actually needed to inspect.
Another user described something much closer to my experience:
“Token usage has been insanely bad for me the last two weeks. I’m hitting the weekly limit on pro in two days of fairly gentle usage” — Reddit user (Score: 3)
The important point is not whether those exact workloads match yours. The useful question is:
What work is Codex doing that never needed to happen?
That is where I found most of my savings.
How Codex Usage Works in 2026
One important correction to my original understanding: Codex usage is not simply a fixed bucket of “weekly tokens.”
Depending on your plan, Codex can be constrained by both a five-hour usage window and a weekly usage window. You need remaining allowance in the applicable windows to continue working.
The amount of useful work you get from that allowance also varies with:
- the model you choose
- reasoning settings
- task complexity
- how much repository context the agent needs
- how long the task runs
- how many tools and files it needs to inspect
OpenAI describes the published per-window figures as estimates rather than fixed message limits. In practice, a tiny mechanical edit and a cross-repository debugging session can consume very different amounts of allowance.
So I no longer try to model my plan as:
weekly_limit = 500000_tokensInstead, I use the product’s Settings → Usage view and Codex’s /status command when available to check the current session, context, and limits.
That also means the token numbers later in this article are my workload estimates, not OpenAI’s billing formula.
Where My Usage Was Going

Here is how I originally thought about a typical “simple” task:
Task: "Fix the bug in auth.py"├─ Read auth.py ~5,000 tokens├─ Read related imports ~9,000 tokens├─ Analyze codebase structure ~2,000 tokens├─ Generate fix ~1,000 tokens├─ Write edited file ~500 tokens└─ Context / reasoning overhead ~1,000 tokensEstimated total ~18,500 tokensThe exact number varies by model, tokenizer, tool behavior, caching, and how much context Codex actually reads. The useful part of this example is the shape of the cost.
The final code change might be tiny.
The expensive part can be everything Codex does before and around that change:
discover files→ read files→ understand dependencies→ form a hypothesis→ edit→ test→ reread→ debugIn agentic coding, input context and repeated reasoning can cost much more than the few lines of code you actually wanted changed.
That changed how I optimize Codex.
I now focus on three levels:
- Reduce context
- Reduce unnecessary generation
- Reduce rework
Here are the changes that mattered most.
1. Reduce Context Before Optimizing Anything Else
My first rule used to be:
One project at a time.
I still follow it, but I would phrase the reason more carefully now.
Simply having a repository visible in an editor does not necessarily mean every file is automatically sent to the model. The real problem is giving Codex a workspace or instruction that encourages it to inspect far more than the task requires.
The expensive version of a prompt looks like this:
Find why authentication sometimes fails.
Look through the project and fix it.That gives the agent permission to explore widely.
I now give it an explicit starting scope:
Investigate the login failure.
Start with:- src/auth/login.ts- src/auth/session.ts
Do not inspect unrelated directories unless these files point to them.
First explain the likely cause.Do not edit code yet.This does two useful things:
- limits unnecessary repository exploration
- separates diagnosis from implementation
For large monorepos, I also avoid casually asking Codex to “understand the whole project” unless that is actually required.
Generated files, build output, vendored dependencies, huge logs, and unrelated packages should not become part of the investigation by default.
My rule now is:
Scope first, code second.
2. Search GitHub Before Writing Thousands of Lines From Scratch

This became one of my biggest workflow changes.
When I started a new project, I used to ask something like:
Build me an admin dashboard.or:
Create an AI knowledge base with authentication, RAG, file upload,chat history, an admin panel, and Docker deployment.Codex could do a lot of that work, but it also meant generating a large amount of infrastructure that already existed elsewhere.
And the real cost was not just the first generation.
Every generated module could later need to be:
generated→ read again→ modified→ tested→ debugged→ rereadSo I now search for a mature open-source base before asking Codex to implement common infrastructure.
For example, an AI knowledge-base project may require:
- file upload and parsing
- embeddings
- vector search
- RAG
- chat UI
- authentication
- user management
- admin tools
- deployment configuration
If an existing project already covers 70–80% of that, my Codex workflow can become:
search→ evaluate→ fork→ remove unnecessary features→ adapt business logic→ connect my services→ deployinstead of:
design everything→ generate everything→ debug everything→ maintain everythingThe Prompt I Use Before Starting a New Project
I want to build XXX.
Do not write code yet.
First, search for existing open-source projects that could be useddirectly or adapted for this requirement.
Evaluate:
1. Is the project actively maintained?2. How recent are its commits and releases?3. Are Issues and PRs active and healthy?4. How difficult is deployment?5. Does the technology stack fit my project?6. Which features can be reused directly?7. Which modules would require modification?8. Is the license suitable for my intended use?
Then recommend one of:
- use directly- fork and customize- build from scratch
Finally, propose the smallest MVP.
Wait for confirmation before writing code.I do not choose a project purely because it has the most stars.
I check:
- recent commits
- releases
- open issues
- PR activity
- dependency freshness
- license
- deployment quality
- stack fit
- how much code I can realistically keep
The key idea is simple:
Less code generated today means less code Codex has to read, modify, and debug tomorrow.
3. Use the Cheapest Model That Can Reliably Finish the Task
Another mistake I made was treating “strongest model” as the default for every task.
That is convenient, but it is not always efficient.
Current Codex model availability and limits change over time, so I avoid hard-coding a permanent model hierarchy into my workflow. But the principle is stable:
Use the least expensive model that can reliably complete the task.
For example, lighter/faster models such as Luna-class options are often a better fit for:
- formatting
- extraction
- repetitive edits
- simple transformations
- documentation cleanup
- straightforward code changes
A mid-tier coding model such as Terra is a more natural default for:
- normal bug fixing
- tests
- common feature work
- day-to-day repository changes
I reserve stronger models such as Sol or Astra-class options for tasks such as:
- architecture decisions
- difficult cross-module debugging
- ambiguous failures
- large migrations
- race conditions
- complex refactoring
The same applies to reasoning effort.
This task:
Rename a field across 12 files and update the affected tests.usually does not need the highest reasoning setting.
This one might:
Find the root cause of an intermittent race conditionacross three services and propose the smallest safe fix.Model routing is one of the easiest ways to make a limited allowance go further.
4. Plan First, Then Let Codex Touch the Code
One of the most wasteful workflows is letting the agent start editing before you know whether its interpretation is correct.
I now use a plan-first prompt for anything non-trivial:
Do not modify files yet.
First:
1. Identify the likely root cause.2. List the files that need to change.3. Propose a solution in no more than 5 steps.4. Identify possible side effects.
Wait for approval before implementation.If the plan is wrong, I have spent a small amount of allowance discovering that.
That is much cheaper than:
scan repository→ change 10 files→ run tests→ discover bad assumption→ undo→ try againWhen available in my Codex host, /plan makes this workflow even easier.
My preferred sequence is:
Understand↓Plan↓Approve↓Implement↓Testnot:
Prompt↓Agent starts changing everything↓Discover wrong assumption↓Undo↓RetryPlanning is not overhead when it prevents an entire wrong implementation.
5. Keep AGENTS.md Small and Local
AGENTS.md is useful because Codex can automatically load project instructions from the repository hierarchy.
That also means it should contain high-value instructions—not a history of every decision ever made.
My root AGENTS.md now focuses on five things:
ProjectCommandsConstraintsDefinition of DonePermissionsFor example:
# ProjectNext.js + TypeScript + PostgreSQL
# Commandsnpm testnpm run lint
# Constraints- Do not modify migrations without approval- Do not add dependencies unless necessary- Preserve public APIs
# Definition of Done- tests pass- lint passes- no unrelated changes
# PermissionsAsk before deleting files or changing deployment config.I avoid putting temporary task details into the root instructions, such as:
- today’s bug description
- meeting notes
- long architecture history
- temporary migration plans
- thousands of lines of coding policy
For a monorepo, local instruction files are much cleaner:
/├── AGENTS.md├── frontend/│ └── AGENTS.md└── backend/ └── AGENTS.mdCodex reads instructions from the repository root toward the current working directory, so more local rules can specialize the broader rules.
That means a CSS task does not need to carry every backend database rule with it.
6. Split Reasoning-Heavy Work, Batch Repetitive Work
At first, these two recommendations sounded contradictory:
- break large tasks into smaller tasks
- batch similar tasks together
They are both correct.
The distinction is reasoning vs repetition.
Split work when each stage needs validation
For example:
architecture→ validate
implementation→ validate
tests→ validate
documentationI do not want Codex writing tests and documentation for an implementation that I have not reviewed yet.
Batch work when the operations are mechanically similar
For example:
- rename the same API field in 12 files- update 15 imports- add the same validation rule to several endpoints- convert a set of similar functionsPreviously I might have sent 10 separate requests:
Request 1 → load context → edit function ARequest 2 → load context → edit function BRequest 3 → load context → edit function C...Now I batch the related mechanical work when the same context applies.
In my own testing, a group of changes that might have cost roughly 50,000 tokens as separate tasks sometimes fell to around 30,000 when batched. That is an observation from my workflow, not a guaranteed Codex savings rate.
My rule is:
Batch repetitive work. Separate reasoning-heavy work.
7. Compact or Reset Stale Context
Long-running Codex conversations are useful because the agent remembers what has already happened.
But old context eventually becomes baggage.
I watch for a few signs:
- Codex starts referring to an old requirement that no longer applies
- the current task has drifted far from the original task
- a small change is carrying dozens of turns of history
/statusshows that context has grown significantly
When that happens, I decide whether I need to keep the history.
Useful options in supported Codex environments include:
/status— inspect session/context/limits/compact— compress a long conversation/fork— start a new main chat from the current history- a clean new thread — when the old history no longer matters
I do not reset the conversation after every small edit, because then Codex may have to rediscover important project state.
The goal is not “minimum context at all times.”
It is:
Keep useful state. Remove obsolete state.
8. Stop Repeating Failed Attempts
One of the easiest ways for an agent to burn allowance is to get stuck in a loop:
attempt→ fail→ small variation→ fail→ another small variation→ failI now use a one-retry rule.
The first failure is useful: it gives us new evidence.
A second attempt is reasonable if that evidence changes the hypothesis.
But if the same type of failure happens again, I stop execution and switch back to diagnosis.
I use a prompt like this:
Stop changing code.
Explain:
1. What we know2. Which assumptions may be wrong3. What evidence is missing4. What diagnostic command should be run next5. What alternative approach exists
Do not retry the same solution.My shorthand is:
No new evidence, no new retry.
When I Don’t Use Codex
Another optimization is simply recognizing tasks that do not need an agent.
I usually handle these with my IDE, CLI, or a lighter tool when possible:
- simple search-and-replace
- obvious renames
- JSON formatting
- deterministic formatting
- generated boilerplate
- trivial documentation cleanup
grep,sed, or IDE refactors
Codex CLI can also run with local open-source providers through --oss; OpenAI currently documents local providers such as Ollama and LM Studio.
That makes local models useful for some high-volume, low-risk work where I do not need the strongest reasoning model.
For example:
codex --oss --local-provider ollamaI would rather save my included Codex allowance for work where repository-aware reasoning is actually valuable.
My Low-Usage Codex Prompt Template
This is the structure I use for many debugging and maintenance tasks:
Task:Fix [specific problem].
Scope:Only inspect:- file A- file B
Do not inspect unrelated directories unless required.
Constraints:- preserve existing API- do not add dependencies- do not refactor unrelated code
Process:1. Diagnose the problem.2. Explain the root cause.3. Propose the smallest fix.4. Wait for confirmation before editing.
Validation:Run [specific test command].
Stop if:The same approach fails twice.The important fields are:
TaskScopeConstraintsProcessValidationStop conditionThe more clearly I define those, the less room Codex has to spend time exploring something I did not ask for.
My 30-Second Codex Checklist

Before I send a non-trivial task, I now ask:
- Can I do this faster manually?
- Does GitHub already have a mature implementation?
- Am I using an appropriate model for the task?
- Is the task scope specific?
- Does Codex really need to inspect the whole repository?
- Should I ask for a plan before editing?
- Can repetitive edits be batched?
- Is this conversation carrying obsolete context?
- Am I repeating a failed approach without new evidence?
It takes less than a minute and prevents a surprising amount of wasted work.
Results After One Month
After changing my workflow, these were my observed results:
| Metric | Before | After |
|---|---|---|
| Days until I typically hit the limit | 2 | 6-7 |
| Estimated tokens per typical task | ~20,000 | ~8,000 |
| Parallel agents | 3-4 | 1 |
| Weekly overrun | Frequent | Rare / none in the measured period |
These are not official Codex benchmarks, and they are not a promise that everyone will save the same amount.
They reflect my workload and my usage habits.
For my workflow, the estimated usage per typical task fell by roughly 60%.
The important part was not a single magic setting. It was eliminating unnecessary work before the model started doing it.
8 Ways I Accidentally Burned Codex Usage
1. Asking Codex to inspect the entire repository
Broad instructions encouraged broad exploration.
2. Using vague prompts
“Make authentication better” requires far more interpretation than a precise change request.
3. Using the strongest model for trivial work
A mechanical edit does not always need the most capable model or highest reasoning setting.
4. Running too many agents in parallel
Parallelism is useful when tasks are truly independent. I was using it before validating whether the first direction was correct.
5. Building common functionality from scratch
Authentication, dashboards, RAG shells, admin panels, and SaaS foundations often have mature open-source starting points.
6. Carrying stale context across unrelated tasks
Useful memory slowly turned into irrelevant baggage.
7. Letting an agent retry the same failure
Without new evidence, repeated retries were mostly repeated spending.
8. Using Codex for trivial work
Some changes are faster and cheaper with an IDE command or a few lines of shell.
Final Takeaway
I originally thought the answer was to make Codex “use fewer tokens.”
That framing was too narrow.
The better way to think about it is:
Level 1: Reduce context
Only make Codex inspect what it actually needs.
Level 2: Reduce generation
Do not generate infrastructure that can be safely reused from a mature project.
Level 3: Reduce rework
Plan before large edits, validate between stages, and stop repeated failures.
The goal is not to make Codex do less useful work.
The goal is to stop spending your allowance on work that never needed to happen.
If you want the highest-impact starting point, I would make three changes first:
- Scope every task
- Plan before large edits
- Search GitHub before rebuilding common features
Those three changes did more for my Codex usage than obsessing over a fixed token number ever did.
Final Words + More Resources
My intention with this article was to help others share my knowledge and experience. If you want to contact me, you can contact by email: Email me
Here are also the most important links from this article along with some further resources that will help you in this scope:
- 👨💻 Reddit Discussion: Codex Token Usage
- 👨💻 OpenAI: Managing Usage in Work and Codex
- 👨💻 OpenAI: Using Codex with Your ChatGPT Plan
- 👨💻 OpenAI: Model Guidance and AGENTS.md
- 👨💻 OpenAI: Codex Advanced Configuration
Oh, and if you found these resources useful, don’t forget to support me by starring the repo on GitHub!
Comments