What a $100 AI Coding Plan Is Actually Worth: A TokenOps Look at Claude Max vs Codex Pro
A TokenOps measurement: we measured how many API-priced dollars of work a $100/month Claude Max 5x and Codex Pro 5x plan actually delivered, week by week. Claude delivered about $940–1,170 a week, Codex about $270–340 — and Claude's 5-hour limit bought roughly 3.5× less on weekday mornings.
TL;DR
A $100/month plan is not $100 of AI. We measured one engineer's Claude Max 5x and Codex Pro 5x plans week by week, pricing every token at the vendor's own API list price:
- Claude Max 5x: 100% of the weekly limit was worth $940–1,170 of API usage, about $4,500 a month.
- Codex Pro 5x: 100% of the weekly limit was worth $340 on GPT-6 Astra and $270 after the switch to GPT-6.1 Sol, about $1,200–1,500 a month.
- Claude's 5-hour limit buys ~3.5× less on weekday mornings: a median $0.36 per 1% of the limit during Anthropic's 5–11am PT peak versus $1.30 off-peak.
Both plans cost the same $100. What they buy depends on the vendor, the model, and the hour. Measuring that continuously, in one unit, is what TokenOps is for.
Why measure a subscription in API dollars
Most teams pay for AI coding two ways at once: flat seats (Claude Max or Team, ChatGPT Pro or Business) for engineers, and metered API keys for CI jobs, agents, and internal tools. The seat invoice says $100. The API invoice says whatever was used. The two numbers can't be compared, so nobody can say whether a seat is a bargain or a waste, or whether moving an agent workflow from a key to a seat would save money.
The fix is the same move TokenOps makes for every AI line item: put everything on one unit. The unit here is API-equivalent dollars: price every token the plan consumed at the vendor's public list rate. A seat that burns $4,500 of list-price tokens in a month is a different asset from one that burns $90, even though both invoices read $100.
What we measured
Both vendors report how much of each limit window you have used, as a percentage. Our capture records that percentage alongside the token usage of every turn, so for each window we can divide API-priced spend by percentage points consumed.
Codex Pro 5x was worth $340 per week of limit, then $270. For three weeks on GPT-6 Astra, one percentage point bought $3.40–3.42 of API-priced usage, about as steady as a measurement gets. In late September this account's default moved to GPT-6.1 Sol, which OpenAI prices at a fifth of Astra's rate, and a point dropped to $2.69–2.70. Even at the edge of the rounding bound, Sol weeks stay below $2.95. Going back further, in late August on GPT-5.6 Sol, a point bought about $4.85 ($4.62–5.11 across three short windows that used 20 points in total), so the value of a point has fallen about 44% in five weeks without any change to the $100 price. Sol is so much cheaper per token that each point still buys far more tokens; it just buys about 21% fewer API dollars. In other words, the limit isn't metered strictly in proportion to list price, and the model you default to changes what your plan is worth.
Claude Max 5x was worth $940–1,170 per week of limit, about 3–4× Codex for the same $100. The latest week is still in progress and has the widest error bar (only 10 points used so far); the two complete weeks read $940 and $1,100.
Claude's 5-hour limit has a peak price
Since late March 2026, Anthropic drains the 5-hour session limit faster during weekday peak hours, 5 to 11am PT, while leaving weekly limits alone. We split 21 five-hour windows by when their usage happened:
At peak, 1% of the 5-hour limit bought a median $0.36 of API-priced work. Off-peak it bought $1.30. The same session limit runs out roughly three and a half times faster on a weekday morning. For a team whose engineers all start work in the same US time zone, that peak is precisely when most of the work happens.
The practical move: anything that can be scheduled (batch refactors, long agent runs, test-generation sweeps) is worth more off-peak. It doesn't change the weekly budget, but it decides whether an engineer hits the 5-hour wall at 10am.
How confident to be in these numbers
Treat these as one well-measured account, not a market survey.
- One account per vendor. These are our own plans, not customer data. Usage patterns differ, and a heavier cached-context workload will price differently from a lighter one.
- Whole-point rounding. Vendors report limits in whole percentage points, so a window that used 10 points carries up to ±10% error. Every chart shows the ±1-point bound. Weekly figures come from weeks that used 10–100 points; the 5-hour comparison only uses windows that consumed at least 20.
- Readings only go up. Idle sessions sometimes report a stale, lower percentage. We charge each window up to its highest reading, which 99% of snapshots sit within one point of.
- Claude reads slightly low. Claude's limits are shared across Claude Code, claude.ai, and the desktop app. Our dollars cover Claude Code on one machine; usage from other devices added about 6% in the week we cross-checked. Missing usage makes Claude's value per point look smaller, not bigger.
- List prices, current models. Spend is priced at each vendor's published API rate for the model actually used, including cache reads and writes. Promotional or negotiated rates would change the absolute dollars, not the comparison.
Why this is a TokenOps problem
TokenOps is the discipline of observing, attributing, and optimizing AI token spend across an engineering team, the way FinOps does for cloud spend. Subscription plans are where it is hardest, because the invoice is flat while the value underneath it moves. Every number in this post is a TokenOps measurement: put seats and API keys on one unit (API-equivalent dollars), attribute it per engineer and per window, and watch the exchange rate over time. Without that, a 44% drop in what a Codex point buys looks exactly like nothing happened.
What this means for an AI coding budget
- A seat's invoice is not its spend. A hard-used $100 seat can carry $1,200–4,500 a month of API-equivalent work. If you are deciding between seats, an Enterprise plan billed at API rates, or raw API keys, compare on API-equivalent dollars, not on the invoice. Our breakdown of what Claude Code costs a team walks through that comparison.
- The same plan is worth different amounts to different engineers. Value per seat depends on the model mix, the cache-hit rate, and when people work. That is an attribution question, and the invoice can't answer it.
- Vendors change the exchange rate. Peak-hour pricing, model launches, and limit resets all move what a percentage point buys, usually without an invoice change. You only see it if you measure it continuously.
How Promptster helps
Promptster Teams is a TokenOps platform for engineering teams using Claude Code, Codex, and Cursor. It does for your whole team what this post did for one account:
- Prices every seat in API dollars. Plan-limit readings are recorded next to the API-priced cost of the work, so you can see what your $100 and $200 seats actually deliver, and when that changes.
- Attributes spend across the team. Seat and API-key usage land on one ledger by repo and workflow. Each engineer sees their own numbers; managers see team and squad totals, never a per-person ranking.
- Surfaces recoverable waste. Re-read files, oversized context, and model choices that cost more than the task needed show up as recoverable token waste, separate from spend that produced work.
- Tracks the trend. When a vendor changes limits or a default model shifts, the value per seat moves on the chart, not silently in next quarter's budget.
Code and diffs are redacted on the engineer's machine before anything is uploaded; only prompt context and usage telemetry leave it. It is not developer monitoring.
See how Promptster Teams runs TokenOps for your team →
Related: What is TokenOps? · TokenOps metrics · How much does Claude Code cost a team? · AI spend attribution
Frequently asked questions
How much API usage does a $100/month AI coding plan include?
On one account measured over about three weeks (Sep 14 – Oct 6, 2026), 100% of the Codex Pro 5x weekly limit was worth about $340 at OpenAI API list prices while the account ran GPT-6 Astra, and about $270 after it switched to GPT-6.1 Sol — roughly $1,200–1,500 a month. 100% of the Claude Max 5x weekly limit was worth $940–1,170 at Anthropic API list prices, about $4,500 a month. Both plans cost $100 a month.Do Claude's peak hours change how much usage you get?
Yes, for the 5-hour limit. Anthropic drains the 5-hour session limit faster on weekdays from 5 to 11am PT. In our data, 1% of the 5-hour limit bought a median $0.36 of API-priced usage during peak versus $1.30 off-peak, about 3.5× less. Weekly limits are not affected.Did Codex's value change when the model changed?
Yes. On GPT-6 Astra, 1% of the Codex weekly limit bought a steady $3.40 of API-priced usage across three weeks. After the account's default moved to GPT-6.1 Sol, it bought about $2.70, roughly 21% less, outside the rounding error. Sol is far cheaper per token, so each point still covers more tokens and more work; it just covers fewer API dollars. No fast mode was used in either period.How was 'API-equivalent value' measured?
For every limit window, we took the token usage the coding agent recorded locally, priced it at the vendor's public API list price, and divided by how many percentage points of the limit that window consumed. Because the limit is reported in whole percentage points, each figure carries a ±1-point rounding bound, which the charts show.Should an engineering team buy subscription seats or pay API rates?
At these measured rates, a heavily used $100 seat delivers several times its price in API-equivalent work, so seats are cheaper for engineers who use them hard. The catch is visibility: a flat seat hides the API-equivalent spend underneath it, which is the number you need to compare seats, API keys, and enterprise plans on equal terms.