Skip to content

How an AI-Native Shop Found That 80% of Its Agent Cost Was Context Passed Between Agents

ops.ai ships nearly all of its work with autonomous Cursor and Codex agents. Session data showed that about 80% of the agent cost was context handed from one agent to its sub-agents. Here is what they changed, what happened to throughput, and what other teams running agents can take from it.

Paarth Jamdagneya
tokenopsai agent costcontext carrysub-agentsagent delegationspec-driven developmentengineering throughputcustomer story

TL;DR

ops.ai's biggest AI cost was context its agents handed to each other. ops.ai is an AI-native software consultancy whose work is shipped almost entirely by autonomous Cursor and Codex agents. Promptster's session data showed:

  • About 80% of the agent cost was context passed between agents. Agents opened large numbers of sub-agents and handed their full context to each one.
  • After ops.ai fixed how its agents split up work, it closed 611 Linear issues in September, up from 104 in July, and the median issue went from started to done in 4.8 hours, down from 23.5.
  • Sessions that started from a written spec shipped 75% of the time. Ad-hoc sessions shipped 34% of the time.

None of that shows up on an invoice. It shows up in the sessions. The full write-up is in the ops.ai customer story.

The setup: agents that run for 16 hours

ops.ai designs, builds and runs enterprise-scale products for its clients, from high-traffic e-commerce platforms to mobile apps, as well as its own feature-flag platform, Toggly. By the founder's count, the team completed about 1,100 pull requests in September. It runs 3–4 production deploys a day, with no production outage in two months.

Nearly all of that is agent work. ops.ai runs agents for up to 16 hours without a single prompt. They follow a "constitution" of delegation rules and about 200 pages of design specs.

The team could see the runs were long. It could not see what the agents spent their time and tokens on.

What the sessions showed

Promptster's delegation breakdown showed agents opening large numbers of sub-agents and handing their full context to each one. Contradictions in the constitution were driving it.

That handoff was about 80% of the cost, and it confused the agent reviewing the work.

"A lot of the process improvements I got by using Promptster. I wasn't aware of how much waste was going into context getting passed around, and the agents splitting at the wrong place."

Alex, Founder & CTO, ops.ai

The sessions showed it. The invoice never could: it has no line for context one agent copied into another.

What changed

ops.ai changed how its agents split up work, size it and iterate on it, and removed the contradictions from its rules. By the founder's account, spend dropped significantly.

Throughput is where the change is measurable:

  • 6× more work finished. ops.ai closed 611 Linear issues in September, up from 104 in July, before the change.
  • 5× faster, at six times the volume. The median issue went from started to done in 4.8 hours, down from 23.5.

The next workflow change: start from a spec

Promptster classifies every session by how the work was driven: from a written spec, from a plan, or by ad-hoc prompting. It then follows each one through to a merged PR.

At ops.ai, coding sessions that started from a written spec ended in a merged PR 75% of the time. Ad-hoc sessions shipped 34% of the time. The gap held when comparing sessions of the same size.

So ops.ai's next move is to start every coding task from a spec, and to watch the ship rate on the same dashboard that caught the delegation problem.

What other teams running agents can learn

  1. Look at what agents hand each other. When an agent delegates, the cost of the sub-agent includes everything it was given to read. If delegation rules are vague or contradict each other, agents split work at the wrong place and copy their whole context into every split. Contradictions in ops.ai's rules were what drove the handoff that made up about 80% of its agent cost.
  2. Grade a workflow by whether it ships. A session that never reaches a merged PR is spend with nothing to show for it. Comparing ship rates by how the work was driven (spec, plan or ad-hoc) gave ops.ai a 75% vs 34% gap to act on, and comparing sessions of the same size kept the comparison fair.
  3. Measure throughput in finished work. ops.ai's before-and-after is issues closed and the median time from started to done. Both are read from the work tracker, so they count work that finished, whatever the agents did along the way.
  4. Price every change in one unit. At ops.ai that came to about $10 of AI usage per merged PR and about $12 per Linear issue closed, with every token priced at the vendor's API list rate whatever the plan. One model was 59% of the month's spend, and a single 14-day agent session cost $2,506, which points at the next saving: a cheaper model for routine execution.
  5. Capture without changing how people work. ops.ai kept its own Cursor and Codex subscriptions. Capture ran passively, including on scheduled agents, with nothing routed through a proxy. A measurement that needs a new workflow changes the thing it is measuring.

How Promptster helps

Promptster Teams reads the sessions your engineers and agents already run in Claude Code, Codex and Cursor:

  • Where spend goes: token usage, sub-agent calls, skills and MCP servers, and top spenders by session, next to how the work was driven.
  • What ships: sessions followed through to merged PRs and closed issues, so cost per shipped change is a number you can track.
  • Where agents wait: hours agents sat waiting on a human, the next bottleneck after cost.

Code never leaves the machine. Before anything is sent, the open-source capture tool strips diffs, file contents, command output and the model's reply text, and redacts secrets.

Read the full ops.ai customer story →


Related: What is TokenOps? · TokenOps metrics · How much does Claude Code cost a team?

Frequently asked questions

  • How much of ops.ai's agent cost was context passed between agents?
    About 80%. Promptster's delegation breakdown showed ops.ai's agents opening large numbers of sub-agents and handing their full context to each one. That handoff was about 80% of the cost, and it confused the agent reviewing the work. Contradictions in the team's delegation rules were driving it.
  • What changed at ops.ai after the delegation fix?
    ops.ai closed 611 Linear issues in September, up from 104 in July, before the change: about 6× more work finished. The median issue went from started to done in 4.8 hours, down from 23.5, about 5× faster at six times the volume. By the founder's account, spend dropped significantly.
  • Do spec-first agent sessions ship more often than ad-hoc ones?
    At ops.ai, yes. Coding sessions that started from a written spec ended in a merged PR 75% of the time. Ad-hoc sessions shipped 34% of the time. The gap held when comparing sessions of the same size.
  • What does each shipped change cost ops.ai in AI usage?
    About $10 of AI usage per merged PR and about $12 per Linear issue closed, across 670+ PRs and 551 issues Promptster matched in the 30 days to Sep 28, 2026. Cost prices every token at the vendor's API list rate, whatever the plan.
  • Did ops.ai have to change its workflow to get this data?
    No. Promptster captures Cursor and Codex sessions passively, including scheduled agents, with no workflow change. ops.ai kept its own Cursor and Codex subscriptions, with nothing routed through a Promptster proxy.
Attribute · optimize · operate

See where your tokens go,
not just what they cost.

Your team's AI-coding spend went from zero to a real line item in eighteen months — unattributed, unbudgeted, invisible behind one vendor invoice. Promptster Teams is the TokenOps platform: it attributes spend per developer, separates recoverable waste from real leverage, and puts the whole loop on a budget.