Skip to main content

What is Tokenmaxxing?

Finn
Finn Developer

Our development team has hired some new colleagues recently. They’ve helped us ship code faster than ever, though none of them appear directly on the payroll. Of course, I’m talking about AI agents. We use them to write code, scope new features, review our work and bounce ideas around. These days, I rarely go more than a few hours without typing a prompt.

I’d happily put them to work all day, freeing my own brain for the more important parts of the job (like responding to support tickets 😩). But agents are hungry, and their favourite food is tokens. This post is about tokenmaxxing: getting more useful work from agents without sending costs through the roof. I’ll share what I’ve tested, what has worked, and how to keep the resulting code clean and maintainable.

What even is “TokenMaxxing”?

TokenMaxxing means getting the most useful work from an AI agent for the tokens you spend. Tokens are essentially the chunks of text that are processed. An agent reads files, runs tests, revises code and tries again, with each step potentially triggering another model call. Context, reasoning, retries and subagents add to the cost; even cached input is billed. OpenAI's agent usage guide explains the breakdown more accurately.

We can take the September 2026 GPT-6 Astra API rates to simulate a more expensive example:

  • One coding task using one million input tokens and 100,000 output tokens would cost about $15.

  • Five developers each running five similarly-sized tasks a day, over 20 working days, would spend $7,500 a month, or $90,000 a year.

Caching, model choice, tools and subscription terms drastically change the cost of usage, so this is an illustration rather than a typical bill.

For a small or medium-sized business, unchecked usage could eat into the budget for another developer. Used well, an agent can help one developer investigate, implement and test more work while they remain responsible for quality. If that extra output holds up in review, the team may be able to delay its next hire and save some budget! That is the overarching goal here: more valuable work per token.

So, how do we do it?

Our developers have AI credits through our JetBrains organisation, giving us access to agents in our IDEs and a budget to work within. At first, usage was uneven: some of us burned through credits on broad requests, while others barely used them. A JetBrains article on spec-driven development gave us a more deliberate way to work. We adapted its example prompt to suit our team, and we now begin our coding session by writing a specification.

We give an agent a high-level feature request alongside our `spec-driven-development-prompt.md` file which is attached as context. Before proposing a solution, it inspects the existing code and conventions, then challenges the request. What happens when a user lacks permission? What if a step fails halfway through? Which edge cases have we missed? When an answer would materially change the feature, the agent asks us instead of quietly choosing for us. More often than not, this discussion reveals a use case we had not considered.

A feature request and our prompt in Junie. We start at Phase 0 and work through the specification one stage at a time.

Together, we turn that discussion into three linked documents:

  1. Requirements (`requirements.md`) describes the behaviour we want.

  2. Plan (`plan.md`) explains how it fits the existing system.

  3. Tasks (`tasks.md`) breaks the work into ordered, verifiable steps.


We review the documents before implementation, then work through the tasks in small stages, checking the code and tests as we go. That gives both the developer and the agent a shared record of what has been decided and what remains to be done.

We usually choose a more expensive frontier model for this planning stage. It is worth spending more tokens to uncover a missing requirement or a risky assumption before either becomes code. Once the work is clearly scoped, we can choose an appropriate model (at this stage, usually a far cheaper one) for each task and avoid paying for repeated attempts to solve the wrong problem.

The benefits

Clarity before coding. Requirements give us a shared definition of what the feature should do, while the plan lets both the developer and the agent question the scope and approach. It is much cheaper to remove an unnecessary step here than to spend tokens building it and then ask the agent to undo it.

Useful history. The documents outlive the conversation. A developer returning to the feature later can see what was intended, why a particular approach was chosen and how the work was divided, without having to reconstruct those decisions from the code alone.

The drawbacks

More to maintain. Three planning documents can feel like a lot of files for a small change. If we let the process grow beyond the problem, the agent can over-engineer a simple task and leave us with more code to maintain. The specification needs to stay proportional to the work.

Review takes time. We still need to check the documents, read the code and verify that the result meets the requirements. If we change direction halfway through, the plan and task list need updating too. That time is well spent on complex features, but it can outweigh the benefit for a quick, well-understood fix.

Conclusion

TokenMaxxing comes down to deciding where an agent's effort is worth paying for. We spend more on the thinking that prevents costly mistakes, then give the agent focused tasks and review what it produces. When the work is small, we keep the process light.

The developer still decides what matters and takes responsibility for the code that ships. Done well, this approach lets a small team make better use of its time and credits, deliver more without automatically adding headcount, and leave behind software that others can maintain. That's when an agent earns its tokens.