Our development team has hired some new colleagues recently. They’ve helped us ship code faster than ever, though none of them appear directly on the payroll. Of course, I’m talking about AI agents. We use them to write code, scope new features, review our work and bounce ideas around. These days, I rarely go more than a few hours without typing a prompt.
I’d happily put them to work all day, freeing my own brain for the more important parts of the job (like responding to support tickets 😩). But agents are hungry, and their favourite food is tokens. This post is about tokenmaxxing: getting more useful work from agents without sending costs through the roof. I’ll share what I’ve tested, what has worked, and how to keep the resulting code clean and maintainable.
What even is “TokenMaxxing”?
TokenMaxxing means getting the most useful work from an AI agent for the tokens you spend. Tokens are essentially the chunks of text that are processed. An agent reads files, runs tests, revises code and tries again, with each step potentially triggering another model call. Context, reasoning, retries and subagents add to the cost; even cached input is billed. OpenAI's agent usage guide explains the breakdown more accurately.
We can take the September 2026 GPT-6 Astra API rates to simulate a more expensive example:
One coding task using one million input tokens and 100,000 output tokens would cost about $15.
Five developers each running five similarly-sized tasks a day, over 20 working days, would spend $7,500 a month, or $90,000 a year.
Caching, model choice, tools and subscription terms drastically change the cost of usage, so this is an illustration rather than a typical bill.
For a small or medium-sized business, unchecked usage could eat into the budget for another developer. Used well, an agent can help one developer investigate, implement and test more work while they remain responsible for quality. If that extra output holds up in review, the team may be able to delay its next hire and save some budget! That is the overarching goal here: more valuable work per token.
So, how do we do it?
Our developers have AI credits through our JetBrains organisation, giving us access to agents in our IDEs and a budget to work within. At first, usage was uneven: some of us burned through credits on broad requests, while others barely used them. A JetBrains article on spec-driven development gave us a more deliberate way to work. We adapted its example prompt to suit our team, and we now begin our coding session by writing a specification.
We give an agent a high-level feature request alongside our `spec-driven-development-prompt.md` file which is attached as context. Before proposing a solution, it inspects the existing code and conventions, then challenges the request. What happens when a user lacks permission? What if a step fails halfway through? Which edge cases have we missed? When an answer would materially change the feature, the agent asks us instead of quietly choosing for us. More often than not, this discussion reveals a use case we had not considered.