Make your AI coding plan go further. From the first request to the last token.
TokenMaestro conducts the work between you, your AI agent and your computer: a local model structures your request, it asks what’s missing before the agent starts, heavy work goes to tools installed on your machine, and each task gets the right model and effort. Every saving is measured on your machine, never estimated.
First release: the free X-ray, for Claude Code. Next: the full orchestra, and Codex, Gemini and other agents.
Free X-ray · No account · Nothing leaves your computer
The orchestra: fewer tokens from the start
Most token tools compress what the agent has already spent. TokenMaestro works before the spend: on the request, the understanding, the tools and the model.
A well-built request
A small open model runs on your computer and turns what you typed into a clear, structured request before it reaches the agent. It costs no tokens.
Questions before the work
When a request is ambiguous, TokenMaestro asks you first, so the agent doesn’t spend tokens guessing and redoing.
Local tools for the heavy lifting
Reading text in images, converting documents, cutting video: tools installed on your machine do it, set up by TokenMaestro with your permission. The agent only gets the result.
The right model and effort
Simple tasks go to a lighter model at low effort; the heavy model and high effort only when the task needs them.
Lean sessions
Compact or start fresh before the context balloons, and before a break lets the cache expire.
Every saving measured
The X-ray shows where every token went. Each fix is measured on your own work, on and off; if it doesn’t pay off, it is switched off.
Your agent, your choice: Claude Code first; Codex, Gemini and others next. TokenMaestro conducts; you choose the instruments.
Start with the X-ray: what your plan was worth, and where it went
One command reads months of logs in seconds and writes a single dashboard file: value at API prices by day, model, effort level and project; days you hit the limit; and a cross-check against Claude Code’s own counters.
What it finds: the leaks, with a label on every number
“Measured” means the logs show it. “Ceiling” means the most a fix could recover — not a saving. In the founder’s logs:
Huge contexts
of the value went to turns with more than 400k tokens of context. Each step re-reads everything before it.
Cache rewritten after breaks
The cache expired 461 times during pauses, and the whole context was written again at full price.
Tokens before the first answer
Instructions, tools and memory every session starts with — re-read on every turn after that.
The weight map
What your agent read the most, by project and file. Area is the volume that entered the context; color shows how much of it was re-read with no change to the file.
Next: the same map drawn over your code’s structure, with the path your agent took in each session.

Why you can trust the numbers
Most token tools report savings they estimate themselves. Independent measurements of popular ones found no savings, and sometimes higher cost. We started from the other end.
The log traps, handled
Claude Code writes several records per message, the early ones with partial counts; resumed sessions replay messages; subagents write separate files. Each case has a test.
Checked against Claude Code
On 8 comparable sessions, our cache-read count came out 4.9% below Claude Code’s own counters, never above. The numbers are a floor.
No savings claims until measured
No percentage goes on this page until a paired measurement is published with the raw data and a way to reproduce it.
Download
The X-ray is a small command-line tool. Paste one line in your terminal to install it, or download the file.
The first public release is being prepared. The download links for Windows, macOS and Linux go here when it is out.
TokenMaestro runs on your computer: Windows, macOS or Linux. Open this page there to download it.
Pricing
X-ray
- Value and usage dashboard
- Leaks, with evidence labels
- Weight map
- Cross-check with Claude Code’s counters
- Missing-tools check
Pro Coming
Founding members: $49 for the first year, first 500 people.
- Structured requests from a local model
- Questions before the agent starts
- Local tools installed and tuned
- The right model and effort for each task
- Savings measured on your own work
- Limit meter; Codex, Gemini and other agents as they arrive
Purchase opens with the first public release.
Putting off one plan upgrade by a single month pays for a year of Pro.
FAQ
What does TokenMaestro read?
The session logs Claude Code already writes on your machine (~/.claude/projects): the token counts the API returned for each message, and which tools ran and how much they returned. Nothing is sent anywhere.
Is the dollar value what I paid?
No. It’s what your usage would have cost at API prices. On a flat-rate plan it shows how much your plan delivered, not a bill.
Claude Code already has /usage. Why this?
/usage tells you how much of your limit is gone. TokenMaestro tells you where it went, what it was worth, and what’s leaking.
How much will I save?
We don’t promise a number before measuring. The X-ray shows how much is leaking in your case. When fixes ship, each one is measured on your own work.
Do I need an account?
No. The X-ray needs no account and no email. Pro uses a license key, sent by email when you buy.
Does it work with Codex, Cursor or Gemini?
The X-ray reads Claude Code first. Codex and Gemini come next: you choose the agent, TokenMaestro conducts the rest.
Windows, macOS, Linux?
Yes. The X-ray runs on all three.
Do I need to install models or tools?
No. When a task calls for one, TokenMaestro shows what’s missing, explains why and installs it with your permission. The local model is small and open; if your computer can’t run it well, TokenMaestro works without it. (Pro, coming.)