All posts
AI Engineering·8 min read·Updated 18 Sept 2026

How to Reduce Claude Code Token Usage: 9 Fixes That Actually Work (2026)

Quick answer
To reduce Claude Code token usage, run /context to see where tokens go, /clear between unrelated tasks, /rewind instead of /compact when a turn goes wrong, route routine work to a smaller model, lower the effort level, keep CLAUDE.md short, and disable MCP servers you are not using.

Nine practical ways to cut Claude Code token usage: /context, /clear, /rewind, /compact, model routing, effort levels, a lean CLAUDE.md, fewer MCP servers and a code index.

Piyush Jangir
Verified author

Founder of StackPicks. Self-taught builder shipping open-source dev tools, marketing, and curator content since 2019. Based in Mumbai, India. Available on GitHub and LinkedIn.

8 min read
How to Reduce Claude Code Token Usage: 9 Fixes That Actually Work (2026)

**Short version:** most wasted tokens in Claude Code come from context you are carrying but no longer need. Measure it with `/context`, then clear, rewind or compact deliberately, route easy work to a smaller model, and stop loading tools you do not use. Everything below is based on Anthropic's own cost management guide plus the tools people are actually installing.

The four Claude Code commands that control context: /context, /clear, /rewind and /compact

1. Measure before you cut: /context

Run /context inside a session. It breaks usage into the system prompt, tool definitions, memory files and conversation history. That one view tells you which of the fixes below matters for you. If tool definitions are large, fix your MCP list. If history dominates, fix how you clear and compact.

For tracking over days rather than one session, codeburn is a free local tracker that reads your usage logs across several coding tools.

2. /clear between unrelated tasks

Every turn resends the accumulated history. Finishing a bug fix and then asking about an unrelated feature in the same session drags the whole bug-fix conversation along. Run /clear when the task changes.

3. /rewind when a turn went wrong

If the last few turns went somewhere you do not want, /rewind to just before them. It cuts those turns off and costs nothing. Reaching for /compact instead rewrites the entire conversation and always costs a call.

4. /compact at phase breaks, not when things degrade

In a long task with natural phases, compact when a phase ends: design done, implementation starting. Compacting only once the session feels sluggish means you already paid for the bloated turns.

5. Route work to the right model

Model routing is usually the largest single saving. Keep a mid-size model as the default, send mechanical work to a small one, and reserve the largest model for hard debugging and design decisions.

6. Lower the effort level for routine work

Higher effort means more thinking tokens. Renames, test scaffolding and small edits rarely need it. Keep high effort for problems where you would want a senior engineer to slow down.

7. Keep CLAUDE.md short

CLAUDE.md loads into every session, so every line is a fixed cost on every request. Keep project rules, commands and pitfalls. Move long background material into files Claude reads only when relevant, or into a skill that loads on demand.

Mention files with `@path/to/file` instead of asking Claude to find them. Letting it explore a large repository means grep and read calls whose output lands in context. For big codebases, a code index such as CodeGraph lets the agent query structure over MCP instead of reading files wholesale.

9. Disconnect MCP servers you are not using

Tool definitions from every connected server load into every session. List what you have with claude mcp list and remove the rest with claude mcp remove <name>. Scope project-only servers with --scope project so they load only where needed.

Which token fix to use for which symptom in Claude Code

What about token-saving skills?

Skills such as caveman, built on the idea of saying less to spend less, have become some of the most-starred repositories on GitHub this year. They can help, but measure your own before-and-after with /context or a tracker rather than trusting a headline number. Your workload is not the author's.

The habit that saves the most

Interrupt early. When Claude heads down the wrong path, press Esc and redirect instead of letting it finish. A wrong approach that runs for twenty turns costs twenty turns of context, and then you pay again to undo it.

Frequently asked questions

What uses the most tokens in Claude Code?+

Usually the conversation history, not your prompt. Every turn resends the accumulated context, so a long session that has drifted across several tasks carries all of it forward. Tool definitions from connected MCP servers and a large CLAUDE.md are the next biggest fixed costs, because they load in every session. Run /context to see the breakdown for your own session before changing anything; the fix depends on which of those dominates.

Is /compact or /clear better for saving tokens?+

They solve different problems. /clear throws the history away, which is right when you are starting an unrelated task. /compact summarises the history so you keep the thread, which is right at a natural break inside one long task, but it costs a summarisation call each time. If the last few turns simply went wrong, /rewind is cheaper than either: it cuts those turns off and costs nothing.

Does switching models really reduce Claude Code costs?+

Model routing is typically the largest single saving. Using the biggest model for every edit spends frontier-model rates on work a smaller model handles fine. A sensible split is a mid-size model as the default, a small model for mechanical tasks such as renames and test scaffolding, and the largest model only for genuinely hard reasoning, debugging or architecture decisions.

Do MCP servers increase token usage?+

Yes. Each connected server adds its tool definitions to the context, and they load whether or not you use the tools in that session. Twelve servers connected "just in case" is a fixed tax on every request. Keep servers you use daily at user scope, move project-specific ones to project scope, and remove the rest with claude mcp remove.

Are there tools that track Claude Code token usage?+

Yes. Claude Code itself reports usage through /context and its cost commands. For tracking across sessions and across tools, open-source local trackers such as codeburn read your usage logs and show spend over time without sending data anywhere. Measure first, then change one thing at a time, so you know which fix actually moved the number for your workflow.

More in AI Engineering

How to Reduce Claude Code Token Usage: 9 Fixes That Actually Work (2026) — StackPicks — StackPicks