How it works
Long Claude Code sessions get expensive and unfocused for a simple reason: every turn re-reads everything that came before. rethinker watches for the moments when carrying that history stops paying off, and makes starting fresh cheap.
The problem
A coding agent sends the whole conversation back to the model on every turn. Prompt caching makes those re-reads cheaper, but not free, and a cache entry only lives for a few minutes without use. After a coffee break, the next turn writes the entire context to the cache again.
As a session grows, each turn costs more and the model has more noise to sort through. The obvious fix, making prompts shorter, does not reliably help: in one recent study, a method that cut tool-output tokens by 38.4% raised the billed cost by 6.8% (arXiv:2607.12161). What matters is when you carry history forward and when you start clean.
The flow
rethinker runs alongside your Claude Code session. It also works with Codex CLI and GitHub Copilot.
The signals
Each signal maps to one concrete suggestion.
| Signal | What rethinker watches | What it suggests |
|---|---|---|
| Large context | The session's context passing a size where every turn gets expensive | Hand off to a fresh session |
| Cold cache | Time since the last turn compared with the prompt-cache lifetime | Hand off before the next turn re-writes the whole context |
| Runaway loop | The same command run again and again with no edit in between | Pause and rethink the approach |
| Unclear prompt | Requests likely to cost extra turns of back-and-forth | A clearer version of the ask |
The handoff note
Short enough to start a session cheaply, complete enough that nothing important is lost.
Illustration of the note format on an example task, not captured output.
How we measure cost
We measure in money, per task. For each task we take Claude's own usage data, split into input, cache writes, cache reads and output, and price each part at the published API rates. rethinker's own Claude calls count against it too. Savings only count if the same kind of task gets cheaper, not if a prompt gets shorter.
We have not published results yet. When we do, the method and the raw numbers will be published together.
What's next: legacy modernization
The same discipline (small steps, clear handoffs, honest measurement) is what large rewrites need. We are designing a modernization workflow in which Claude agents map a legacy application into an inventory graph of screens, endpoints, data and dependencies, split the work into small slices with acceptance criteria and a rollback plan, and draft each slice for engineers to review and own. It is in development and not yet available.
Try it
rethinker is in private beta with a small number of external teams. To request access or a live demo, write to sales@rethinker.dev.