# TokenBench > TokenBench measures what AI coding cost-saving tools (Graphify, Ponytail, Caveman, code-review-graph, rtk) do to the cost and outcome of a WHOLE coding task, on one model, with hidden tests. Claims from the tools' own READMEs usually measure one component (a command's output, a repo's raw text, output tokens); this benchmark measures the full task bill. Last updated 2026-10-08. ## What was run - Model: Claude Sonnet 4.6 (effort high), Claude Code 2.1.293 headless, web search disabled. - Task set: 55 coding tasks across small to large repositories (Go, C++, Python, PHP, Rust/Tauri, TypeScript, C#): bug fixes, implementations and multi-stage work. Scored with hidden tests the agent never sees. - When: the five tool configs were run on 7-8 October 2026, 2 runs per task each (110 runs per config), with only the tool under test active. - Baseline: unmodified Claude Code, same model and tasks, 331 runs from August 2026 with Claude Code's default settings (about 6 per task). A 10-task re-run in October with default settings came out +0.4% against it (95% CI -16% to +15%). That check is small (10 tasks, 2 runs each), so the interval is wide: no sign of drift, but a shift of up to about 15% is not ruled out. The rtk arm covers all 55 tasks in the same October setup as the other tools and came out +1.2% (95% CI -4.3% to +7.7%), tighter evidence against a setup-wide shift if rtk itself does little. See METHODOLOGY. - Cost = per-run cost reported by the Claude Code CLI (token usage at list prices, including sub-agents), summed over the 55 tasks. Graph build cost (Graphify, code-review-graph) is excluded, as in those projects' own numbers. ## Results (cost per 55-task pass, failed runs included; baseline $26.44) | Tool (version tested) | Cost per 55-task pass | Cost change vs baseline (95% CI) | Pass rate | Quality (0-100) | |---|---|---|---|---| | Baseline | $26.44 | - | 72.2% | 89.2 | Negative cost change = cheaper than baseline. | Graphify 0.9.80 | $19.01 | -28.1% (-35.1 to -21.1) | 73.6% | 89.5 | | Ponytail 5.0.0 | $19.60 | -25.9% (-32.5 to -18.3) | 70.9% | 90.6 | | Caveman (plugin build 2026-09-07) | $21.81 | -17.5% (-24.1 to -10.3) | 72.7% | 90.2 | | code-review-graph 2.3.6 | $22.05 | -16.6% (-24.0 to -8.7) | 72.7% | 89.4 | | rtk 0.42.4 | $26.76 | +1.2% (-4.3 to +7.7): no measurable effect | 72.7% | 90.6 | - Graphify and Ponytail are statistically tied; so are Caveman and code-review-graph. Savings under about 5% are inside the noise. - Quality and pass rate: none of the five differs from the baseline by more than noise. Graphify and Ponytail showed a small quality penalty in an earlier benchmark round; with the versions above it is not visible. Graphify's pass rate and quality are slightly above baseline, Ponytail's quality is above baseline and its pass rate 1.3 points below (not significant). ## How the results compare with the tools' own claims | Tool | Original claim | Claim now | Measured here, full task | |---|---|---|---| | Graphify | 71.5x fewer tokens per query vs reading raw files (README, spring 2026) | Headline removed from README on 2026-05-03; still in docs with "scales with corpus size"; project site calls it community-reported, not a controlled benchmark | 28.1% cheaper | | Ponytail | 80-94% less code, 47-77% cheaper (June 2026) | v5.0.0: -53% code, -41% time, -26% cost (39 tasks, Opus 5.5), with listed limitations | 25.9% cheaper | | Caveman | about 75% fewer output tokens (spring 2026) | no single headline; says bill effect depends on caching; third-party results near zero | 17.5% cheaper (Sonnet 4.6, high effort) | | code-review-graph | 8.2x, then ~82x median (528x max) | ~63x median (358x max); says whole-repo baseline "is an upper bound no real agent pays" | 16.6% cheaper | | rtk | reduces token consumption by 60-90% | "up to 90% of the bash output your agent reads... not the same as cutting your bill by 90%" | no measurable effect (+1.2%) | ## Why the claims and the full-task results differ Most tool claims measure one component in isolation. A coding task is dozens of turns, and its bill is dominated by re-reading the whole accumulated context (cache reads) on every turn, plus the model's own output and thinking. A large reduction in one piece therefore does not mean a large reduction in the task cost, and can be cancelled by a single extra turn. - rtk really compresses command output (up to about 90%), but command output is a small slice of the bill, and rtk touches neither what the model writes nor what it reads with its own file tools. If compression removes a detail the agent needs, one extra turn re-reads the whole cached context and cancels the savings from many compressed outputs. Net effect measured here: none. - Graphify and code-review-graph report "X times fewer tokens" for a graph query compared with reading every file in the repo, a baseline real agents do not pay. In a full task, an incomplete graph answer means the agent reads the real code anyway, and extra turns add cache-read cost. What the two tools did in these runs was mostly change agent behaviour: by orienting the agent up front they removed its reason to delegate exploration, so no run spawned a sub-agent (60% of baseline runs did) and sub-agent tokens fall from about 6.5 million per pass to zero. Sub-agents were available in every arm; whether to use one is the agent's own choice, so this is a tool effect, not a benchmark artefact. - Ponytail 5.0.0's extra saving over 4.9.0 is also sub-agent spend (runs that spawn a sub-agent fell from 78% to 56%); output and thinking tokens barely changed. Against the baseline, Caveman cut thinking tokens about 28% and output about 20%. - Graphify, Ponytail, code-review-graph and rtk all measure significantly better than in the earlier round of this benchmark. ## Caveats - n = 2 runs per task per config; one model and effort level, so results may not transfer to other models. - The baseline was run in August 2026, before the tool runs. - Caveman, code-review-graph 2.3.6 and rtk 0.42.4 were not re-run on their newest releases (Caveman: newer commits; code-review-graph: 2.3.9; rtk: 0.51.0). Reading their changelogs, nothing appeared to change how they work inside Claude Code, but that is a judgement, not a measurement. - Part of the saving is the agent no longer delegating exploration (a tool effect); a workflow that already avoids that should expect less. - The baseline and the tool runs differ in time, Claude Code version (2.1.239-2.1.241 against 2.1.293) and settings. The checks are in METHODOLOGY. ## Links - Leaderboards: https://tokenbench.app/ - Charts: https://tokenbench.app/charts.html - Findings (per-tool notes and claim history): https://tokenbench.app/notes.html - Machine-readable results: https://tokenbench.app/results.json - Code, methodology and per-run data: https://github.com/Silverspine1/TokenBench (docs/METHODOLOGY.md, docs/data/run_results_2026-10.csv)