I built TokenBench because the claims made by the tools' makers didn't match my experience using them.

TokenBench. Tasks span small to large repositories. Greenfield (build-from-scratch) benchmarks aren't included yet — they're next. Open source: github.com/Silverspine1/TokenBench My benchmark, my setup, my interpretation. Your situation is different. Do your own testing before relying on these numbers.

Cost saving

Provider cost to clear the benchmark, per config — ranked cheapest first, with saving versus the baseline config.

#ConfigCost / taskTotalSaving vs baseline
1crgraph$0.261$19.04−23.1%
2caveman$0.275$20.05−19%
3ponytail$0.313$22.82−7.8%
4ponytail+graphify$0.316$22.76−6.8%
5Baseline$0.339$24.750%
6graphify$0.348$25.43+2.7%
7rtk$0.397$28.96+17%

Value

Quality earned per dollar spent — ranked best value first. Quality folds tests, correctness and cleanliness into one score; Value divides it by cost. Higher is better.

#ConfigValuequality per dollarQualitytests · correctness · cleanlinessSuccess
1caveman31285.767%
2crgraph30479.354%
3ponytail+graphify26383.160%
4ponytail25880.859%
5Baseline24984.363%
6graphify23481.658%
7rtk20982.865%

Suggested configs

The short version, if you'd rather not read the tables above.

Best value
caveman
312 quality per $
Most quality per dollar — the all-round pick.
Cheapest
crgraph
$0.261 per task
Lowest spend per task, if budget is the constraint.
Highest quality
caveman
quality 85.7
Best results regardless of cost.