Provider cost to clear the benchmark, per config — ranked cheapest first, with saving versus the baseline config.
| # | Config | Cost / task | Total | Saving vs baseline |
|---|---|---|---|---|
| 1 | crgraph | $0.261 | $19.04 | −23.1% |
| 2 | caveman | $0.275 | $20.05 | −19% |
| 3 | ponytail | $0.313 | $22.82 | −7.8% |
| 4 | ponytail+graphify | $0.316 | $22.76 | −6.8% |
| 5 | Baseline | $0.339 | $24.75 | 0% |
| 6 | graphify | $0.348 | $25.43 | +2.7% |
| 7 | rtk | $0.397 | $28.96 | +17% |
Quality earned per dollar spent — ranked best value first. Quality folds tests, correctness and cleanliness into one score; Value divides it by cost. Higher is better.
| # | Config | Valuequality per dollar | Qualitytests · correctness · cleanliness | Success |
|---|---|---|---|---|
| 1 | caveman | 312 | 85.7 | 67% |
| 2 | crgraph | 304 | 79.3 | 54% |
| 3 | ponytail+graphify | 263 | 83.1 | 60% |
| 4 | ponytail | 258 | 80.8 | 59% |
| 5 | Baseline | 249 | 84.3 | 63% |
| 6 | graphify | 234 | 81.6 | 58% |
| 7 | rtk | 209 | 82.8 | 65% |
The short version, if you'd rather not read the tables above.
Built and run by one person. Runs cost real money — donations pay for the greenfield benchmarks coming next.
Built & maintained by Silverspine1