tokenwise.sk
← Back

The Arena

We benchmark AI models on real factory tasks, so each role runs on the model that earns it.

Model leaderboard

  • AnalystSonnet
    0,63 €
  • UI ArchitectSonnet
    0,65 €
  • DeveloperOpus
    2,45 €
  • ReviewerOpus
    0,67 €
  • SecurityOpus
    0,37 €
  • PerformanceSonnet
    0,19 €
  • TesterSonnet
    1,55 €
  • DeployHaiku
    not run yet

Cost per task is measured — total spend per role divided across 21 completed tasks. It updates as the factory works.

Provider comparison

Not yet measured
  • Claude

  • GPT

  • Gemini

We use the right model for each role, not one model for everything.

Every task so far has run on Claude. When the multi-model competition has scored enough issues, these cards fill in from real measurements.

Methodology

Every task passes the same gates, so the scores compare like with like:

  • Code review pass rate on first attempt
  • Security audit findings per change
  • End-to-end tests green before merge
  • Cost per merged pull request

We count cost per successful task, not tokens consumed. Failed attempts count against the model that failed.

Start a project →Get a consultation (coming soon)