We benchmark AI models on real factory tasks, so each role runs on the model that earns it.
| Role | Model | Cost / task | Quality |
|---|---|---|---|
| Analyst | Sonnet | 0,63 € | |
| UI Architect | Sonnet | 0,65 € | |
| Developer | Opus | 2,45 € | |
| Reviewer | Opus | 0,67 € | |
| Security | Opus | 0,37 € | |
| Performance | Sonnet | 0,19 € | |
| Tester | Sonnet | 1,55 € | |
| Deploy | Haiku | not run yet |
Cost per task is measured — total spend per role divided across 21 completed tasks. It updates as the factory works.
Claude
—
GPT
—
Gemini
—
We use the right model for each role, not one model for everything.
Every task so far has run on Claude. When the multi-model competition has scored enough issues, these cards fill in from real measurements.
Every task passes the same gates, so the scores compare like with like:
We count cost per successful task, not tokens consumed. Failed attempts count against the model that failed.