Free tool
What would your LLM workload cost — API or your own hardware?
Four deployment options, one workload, editable assumptions. The interesting number is the break-even: the volume at which a flat self-hosted setup becomes cheaper than paying per token. Everything runs in your browser.
Your workload
A request ≈ one question over one document chunk set, one extraction, one draft. 1,000 tokens ≈ 750 words.
Price assumptions (edit to match your quotes)
Flagship API (frontier model)
Efficient API (small/fast model)
EU-hosted open-weight API
Self-hosted (open-weight on your hardware)
Monthly cost at your volume
Defaults are indicative list-price ballparks (reviewed 28 August 2026) — edit them to match real quotes. Not in the math: engineering to build the pipeline (either way), model quality differences, and the reasons that often decide it — data residency, confidentiality, the AI Act audit trail. Cheapest and permitted are different questions; the picker next door handles the second one.
Default prices last reviewed: 28 August 2026. Reviewed quarterly.
The spreadsheet answer is not the architecture answer
We run open-weight models in production and sell nothing per token — so our read on API vs self-hosted has no margin in it. A 30-minute debrief gets you an honest one for your workload.

Mohamed Ben Haddou
Founder & CEO · Independent AI Expert for the European Commission
You talk to Mohamed, not a sales team — and the person who scopes your work is the one who delivers it.
