Free tool

What would your LLM workload cost — API or your own hardware?

Four deployment options, one workload, editable assumptions. The interesting number is the break-even: the volume at which a flat self-hosted setup becomes cheaper than paying per token. Everything runs in your browser.

Your workload

A request ≈ one question over one document chunk set, one extraction, one draft. 1,000 tokens ≈ 750 words.

Price assumptions (edit to match your quotes)

Flagship API (frontier model)

Efficient API (small/fast model)

EU-hosted open-weight API

Self-hosted (open-weight on your hardware)

Monthly cost at your volume

    Defaults are indicative list-price ballparks (reviewed 28 August 2026) — edit them to match real quotes. Not in the math: engineering to build the pipeline (either way), model quality differences, and the reasons that often decide it — data residency, confidentiality, the AI Act audit trail. Cheapest and permitted are different questions; the picker next door handles the second one.

    Default prices last reviewed: 28 August 2026. Reviewed quarterly.

    The spreadsheet answer is not the architecture answer

    We run open-weight models in production and sell nothing per token — so our read on API vs self-hosted has no margin in it. A 30-minute debrief gets you an honest one for your workload.

    Mohamed Ben Haddou

    Mohamed Ben Haddou

    Founder & CEO · Independent AI Expert for the European Commission

    You talk to Mohamed, not a sales team — and the person who scopes your work is the one who delivers it.