“Legal and the DPO vetoed the cloud AI tool”
And the business still wants the capability — so the answer cannot stay “no”.
Mentis · Services · Sovereign AI architecture & MLOps
Reference architectures for on-premise LLM inference and EU-hosted stacks, plus the pipelines, monitoring and retraining that keep models working — designed for organisations whose legal, compliance or clients rule out US clouds.
When to call us
And the business still wants the capability — so the answer cannot stay “no”.
It works, nobody can reproduce it, and it will break the day that person is on holiday.
Per-token costs climbing, data residency unclear, and no path back out.
What you get
Inference, vectors, logs and models on your premises or in EU-only hosting — data residency as a fact, not a clause.
GPU sizing and open-weight model choice against your real load — serious capability without an overbuilt data centre.
Pipelines, runbooks, monitoring and training so the stack is yours to run — or ours to run for you.
What we deliver
On-premise, EU-hosted or hybrid — components, data flows, security zones, the sovereignty rationale your DPO can sign.
GPU/CPU sizing, model selection (open-weight LLMs, embeddings), licensing, three-year cost compared with cloud equivalents.
LLM serving, vector database, API gateway, authentication — installed, secured and load-tested on your infrastructure.
Training and evaluation pipelines, model registry, CI/CD for models and prompts, reproducible environments.
Performance, drift, cost and usage dashboards; the record-keeping the AI Act expects for high-risk systems.
Operations documentation, incident playbooks and hands-on training for your IT team.
Fixed scope, fixed price per phase — you decide at each step.
Constraints, workloads, existing estate — and a reference architecture with a cost model.
Stack installed, secured, integrated with identity and your systems, load-tested.
Monitoring, runbooks, training — your team operates, we stay on call.
Regulation & sovereignty
GDPR transfer rules, DORA’s ICT third-party requirements in finance, professional secrecy in legal and healthcare, and the AI Act’s logging and record-keeping duties for high-risk systems all become simpler when inference, data and logs stay on infrastructure you control. We design the architecture so the compliance argument is in the diagram, not in a contract clause.
Yes, for most enterprise workloads. Open-weight models in the 7–70B range on one to a few GPUs serve assistants, extraction and summarisation for hundreds of users; we size against your real load and show the numbers.
It depends on volume and confidentiality value. We produce a three-year comparison — hardware, hosting, people — against cloud equivalents, so the decision is financial and regulatory, not ideological.
Often the best design: sensitive workloads on-premise or EU-hosted, non-sensitive ones on cloud APIs through EU regions, with routing rules that encode the policy.
No. EU regions, private endpoints and customer-managed keys cover many cases; for the rest, a sovereign enclave for the sensitive workloads sits next to your existing estate.
A 4-week, fixed-price diagnostic across six dimensions: maturity radar, scored opportunity portfolio, AI Act exposure register and a 90-day plan — the entry point for most clients.
Learn more →ServiceKnow exactly what the EU AI Act asks of you — and be able to prove it
Learn more →ServiceGenerative AI that answers from your documents — inside your walls
Learn more →ServicePredictions you can explain, defend and act on
Learn more →ServiceOptimization that moves the P&L — routes, schedules, resources
Learn more →ServiceFrom a model that works to a product your people actually use
Learn more →Tell us the problem, not a spec. You get an honest read on feasibility, data, compliance exposure and a first step — within one business day.