Baseline · Monitor · Comply

Evidence that your AI service performs as you say it does.

Mavai® establishes a statistical baseline for each service you deploy — the model, its prompts, the success rate they achieve and the circumstances it was measured under — then monitors the service against that baseline for the life of the project. The result is evidence your own engineers act on, and that regulators and risk committees can read.

Built on open-source tools: punit for Java, feotest for Rust, baseltest for Python.

The Problem

If you deploy a service built on a language model under the EU AI Act, ISO/IEC 42001 or a sector supervisor’s guidance, you are required to show that it performs within acceptable bounds — not once, at go-live, but continuously, for as long as it runs. The EU AI Act, ISO/IEC 42001 and sector supervisors such as FINMA all ask the same question in different words: where is the evidence?

Traditional unit testing assumes deterministic outcomes. In reality, that assumption never withstood scrutiny. But AI means we have no choice but to manage uncertainty professionally, and that means statistically.

Mavai in a Nutshell

Baseline

Any one call can be judged right or wrong. What nobody knows in advance is how often the service gets it right. A baseline measures that rate over a stated number of runs and records it together with the model, the prompts, and the circumstances it was measured under.

Monitor

Hold the service to its baseline for as long as it runs. Every release, every model or prompt change, and on a schedule in between, the live service is tested against the baseline, and a fresh sample whose success rate falls below the bound the baseline implies is flagged as degradation, with a stated confidence.

Comply

The baseline is the record, monitoring is the evidence, and the method is documented in public: the Statistical Companion sets out the statistics and the open-source frameworks implement it line by line. Together they are the technical documentation the EU AI Act asks for, and what any supervisor, auditor or standard can read.

Who This Is For

Regulation is why more teams are asking the question now. The method is the same wherever a service’s behaviour is a rate rather than a value, supervised or not. We bring the implementation expertise; you bring the service.

Banking

Advisory, credit and client-facing assistants under FINMA’s expectations on AI governance and model risk, and the EU AI Act where it reaches.

Healthcare

Clinical-support and patient-facing services where the AI Act’s high-risk obligations meet medical-device post-market surveillance.

Insurance

Underwriting, claims and pricing models whose fairness and stability must be demonstrable to the supervisor and the policyholder alike.

Public Sector & Supervisory Bodies

Authorities that deploy AI themselves, or that must judge the evidence others submit, and need an independent, reproducible measure.

Current Engagements

We are baselining and monitoring live AI services in regulated sectors today. What each baseline tells its team about the service comes before any submission does. None of these has yet been part of a completed regulatory submission, so we describe what is being measured and where each engagement stands, not results. Interim figures will be published as they firm up.

Ask about an engagement in your sector

Insights

Clean Code Was Never More Important

AI amplifies the judgement already present in an organisation; it does not supply the judgement that is missing. Clean code is not prettiness - it is the structure that lets a competent human understand the problem. In the AI era, it is the steering wheel.

Shewhart, Toyota, and the Probabilistic Turn

LLMs break software’s binary view of correctness. The discipline that replaces it already exists - it was built at Bell Labs, refined in Nagoya, and has been waiting for software to need it.

All signals

Let’s talk

Compliance changes how a project is tested, not just what is documented at the end. A first conversation covers where your service lifecycle stands today and what a baseline-and-monitor regime would mean for it.

Let’s talk