The Problem
If you deploy a service built on a language model under the EU AI Act, ISO/IEC 42001 or a sector supervisor’s guidance, you are required to show that it performs within acceptable bounds — not once, at go-live, but continuously, for as long as it runs. The EU AI Act, ISO/IEC 42001 and sector supervisors such as FINMA all ask the same question in different words: where is the evidence?
Traditional unit testing assumes deterministic outcomes. In reality, that assumption never withstood scrutiny. But AI means we have no choice but to manage uncertainty professionally, and that means statistically.