GovBench
An AI-governance benchmark that publishes its own resolution limit — and the experiments that refute its own architecture.
Ask SOV
Runs in your browser · nothing is sent anywhere
Deterministic reading of Articles 5, 6, 43 and 50 against Annex III — not legal advice, and it does not perform a conformity assessment. It answers what it can decide from the statute and says so when a question is outside that. See what we measure.
click to launch global
Measurement Lens
Frozen corpus · live axis · deterministic
0 of 15 scored dimensions have a resolved winner
Every dimension is statistically tied across all 10 models on Wilson intervals. At current item counts the minimum detectable effect is ≈63 points, and observed margins are 1–15. MMLU's own floor is 100 items per subject; Miller (arXiv:2411.00640) puts it at ~1,000 per comparison.
So: do not rank models on these numbers. Use them to find failure cases.
What is actually measured
Not a model — the composed pipeline (gate → retrieve → answer → verify → attest → mark) against the same items answered by a raw base model. n=193, paired, judged by an analysis written before the run existed.
Intervals are cluster-robust. Items inside a dimension share a rubric and a grader, so treating 193 items as 193 independent draws overstates precision — the measured design effect is 1.92, giving an honest effective n of ≈100. Every row below is computed from the same run and they partition it: 6 + 14 + 173 = 193.
| Layer | n | Δ vs base | 95% CI |
|---|---|---|---|
| Deterministic gate | 6 | −20.00 | [−65.26, +25.26] |
| Knowledge base | 14 | +19.64 | [+9.24, +30.04] |
| Tuned model | 173 | +6.50 | [+1.06, +11.95] |
| Whole system | 193 | +6.63 | [+1.05, +12.21] |
Wins 55 · losses 25 · ties 113 · sign test p=0.0011. Dropping the single largest item moves the headline to +6.15, so it does not rest on one case.
Experiments that refuted our own architecture
Published because a benchmark that only reports its wins is worth less than one that reports the controls killing its own thesis. Both of these were built, measured, and switched off.
Retracted 2026-07-29. Previously published as +34.84 — the largest number we had. Re-measured on a clean run it fires 6 times, not 31, and adds nothing: the base model already refuses all four plain-harm items, and its only measurable effects are two false blocks. The earlier figure was measured on a gate that had overfitted to its own battery; fixing the overfitting removed the benefit.
No effect. Routing ships OFF.
Significant harm.
Harm removed, benefit not shown. Retrieval ships OFF.
What this is not
This governs provenance, not correctness. The pipeline has shipped a wrong legal answer carrying a valid EU AI Act Article 50 marking and a clean signed receipt. An attested answer is attested, never verified — different word, different guarantee.
UNCERTIFIED is the default. No competent authority exists to confer EU AI Act conformity, so neither can we. Nothing here is a certification.
The models tested are system-prompt variants over one shared base, not separately trained weights — which is precisely why the routing experiment above came back null.
Run it yourself
pip install inspect-ai
inspect eval govbench_inspect.py --model ollama/qwen2.5:0.5bFailed runs are recorded as absent, never as zero — a model we could not reach is missing from the board, not scored badly on it.
The core instrument, measured
source: flywheel 6-model sweeps + gate1, both substrates, 2026-08-01
Frequently asked questions
GovBench — governance measurement with the method on the table.
What does GovBench measure?
Model behaviour on governance-shaped tasks: refusal correctness on prohibited-practice items, honest answers on benign-adjacent items, and the token cost per correct verdict — because a gate that refuses everything is not safe, it is useless.
How is overfitting prevented?
Items are split into practice and held-out sets by a fixed public salt (csoai-flywheel-v1). Training fuel is exported from practice items only, and the export path raises a hard error if a held-out item ever reaches it. The practice-vs-held-out gap is printed on every run.
Are the results cross-checked anywhere else?
Yes — the identical instrument runs on a second, independent substrate (a Kaggle T4 kernel) with the same salt and scoring. The flagship result replicated across both on 1 August 2026.
What happens when a model cannot be reached?
That lane is recorded as UNMEASURED and excluded from accuracy — never counted as a failure and never as a pass.