External benchmark · 170/170 held-out scenarios · losses published

Measured against the EU AI Act Benchmark — including where we fail.

We ran our engines against the EU-funded AI Act Evaluation Benchmark — all 170 held-out scenarios — and we publish everything: the score, the interval, the per-scenario data, and the caveats. Measured, not claimed.

The number

Run 2026-08-02, $0. Two independent substrates — Apple M4 (local ollama) and NVIDIA T4 (Kaggle) — same frozen split, same harness, same weights. The intervals overlap almost exactly: this is a replicated measurement, not an anecdote.

Engine (variant)M4 set-F1 [95% CI]T4 set-F1 [95% CI]Exact matchScenarios
clan-refusal-gate0.153 [0.128, 0.182]0.151 [0.127, 0.180]0/170170/170 measured, both substrates
clan-law-refusing0.156 [0.131, 0.184]0.147 [0.123, 0.176]0/170170/170 measured, both substrates

artefacts: coai-dashboard/benchmark-results/aiact_benchmark/aiact_20260802_071146.json (M4) · coai-dashboard/benchmark-results/aiact_benchmark/aiact_t4_20260802_115217.json (Kaggle T4) — per-scenario scores, CIs, and run metadata included.

What this measures

Honest scope — what the stick is, and whose stick it is.

The task

Given a described AI system, list which EU AI Act articles apply — scored as set-F1 against the benchmark's gold article lists.

The benchmark

AI Act Evaluation Benchmark (davidath), arXiv 2603.09435, data CC-BY-4.0 — an EU-funded measuring instrument. We use it as a stick to measure ourselves; we did not build it and do not maintain it.

The split

Frozen split v1 — held-out selected by hash of the scenario text (sha256 % 2), fixed before any run, reproducible by anyone. 170 of 339 scenarios are held-out; we scored all 170, never a cherry-picked subset.

  • The benchmark's gold labels are LLM-generated (upstream disclosure) — we measure agreement with the benchmark, not legal truth.
  • Article retrieval is one task. It does not measure compliance judgement, obligation drafting, or deployment safety.
  • These are small council-tuned variants (≤4B class). The size ladder (0.5B → 8B, second substrate: Kaggle T4) publishes next; we expect the ordering to change and we will publish that too.

The statistics — why you can trust the interval

BCa bootstrap

2,000 resamples, fixed seed — near-nominal coverage for skewed score distributions at this sample size; plain percentile bootstrap undercovers.

Wilson score interval

For the exact-match proportion — correct behaviour near 0 and 1 where bootstrap fails. This run: 0/170 exact, 95% Wilson [0.000, 0.022], both engines.

Per-scenario scores published

Every per-scenario set-F1 is in the artefact — anyone can recompute our CIs or run paired tests against us.

Three-outcome honesty

A scenario is MEASURED, UNPARSEABLE (scored 0 and counted), or UNMEASURED (infrastructure failure — counted, never folded into the score). This run: 170 measured, 0 unparseable, 0 unmeasured, both engines.

The harness

Open, deterministic, stdlib-only: scripts/eat_aiact_benchmark.py in the coai-dashboard repo. Same prompts, same hashing, same statistics on every substrate — Apple M4 (this result) and Kaggle T4 (landing) — because two independent substrates measuring the same thing is the difference between a result and an anecdote.

No harmonised AI Act standards are published in the OJEU as of August 2026; we anchor on ISO/IEC 23894, TR 24027, EN ISO/IEC 42001 and CEN/CLC/TR 17894, and we say so.

What this page does not claim

  • Not a compliance certification. Article retrieval is one task, scored against one benchmark.
  • Not legal truth. The gold labels are LLM-generated; we measure agreement with the benchmark.
  • Not a leaderboard position. We publish our own engines only; no competitor was run on this page.
  • Not final. The second substrate (Kaggle T4) has now replicated these intervals (see table). The size ladder (0.5B → 8B) publishes next — ordering may change, and that will be published too.