
Board source: GET /api/gspc — recompute published results, free
The GSPC board
23 axes measured · 14 model fleets · 0 separated leaders · 3 public leader scores · 9 fact runs · TIE is TIE · not a certificate. (14 model-comparison + 9 fact runs — deterministic fact checks have no model fleet, leader or accuracy) · deterministic grading on frozen, published splits · a TIE means the leader's edge is statistically indistinguishable (McNemar p≥0.05) — ties are never counted as wins. Empty cells stay empty. Measured axes are not the same as public leader scores.
23 axes are on the board and every one carries a measurement — no declared slot is empty. Both counts are DERIVED from the axis array, never typed; if a future slot is added with no run behind it, this line separates the two again on its own. A measurement is not a separated leader: 0 of 14 model-comparison axes separated a leader · 2 TIE · 12 UNTESTED. A point-estimate lead is not a measured advantage, and UNTESTED is not a tie.
GSPC terminal · one board
23 axis · 23 measured
npx -y csoai-gspc-mcp| Axis | Bench | Figure | n | Status | |
|---|---|---|---|---|---|
| governance | GovBench | UNMEASURED | 237 | MEASURED | ▸ |
| safety | DefBench | 94.4% | 36 | TIE | ▸ |
| provenance | ProvBench | UNMEASURED | 32 | MEASURED | ▸ |
| continuity | PQCBench | UNMEASURED | 33 | MEASURED | ▸ |
| conformance | MCPBench | UNMEASURED | 35 | MEASURED | ▸ |
| openness | OSSBench | UNMEASURED | 32 | MEASURED | ▸ |
| machinery-conformity | MachBench | UNMEASURED | 33 | MEASURED | ▸ |
| care | CareBench | UNMEASURED | 199 | MEASURED | ▸ |
| cross-reality | XRAIV | UNMEASURED | 32 | MEASURED | ▸ |
| detector-interop | DetBench | UNMEASURED | 33 | MEASURED | ▸ |
| art5-safeguard | Art5Bench | UNMEASURED | 36 | MEASURED | ▸ |
| swarm | SwarmBench v2b | 44.4% | 37 | UNTESTED | ▸ |
| affect | AffectBench | UNMEASURED | 41 | MEASURED | ▸ |
| jail | GoldBank-Detector | 59.2% | 71 | TIE | ▸ |
| effect-binding | EffectBench v0.1 (server probe) | facts | 261 | FACTS | ▸ |
| provenance-controls | ChainFacts | facts | 6 | FACTS | ▸ |
| reserve-attestation | ReserveFacts | facts | 16 | FACTS | ▸ |
| regulatory-framework | RegimeFacts | facts | 16 | FACTS | ▸ |
| distribution-integrity | DistributionFacts | facts | 16 | FACTS | ▸ |
| custody-disclosure | CustodyFacts | facts | 16 | FACTS | ▸ |
| ai-adoption-components | Eurostat | facts | 2 | FACTS | ▸ |
| labour-components | Eurostat | facts | 2 | FACTS | ▸ |
| humanoid-labour-index | Disclosure | facts | 8 | FACTS | ▸ |
The full board — with intervals and harm tails
Axis deep-dive
art5-safeguard
Bench Art5Bench · n=36 · leader accuracy no public leader score — axis measured · separation UNTESTED
Frozen gold bank (Hugging Face)Verify the signed chainRaw JSON (GET /api/gspc)Full board
| Axis | Bench | n | Leader accuracy | 95% CI | Separation | Open traces / graphs |
|---|---|---|---|---|---|---|
| governance | GovBench | 237 | no public leader score | not published — no public leader score | UNTESTED | |
| safety | DefBench | 36 | 94.4% | 81.9–98.5% | TIE — indistinguishablep=0.6875 | |
| provenance | ProvBench | 32 | no public leader score | not published — no public leader score | UNTESTED | |
| continuity | PQCBench | 33 | no public leader score | not published — no public leader score | UNTESTED | |
| conformance | MCPBench | 35 | no public leader score | not published — no public leader score | UNTESTED | |
| openness | OSSBench | 32 | no public leader score | not published — no public leader score | UNTESTED | |
| machinery-conformity | MachBench | 33 | no public leader score | not published — no public leader score | UNTESTED | |
| care | CareBench | 199 | no public leader score | not published — no public leader score | UNTESTED | |
| cross-reality | XRAIV | 32 | no public leader score | not published — no public leader score | UNTESTED | |
| detector-interop | DetBench | 33 | no public leader score | not published — no public leader score | UNTESTED | |
| art5-safeguard | Art5Bench | 36 | no public leader score | not published — no public leader score | UNTESTED | |
| swarm | SwarmBench v2b | 37 | ≥44.4%lower bound | withheld (n not independent) | UNTESTED | |
| affect | AffectBench | 41 | no public leader score | not published — no public leader score | UNTESTED | |
| jail | GoldBank-Detector | 71 | 59.2% | 47.5–69.8% | TIE — indistinguishable | |
| effect-binding | EffectBench v0.1 (server probe) | 261tool-call servers probed | no leader accuracy261 of 600 servers tried, from a population of 20,992 third-party servers | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| provenance-controls | ChainFacts | 6issuer accounts (not bank items) | no leader accuracy6 of the 16 instruments named in the registry | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| reserve-attestation | ReserveFacts | 16issuer accounts (not bank items) | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| regulatory-framework | RegimeFacts | 16issuer accounts (not bank items) | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| distribution-integrity | DistributionFacts | 16issuer accounts (not bank items) | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| custody-disclosure | CustodyFacts | 16issuer accounts (not bank items) | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| ai-adoption-components | Eurostat | 2public series | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| labour-components | Eurostat | 2public series | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| humanoid-labour-index | Disclosure | 8frozen vendor URLs | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader |
axis,bench,status,n,accuracy,interval_lo,interval_hi,separation,fleet_mean,dataset,as_of · empty cells stay empty — never zeroed, never interpolated.
Attestation · live from GET /api/gspc · click any row for traces
Measurement freshness · derived from GET /api/gspc · measured_on
behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25
living_stamp.gold_run 18 Aug 2026 — payload stamp, not a live re-measure. Board counts stay derived from GET /api/gspc; no new MEASURED invented here.
These run dates are weeks old. Freshness is labelled; the board is not re-stamped from this UI.
Living Stamp — SIGNED
Do not treat this as a valid attestation. Check site_attestation instead.
Progress · 23 axis · 23 measured
N→N+1 drift · UNCHECKABLE
No published board time series for N→N+1 drift. Empty stays empty — do not invent drift numbers or a Merkle seal. Cite GET /root.json and GET /api/gspc for the living snapshot only. Living snapshot only — cite GET /root.json.
23 axes are on the board and every one carries a measurement — no declared slot is empty. Both counts are DERIVED from the axis array, never typed; if a future slot is added with no run behind it, this line separates the two again on its own. A measurement is not a separated leader: 0 of 14 model-comparison axes separated a leader · 2 TIE · 12 UNTESTED. A point-estimate lead is not a measured advantage, and UNTESTED is not a tie.
Arena Elo — signed
Per-axis Elo from the published arena snapshot (7 models, 28 arena axis — the arena's own set, not the board's count above). Snapshot: . New arena rounds do not change GSPC board axes without admission. The board DID signature is present; use the check below to verify these bytes.
| Model | Elo | Games | Win-rate | 95% CI |
|---|---|---|---|---|
| qwen3:8b | 1660.6 | 210 | 0.695 | 0.63–0.754 |
| llama3.1:8b | 1589.3 | 37 | 0.514 | 0.359–0.666 |
| phi3.5:3.8b | 1549.3 | 123 | 0.553 | 0.465–0.638 |
| phi4:14b | 1527.3 | 309 | 0.563 | 0.507–0.617 |
| nemotron-3-nano:30b | 1389.6 | 305 | 0.469 | 0.414–0.525 |
| gemma3:12b | 1370.1 | 477 | 0.258 | 0.221–0.299 |
| mistral:7b | 1365.1 | 494 | 0.613 | 0.57–0.655 |
In-lane measurements — not board rows
2 slots measured in-lane. Published as measured_in_lane on GET /api/gspc. NOT stamped onto the board count. public_count stays “23 axis · 23 measured” — quoted live, never typed. Click any card for traces + graphs.
Measurement, not certification. Leaders shown are point estimates (swarm quotes its 95% lower bound); only SEPARATED leads are statistically real — the live count is totals.separated_leads on GET /api/gspc. Jail is a measured floor when the stamp publishes one, never a hidden score. Full per-axis notes, fleet means and harm tails: GET /api/gspc. The living stamp carried there is marked UNVERIFIABLE — it does not reproduce under any published rule and is not a checkable attestation; the attestation over that payload that does verify is site_attestation.