The honesty gate
What we publish against ourselves.
An instrument that will say anything measures nothing. So this page collects the results that embarrass us, the gaps we have not closed, and the exact points where our own cryptography stops proving things. Every number on it is read from a published artifact at load time; none is typed.
1. Our own fine-tunes lose our own arena
We built council fine-tunes on small base models. They are beaten by the bases we started from. This is the most credible thing we can publish, because it contradicts our own product narrative — and because you can fetch the artifact and check it without us.
| Model | Elo | Games | Win rate | Wilson 95% |
|---|---|---|---|---|
| qwen3:8bnot ours | 1660.6 | 210 | 69.5% | 63.0–75.4% |
| llama3.1:8bnot ours | 1589.3 | 37 | 51.4% | 35.9–66.6% |
| phi3.5:3.8bnot ours | 1549.3 | 123 | 55.3% | 46.5–63.8% |
| phi4:14bnot ours | 1527.3 | 309 | 56.3% | 50.7–61.7% |
| nemotron-3-nano:30bnot ours | 1389.6 | 305 | 46.9% | 41.4–52.5% |
| gemma3:12bnot ours | 1370.1 | 477 | 25.8% | 22.1–29.9% |
| mistral:7bnot ours | 1365.1 | 494 | 61.3% | 57.0–65.5% |
Read from /arena/elo_reference.json · generated 2026-09-24T03:36:56Z · Bradley-Terry Elo, K=32, base 1500, per-axis and overall; Wilson 95% CI on win-rate; ranked only with n>=5 decided games; ties counted, not rated; style_controlled_weight honoured when a round carries it.. Nothing in this table is typed into the page. Rows with a small games count carry a wide interval and should be read as such — the interval is printed rather than hidden. Arena Elo is an internal-doctrine figure we publish here, against ourselves; it is never published as a verdict on anyone else's model, and the board at /gspc-arena remains the only measurement surface we stand behind.
What it does mean: adapter-souping weak bases does not beat the base, and the measurement rail works — it caught us. What it does not mean: that the instruments are broken. The instrument that shows us losing is the same one we publish. The honest next step is base model plus statute retrieval, not weight-merging weak specialists.
One ceiling, stated before anyone else states it for us: this instrument governs provenance, not correctness. An attested answer is attested, never verified. Our own fine-tunes are the proof — they are signed, and they still lose.
2. Where our cryptography stops
335 signed measurement cards are catalogued in the index, and all 335 verify against did:web:csoai.org#card-attestation-1. A stranger can run that check offline with public/signed/verify-card.mjs, no account and no permission.
The limit, stated precisely.
0 card bodies are withheld because their signed bodies carry an internal identifier — disclosed, not deleted, and each keeps its position in the chain. Of those, 0 are independently attested by a published card's signed parent link; the rest appear inside the signed chain manifest — non-repudiable as a list, but still our own attestation of our own list rather than an independent signature.
We publish that distinction because the alternative is letting a complete-looking manifest do work a signature has not done — which is the exact class of defect this estate keeps finding in itself. The full position manifest is published at /signed/chain.json as a card-shaped, Ed25519-signed envelope — verify it yourself with /signed/verify-card.mjs, unchanged, under the same pinned key as every card. The derived counts are at /signed/chain-facts.json, each one recomputed from the bytes, and the derivation behind these two numbers is at GET /api/state → card_chain.
Two more limits in the same family. There is no RFC-3161 timestamp authority and no blockchain anchoring behind any card — our records say timestamp_authority: none. A card's trust path is an Ed25519 signature over a SHA-256 hash chain, verifiable offline against did:web:csoai.org — no blockchain and no timestamp authority sits in that path. The /xrpl-attest page is a reader of GET /root.json (signed root envelope; inclusion does not individually sign a leaf). GET /api/xrpl is a reader of that root (writes_board false, live locked 16, same merkle). Historical DEVNET Payment-memo / CredentialCreate hashes are not this feed. XLS-70 Credentials are live on XRPL mainnet as an allowlist primitive; we are not issuing GSPC grades on-ledger. Separately from the card trust path, The current canonical public root has a proof-derived CONFIRMED_BITCOIN OpenTimestamps witness at block 968130. That witness covers the exact public root.json bytes only, not the separate signed-card index. Queued and candidate atoms are not automatically admitted, published, or anchored; a pending calendar stamp, where one exists, does not by itself prove inclusion in a Bitcoin block. There is no Ethereum-chain attestation backend.
3. What we have not measured
23 axes are on the board and every one carries a measurement — no declared slot is empty. Both counts are DERIVED from the axis array, never typed; if a future slot is added with no run behind it, this line separates the two again on its own. A measurement is not a separated leader: 0 of 14 model-comparison axes separated a leader · 2 TIE · 12 UNTESTED. A point-estimate lead is not a measured advantage, and UNTESTED is not a tie. Those slots are published precisely so the gap is visible. A slot is not a measurement and we will not let the larger number stand alone.
A slot stays UNMEASURED when the sample is too small to quote — nothing goes on the board below thirty usable graded items, and a wave queued at twenty-four returned UNMEASURED across every job rather than being quoted — or when the instrument is not frozen and published, or when the legal gold labels are still with counsel. UNMEASURED is not a failing grade for anyone's AI system; it is a disclosure about us.
The largest single one: comparative coverage of the evaluation landscape. We have measured one rating organisation, on one criterion, on one benchmark. No survey of raters exists and no cross-rater comparison is published, and our claims register records that as UNMEASURED at CR-020 rather than letting the one result imply a landscape verdict. Every material claim, with its status.
4. Where we were wrong, appended and never edited
The corrections ledger at GET councilof.ai/api/corrections records what was wrong, how it was caught, and what changed. Entries are appended; none is edited or deleted. Most were caught by our own instrument turned on its owner.
- A verification of ours that could not observe failure. Our prerender check recorded a failed route in a field named
err, while every checker in the repository read a field namederrored— which never existed. The check could not have reported a failure if one had occurred. A verifier that is structurally unable to fail is worse than no verifier, because it also produces a green light. - A verification of our own card store that returned zero. On 26 August we ran the ruling's own test — every card hash must resolve to signed bytes that recompute — across eight candidate stores on pods and Hugging Face. It resolved none of them, and we published that result at
/interop/card-store-verification.jsonwith the honest note that zero verified is a fact about the reachable record, not a claim that the cards do not exist. Today the bodies published under/signed/cards/do verify, against the pinned key, with the verifier we ship — the count is at the top of this section. Both records stand: the dated failure is not deleted because a later run succeeded. - We retracted a guarantee rather than rewording it. We had published a consensus guarantee for our council architecture. The historical DR-0007 narrative named a numeric result, but its cited result artifact is absent from this repository, so that number is unbound and not independently reproducible. The guarantee was withdrawn. The 33-seat structure with its 23-of-33 threshold remains a design figure and is labelled as one everywhere. The latest published point test measured rho=1 and n_eff=1.
- Our own board contradicted our own ruling for two days. An owner ruling set the canonical axis count; the endpoint kept reporting the pre-sweep number because the new axis existed in the ruling and not in the payload the count is derived from. No axis was marked measured to close that gap.
- We repeated a human-versus-machine comparison without checking rule-match. The attribution was careful and the number was correctly labelled as reported, not measured. The defect was publishing the contrast at all without asking whether both sides were scored under the same rule — the very question our first rating-the-raters result exists to ask.
The refutation ledger — claims we published, tested, and killed
5. REPORTED — figures by others, never mixed with ours
Three data states run this estate. MEASURED — a graded run on our frozen instruments, signed. UNMEASURED — honestly empty, published so the gap is visible. REPORTED — a figure published by someone else, cited and dated, carried for context and left unsigned. A REPORTED number never enters our board and is never averaged with a MEASURED one; the human-performance baselines beside our AI figures are REPORTED aggregates from other people's studies. The machine-readable set, each entry with its source URL, capture date and attribution basis, is at GET councilof.ai/api/reported. Scores move: read every figure as of its capture date and follow the source for the live number.
Honesty, stated as questions
The same five disclosures as the sections above — structured so answer engines can cite them without inventing a sixth.
Do Council of AI fine-tunes beat the bases they started from?
No. On the signed arena reference at /arena/elo_reference.json, our council fine-tunes sit below the bases we started from. That finding is read from the published artifact at load time — not typed — and it is the most credible thing we can publish because it contradicts our own product narrative.
What does the cryptographic chain prove — and where does it stop?
A complete-looking manifest is not a signature. We publish the exact limit of what the chain proves about withheld cards, with no RFC-3161 timestamp authority and no blockchain anchoring behind any card (timestamp_authority: none). Verify the signed envelope yourself under /signed/chain.json with /signed/verify-card.mjs.
What board slots have you not measured?
UNMEASURED slots are published so the gap is visible. A slot stays UNMEASURED when the sample is too small (nothing below thirty usable graded items), when the instrument is not frozen and published, or when legal gold labels are still with counsel. UNMEASURED is a disclosure about us, not a failing grade for anyone else's AI system. Live counts: GET /api/gspc.
Where do you publish what you got wrong?
The corrections ledger at GET /api/corrections records what was wrong, how it was caught, and what changed. Entries are appended; none is edited or deleted. Most were caught by our own instrument turned on its owner — including a verifier that could not observe failure and a card-store verification that returned zero.
What is the difference between MEASURED, UNMEASURED, and REPORTED?
MEASURED is a graded run on our frozen instruments, signed. UNMEASURED is honestly empty, published so the gap is visible. REPORTED is a figure published by someone else, cited and dated, carried for context and left unsigned. A REPORTED number never enters our board and is never averaged with a MEASURED one. Machine-readable set: GET /api/reported.
GET councilof.ai/api/gspc and the counts behind this page at GET councilof.ai/api/state — no account, no key, no permission.