Independent measurement · signed evidence · free to re-check

We measure how AI systems behave,and publish the evidence so you can check it yourself.

Frozen, published tests. Answers are graded by fixed rules, never by another AI. Issued measurement cards are signed; unsigned supporting runs are labelled. Re-check the evidence for free, without an account. Unmeasured stays visible.

23 axis · 23 measuredon the board todayevery declared slot and how many carry a real run
1,230graded rows behind iteach row is one answer a rule graded, not a model's opinion
14model fleets testedone fleet per behavioural axis, all answering the same frozen questions
9runs graded from public factsno model, no fleet, no judgement — a rule reads a public record

Measured is not the same as separated.

A slot counts as measured when a real run sits behind it. Whether the axis actually told two models apart is a second question, and across the 14 model-comparison axes the answer today is 0 separated, 2 tied and 12 untested. A tie stays a tie and an untested axis stays untested; neither is rounded up into a ranking.

Read live from GET /api/gspc as this page rendered; the runs behind it were made behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25. Nothing on this page is a certificate, and a tie between two models stays a tie.

Why you can check us rather than believe us

Six things you can verify about this business before you trust a single number on it.

The six figures in this section are read from their owning public sources as this page loads. If a read fails, we say so rather than substitute a number. Open each source to check its date and limits before citing it.

We say what the board cannot yet tell apart

0 of 14

model-comparison axes separated a leader — 2 tied, 12 untested

A slot is MEASURED when a run sits behind it. That is not a finding that the axis told two models apart, and we do not print it as one.

Read the board, row by row

Anyone can re-check our signed cards, for nothing

335

signed records verified — kind: measured

Pin our public key, recompute the hash, check the signature. It runs offline, needs no account, and it is free forever. Verification is never sold.

Check a record in your browser

Our worst results are published beside our best

0.29

the lowest fleet mean on the board — the care axis, published like every other

A measurement body that publishes only its wins is a marketing department. One of our own low scores is further down this page as a signed card; supporting runs elsewhere may be unsigned and are labelled as such.

See one of our own low scores

Every claim we got wrong is written down

63

corrections published — ledger signature: VALID

What was wrong, how it was caught, what changed, dated. Signed records are superseded, never quietly edited — editing the bytes would break the signature that makes them checkable.

Read the corrections ledger

Timestamp proofs show what is attested and what is pending

602

proofs containing Bitcoin block-header attestations — 1 calendar-pending, of 603

The manifest parses proof bytes; it does not independently check the block headers against a Bitcoin node or verify every named subject file. A pending proof is a submission, not an anchor.

View the free timestamp preview

We take part where the rules are being written

34

participation records, each linked to its own evidence (manifest 2026-09-22)

Participation is not endorsement and a listing is not adoption. We hold no certification under any scheme, and we show no other body's logo to imply one.

See every record, and what it does not prove

The long version of this argument — the method, the machine surface, the films, and the answers on funding, limits and offline verification — is kept in full on one page. Nothing was deleted to shorten this one.

Distribution, with its limits

The published tools travel. Downloads are only the first signal.

This census covers packages in the wider CSOAI and MEOK publishing estate across PyPI, npm and Hugging Face. It counts registry download events, including mirrors and automated traffic. It does not count unique people, active installations, executions or customers.

≥ 2,640,224

gross download events since first release

836 of 837 package counters answered

≥ 386,765

gross download events in registry 30-day windows

836 of 837 package counters answered

Published census dated 23 September 2026 UTC · Partial read; missing counters are not treated as zero. The PyPI source and a separate sample counter disagree, so these are reported registry events, not verified adoption.

Of the cumulative total: MEOK 2,321,201 · CSOAI 166,079 · joint 45,841 · unattributed 107,103. Package ownership labels come from the published census.

Live from GET /api/gspc

The living board

23 axes measured · 14 model fleets · 0 separated leaders · 3 public leader scores · 9 fact runs · TIE is TIE · not a certificate.

Rows are in board order — layout, not rank. Status, family, separation and leader state are printed as the API serves them. A TIE is a TIE. A withheld leader is a state, not an empty cell. Verify is free; a rank is never sold. Measurement, not certification.

23 rows on this table · 23 MEASURED · gspc 15 · financial 8

Models with a public leader score

One entry per axis whose leader the board publishes. Ordered by point estimate on each model's own frozen bank — layout, not a cross-axis rank. A TIE is not a win.

  1. gemma3:12b (base model)

    leads Safety

    94.4%81.9% – 98.5%

    TIEn 36

    A point lead the test could not separate from the fleet.

  2. qwen2.5:0.5b-instruct (base model)

    leads Jail

    59.2%47.5% – 69.8%

    TIEn 71

    A point lead the test could not separate from the fleet.

  3. qwen2.5:7b (base model)

    leads Swarm

    44.4%

    UNTESTEDn 37

3 public leader scores on the board today, counted from the rows above · the board's own count agrees (3) · 11 model-comparison axes withhold their leader (8 EXCLUDED_OWN_MODEL, 3 NO_SIGNED_CARD) · 9 fact runs have no fleet and no leader · nothing is padded and a TIE is not a win.

Hugging Face measured-model results

Third-party Hub cells from /api/hub-cards. This is a separate benchmark instrument from the GSPC board above. Each signed card is an observation; deterministic fact axes do not score models.

Open published Hub dataset

Partial read · 397 retrieved MEASURED cells · population totals withheld

Feed observed 2026-09-24T05:29:19.377Z

These observations may use different frozen banks or instruments. The feed does not identify a common comparison set, so their scores are not ranked or directly comparable. Check each signed card for its bank and instrument hashes.

Up to nine signed measured observations for Governance, sorted by model name; scores may come from different banks.
ModelObserved scoreEvidence
deepseek-ai/DeepSeek-V353.3%n 30Signed card
deepseek-ai/DeepSeek-V3-032446.7%n 30Signed card
deepseek-ai/DeepSeek-V3.263.3%n 30Signed card
deepseek-ai/DeepSeek-V4-Flash46.7%n 30Signed card
deepseek-ai/DeepSeek-V4-Flash-073183.3%n 30Signed card
deepseek-ai/DeepSeek-V4-Pro36.7%n 30Signed card
farbodtavakkoli/OTel-2.0-LLM-31B-IT60%n 30Signed card
google/gemma-2-2b-it26.7%n 30Signed card
google/gemma-2-9b-it56.7%n 30Signed card

Showing up to nine model-name-sorted observations on this axis. The displayed subset is not a top-nine ranking, separation test, winner claim, compliance verdict, or certificate. Open each signed card to verify its own evidence.

Ask it a question, or paste a record.

Name an axis and the board jumps to it. Paste a signed record and it is checked right here. Nothing leaves this device either way.

Functions: VERIFY · BOARD · AXIS {name} · CENSUS {id} · CORRECT · WATCH {id} · COMPUTE · XRPL · SWIFT · JAIL · TRACE · AIBOM · REPRO · ROOT · PQC · OTEL

The number a vendor would bury

Here is one of our own models, scoring single digits.

We trained it. It has been on the board since the day it was measured, issued as a signed card under our published key, and it is not going anywhere. It is not even the bottom: the signed set runs all the way down to zero. A score that only ever goes up is not a measurement — it is marketing with a chart.

The panel beside this is not a screenshot. Your browser fetched the record, pinned our public key from the published identity document, recomputed the record's fingerprint from its own contents and checked the signature. Nothing was sent to us, and nothing needed our permission.

And you do not have to take our word for which record we chose to show you. Every signed record is listed here, each one a single fetch from its own body, so you can go and find the low ones yourself.

9.7%

VALID

sha256 of the canonical body equals the card id, and the signature verifies under did:web:csoai.org#card-attestation-1.

what was tested
care-refusal-protect
which system
clan-csoai-plain:latest
signed on
2026-08-19T09:24:39.152331+00:00
issued by
CSOAI Ltd (UK 16939677)

record id · 82994353b8f94337746ddf73700b0edc425d695d43910dbfeb53d118d5a09a1c

One thing inside this record is out of date, on purpose. It describes the board as “13 measured of 14 quotable”, which was true when it was signed. Signed bytes are never edited here, so the old wording stays inside the signature and the live board above says what is true today. We supersede; we do not overwrite.

In the room, on the record

The standards that will govern this are being written now. We are in those rooms.

Standards bodies, public registries, scholarly identifiers and filings on the public record. Every entry below names what it proves, what it does not prove, and the evidence you can open for yourself — because a membership logo with nothing behind it is the exact thing this business exists to make unnecessary.

Participation is not endorsement, and a listing is not adoption. Every entry links to its evidence. Manifest as of , unsigned; re-checked by scripts/memberships-check.mjs.