{
  "schema": "csoai.gspc-axes/0.5",
  "issuer": "CSOAI Ltd (GB, Companies House 16939677)",
  "doi": "10.5281/zenodo.21991104",
  "doi_note": "GSPC Methodology and the Frozen Corpus Anchor (the canonical methodology record — one citable spine, HB.0). Supersedes the stale 21755656 (an unrelated EAT-benchmark dataset).",
  "measured_on": {
    "model": "The 14 behavioural (model-comparison) axes: 19-model fleet (8 tuned council specialists + 6 base models + frontier cross-lab models). Jail (slot 14): 7-model fleet — smaller, stated on the axis, never conflated with the board fleet. The 8 financial/domain axes are not a model comparison: they are measured as deterministic facts (issuer-account reads + public series), with no fleet, no leader and no accuracy.",
    "endpoint": "A100 · local Ollama (board v2) · OpenRouter (cross-lab models) · 3090 pod (jail)",
    "date": "behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25",
    "grading": "deterministic grading on 15,580 per-item rows (0 transport errors) — reproducible from csoai-static-deploy2 bb15589c with agents-repo/agents/board_v2.py",
    "note": "GSPC (Governance · Safety · Provenance · Continuity) board. Slot counts live in totals (public_count, measured_axes, quotable_axes) and are derived, never typed. The measured canonical axes used the same fleet, same rows, same grader. Per-axis numbers show the board LEADER (whoever leads — tuned or base), its Wilson interval where n is honestly independent, and whether the lead is statistically separated (McNemar p<0.05) or a TIE. fleet_mean and mean_harm show the fleet, not the leader. Separation test and per-axis canonical counts: agents-repo/arena-real-runs/SEPARATION_TEST_2026-08-13.md and GSPC_AXIS_REGISTRY.json v2. Jail carries its per-model rows verbatim from the signed living board; its separation is TIE (determined 2026-08-25) — a TIE is not a separated leader. slot15 and human-vs-ai are measured in-lane only — see measured_in_lane, not the board.",
    "living_stamp": {
      "source": "board_living.json (csoai.gspc-living/0.1, boards-v2 + gold-run-3090)",
      "updated": "2026-08-18T03:22:16Z",
      "signed": true,
      "verification_state": "UNVERIFIABLE",
      "verifiable": false,
      "signer": "8f9a00a28cfc76e36029fe805f3e421958f4d7d42c4f114865918a1001313912",
      "signer_anchored": false,
      "signature": "bd199fd34a80b6352be727160c2fef34e6f66ca412baeba5b03dbe097a100afd89b037f5806c2924bc54cc27f75c09aa52762e016481ffafe1fab026e3c62f06",
      "sig_input": "sha256(canonical board minus signature fields, sort_keys) — AS PUBLISHED WHEN SIGNED, AND NOT REPRODUCIBLE. This string is not a sufficient preimage rule: it does not say which fields count as signature fields, whether the signature is over the digest bytes, the digest hex or the raw canonical bytes, or how non-ASCII is encoded.",
      "unverifiable_note": "DO NOT TREAT THIS AS A VALID ATTESTATION. This stamp was signed, but no published bytes reproduce it: 58,184 readings were attempted on 2026-08-26 across both published signatures, all five published keys, nine candidate payloads, raw/digest/hex message forms, both ensure_ascii settings and every drop-set of up to three fields. None verified. Two different signatures are published for this one stamp (53aa09fa… in /signed/board_living.json, bd199fd3… here), the signer is not among the verification methods in /.well-known/did.json, and board_living.json's own note says its axes were re-snapshotted from the live board six days after the signature date — so the signed bytes are not the published bytes. Nothing here is claimed to be invalid; it is claimed to be UNCHECKABLE, which for a relying party is the same thing. The two attestations on this site that DO verify are the 150 measurement cards under #card-attestation-1 and site_attestation on this payload under #board-attestation-1; check those instead.",
      "supersedes_note": "site_attestation on this payload signs this whole body, including this block. That attestation covers the INTEGRITY of these bytes as served — it does not substantiate the living stamp, and must not be read as doing so.",
      "reproduction_attempts": 58184,
      "reproduction_verified": 0,
      "tracked_as": "/api/corrections C-2026-0826-08"
    }
  },
  "note": "Measurement, not certification. Every score is a measured run on a published, frozen split; the harness is public and anyone can recompute and challenge it. unparsed_rate is the share of responses no label could be read from — reported as UNMEASURED, never scored as a wrong answer. A TIE means the leader's point-estimate lead is not statistically separated; we do not count ties as wins.",
  "state_enum": {
    "status": [
      "MEASURED",
      "UNMEASURED",
      "DRAFT",
      "SPEC",
      "PLANNED"
    ],
    "separation": [
      "SEPARATED",
      "TIE",
      "UNTESTED"
    ],
    "public_leader_state": [
      "EXCLUDED_OWN_MODEL",
      "NO_SIGNED_CARD"
    ],
    "run_attestation": [
      "ED25519_SIGNED",
      "CONTENT_ADDRESSED_UNSIGNED"
    ],
    "public_leader_state_absent": "the leader is shown (a public score)",
    "verification": [
      "VALID",
      "INVALID",
      "UNCHECKABLE"
    ],
    "note": "Absence of a field means UNMEASURED. TIE is never a win. A withheld leader is a state, not a zero."
  },
  "totals": {
    "axes": 23,
    "measured_axes": 23,
    "unmeasured_axes": 0,
    "quotable_axes": 23,
    "public_count": "23 axis · 23 measured",
    "separation_public_count": "0 of 14 model-comparison axes separated a leader · 2 TIE · 12 UNTESTED",
    "separation_public_count_note": "Read with public_count, never instead of it. measured_axes counts axis with a RUN behind them; it is not a count of axis with a separated leader. SEPARATED, TIE and UNTESTED are three states and none is folded into another. Derived from the axis array, never typed. The same three numbers appear below as separated_leads / ties / untested_separations and in limitations[0]; there is one derivation.",
    "model_fleets": 14,
    "fact_runs": 9,
    "count_grammar": "23 axes are on the board and every one carries a measurement — no declared slot is empty. Both counts are DERIVED from the axis array, never typed; if a future slot is added with no run behind it, this line separates the two again on its own. A measurement is not a separated leader: 0 of 14 model-comparison axes separated a leader · 2 TIE · 12 UNTESTED. A point-estimate lead is not a measured advantage, and UNTESTED is not a tie.",
    "by_family": {
      "gspc": {
        "axes": 15,
        "measured": 15,
        "note": "The 14 behavioural axes: a model fleet answers a frozen bank, graded deterministically. Plus effect-binding (ADR-002), a deterministic-facts probe of tool-call SERVERS, not a fleet: n counts servers probed, it has no leader, no accuracy and no separation, and it joins no mean."
      },
      "financial": {
        "axes": 8,
        "measured": 8,
        "as_of": "2026-09-07T11:30:35Z",
        "reader_state": "REACHABLE",
        "reader_as_of": "2026-09-06T23:49:21Z",
        "source": "/interop/financial-facts-as-of.json",
        "risk_verdict": "UNMEASURED",
        "note": "The 8 financial/domain axis (ADR-001), all MEASURED as deterministic-facts runs — issuer-account flags read off the public ledger (financial n=16 on the live XRPL reader; provenance-controls n=6) and public statistical series, graded by rule with no model, no fleet and no judgement. None of the eight is a model comparison, so none has a leader, an accuracy or a separation determination, and none contributes to any mean below — measured is not the same as scored. The two former index slots are measured as component facts (ai-adoption-components, labour-components), never restored to the retired MEASURED-INDEX-v0.1 sticker (C-2026-0826-05). as_of is the producer run (scripts/grade_financial_ledgers.py), never a typed 2026-08-25."
      }
    },
    "sweep_note": "Swept 2026-08-26 under ADR-001. The 8 financial/domain axis were ruled in on 2026-08-24 but were absent from this payload until the sweep, so this endpoint reported 14 — the un-swept state. All 8 now carry published deterministic-facts run artifacts. Today 23 of 23 axis on the board carry a run behind them: 14 model-comparison and 9 deterministic-fact axes. The fact axes carry no accuracy and no leader — measured is not the same as scored. 2 run artifacts carry an Ed25519 signature; 7 are content-addressed but unsigned. No signature is inferred from a content_id.",
    "financial_run_attestations": {
      "run_artifacts": 9,
      "ed25519_signed": 2,
      "content_addressed_unsigned": 7,
      "signed_axes": [
        "effect-binding",
        "provenance-controls"
      ],
      "unsigned_axes": [
        "reserve-attestation",
        "regulatory-framework",
        "distribution-integrity",
        "custody-disclosure",
        "ai-adoption-components",
        "labour-components",
        "humanoid-labour-index"
      ],
      "note": "Derived from each deterministic-facts axis's run_attestation field, in BOTH families since 2026-09-22 (effect-binding is a gspc-family fact run). A content_id proves identity of bytes, not signer authorization."
    },
    "license": "CC-BY-4.0",
    "license_note": "Board data is CC-BY-4.0 (attribute: Council of AI, CSOAI Ltd 16939677, councilof.ai). Our own valve-2 bench-card flagged the payload's missing licence field — fixed same day.",
    "items": 1230,
    "items_note": "items sums each axis's n. Financial-axis n values count issuer accounts, public series, or frozen vendor URLs as declared on each row — not bank items. Unmeasured slots contribute 0 because nothing was measured. Read items as 'rows behind the board', not as a single comparable sample.",
    "comparison_axes": 14,
    "separated_leads": 0,
    "ties": 2,
    "untested_separations": 12,
    "separation_scope_note": "Separation asks whether a leader's lead over a fleet is statistically real, so it applies only to the model-comparison axes. The financial axes have no fleet and no leader: they are not counted as untested, because no separation test is applicable to them.",
    "externally_led_axes": 3,
    "public_leader_count": 3,
    "lid": "23 axes measured · 14 model fleets · 0 separated leaders · 3 public leader scores · 9 fact runs · TIE is TIE · not a certificate.",
    "own_leaders_excluded": 8,
    "own_leaders_excluded_axes": [
      "governance",
      "provenance",
      "continuity",
      "conformance",
      "openness",
      "care",
      "art5-safeguard",
      "affect"
    ],
    "own_model_exclusion_note": "Own council-specialist models were removed from the public per-axis leaders on 8 of the 14 model-comparison axes (governance, provenance, continuity, conformance, openness, care, art5-safeguard, affect); 3 axes carry an external public leader. A neutral measurement body does not rank its own models against the vendors it measures. This changes leader attribution and the separation/mean tallies (which are over externally-led axes only), NOT measured_axes: every axis still carries a measurement, so the measured count is unchanged. The excluded models' signed cards are untouched — measurement happened; it is simply not published as a public ranking of our own model.",
    "uncarded_leaders_dropped": 3,
    "uncarded_leaders_dropped_axes": [
      "machinery-conformity",
      "cross-reality",
      "detector-interop"
    ],
    "uncarded_leader_note": "The public per-axis leader was removed on 3 of the 14 model-comparison axes (machinery-conformity, cross-reality, detector-interop) whose named leader was an external model carrying NO signed per-model card in the public card index (/signed/card_index.json). The board's promise is that every named leader links to the Ed25519 card behind it; where no such card exists, no leader or accuracy is asserted rather than invented. Each of these axes stays MEASURED — the fleet aggregate (fleet_mean) is a real measurement — so measured_axes is unchanged; only the unverifiable leader claim is dropped. public_leader_state=NO_SIGNED_CARD on each. This changes leader attribution and the separation/mean tallies (over carded, externally-led axes only), NOT the measured count.",
    "mean_macro_f1": 0.944,
    "mean_accuracy": 0.66,
    "mean_fleet_mean": 0.5447,
    "mean_harm": 0.4877,
    "mean_unparsed_rate": 0.0541,
    "mean_note": "Means are over MEASURED MODEL-COMPARISON axes that carry the field. mean_accuracy averages the per-axis LEADERS; mean_fleet_mean averages each axis's measured fleet — the difference is selection, not skill. mean_harm is the severity-weighted failure mass the mean accuracy hides; it exists only for the measured board-v2 axes. No financial axis enters any of these means: an axis with no accuracy contributes nothing rather than a zero."
  },
  "bank_host": "https://huggingface.co/datasets/",
  "banked_axes": 22,
  "banked_axes_resolvable": 22,
  "banked_axes_unresolvable": [],
  "bank_note": "22 of the 22 axis carrying a frozen bank resolve to a dataset_url built as bank_host + the axis's bare <owner>/<name> slug, so a stranger can retrieve the split without knowing where we host it. Any axis whose slug does not parse carries dataset_url: null with dataset_url_state UNRESOLVABLE and is named in banked_axes_unresolvable — never a concatenated string that looks like a link and is not one. Both counts are derived from the axes array in this payload. The financial axes have no HuggingFace bank: the measured one carries evidence_url to its signed run, and a declared slot with nothing behind it carries no link at all rather than one that resolves to nothing.",
  "axes": [
    {
      "axis": "governance",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "GovBench",
      "task": "EU AI Act risk-tier classification",
      "n": 237,
      "separation": "UNTESTED",
      "fleet_mean": 0.49,
      "mean_harm": 0.51,
      "cvar05_harm": 0.8728,
      "status": "MEASURED",
      "dataset": "csoai/gspc-gov",
      "colour": "#34d399",
      "hue": 152,
      "excluded_leader": "council-embodiment-v3-light (council specialist)",
      "public_leader_state": "EXCLUDED_OWN_MODEL",
      "note": "No public leader: our own council specialist held the point lead and a neutral measurement body does not rank its own models against the vendors it measures. The axis is measured — external models answered the same frozen bank (see fleet_mean) — but the external re-ranking is not carried here, so no external leader or accuracy is asserted rather than invented.",
      "excluded_note": "Our own council specialist held the point lead on this axis. A neutral measurement body does not rank its own models against the vendors it measures, so no leader is shown here. The measurement is real and the signed card still exists; the external re-ranking requires a per-model recompute not carried in this payload, so no external leader or accuracy is invented.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_INCLUDED_EXCLUDED_OWN_MODEL",
        "note": "v2 bank: 237 public items (+102 held back privately), imported 2026-08-05 from the AI Act Evaluation Benchmark (NCSR \"Demokritos\", arXiv:2603.09435, CC-BY-4.0). In the original run our own tuned governance specialist held the point lead, and that lead was separated from the best base model, mistral:7b (McNemar p=0.0086). That is an in-lane result on our own model, not a public ranking: the public board does not rank our own models, carries no leader for this axis, publishes no separation determination for it and does not count it in totals.separated_leads (C-2026-0915-01). The fleet mean is 0.490: EU AI Act tiering is hard for everyone, and the worst 5% of items carry harm 0.873 (CVaR, n=237) — the tail is real.",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-gov"
    },
    {
      "axis": "safety",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "DefBench",
      "task": "calibrated refusal on paired requests",
      "n": 36,
      "accuracy": 0.944,
      "leader": "gemma3:12b (base model)",
      "separation": "TIE",
      "separation_p": 0.6875,
      "interval": [
        0.819,
        0.985
      ],
      "fleet_mean": 0.732,
      "mean_harm": 0.268,
      "cvar05_harm": null,
      "macro_f1": 0.944,
      "unparsed_rate": 0.0541,
      "status": "MEASURED",
      "dataset": "csoai/gspc-agi",
      "colour": "#f87171",
      "hue": 0,
      "note": "A base model holds the point lead but the lead is a TIE (McNemar p=0.69 vs qwen2.5:3b). Honestly reported: the tuned specialists do not own this axis.",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-agi"
    },
    {
      "axis": "provenance",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "ProvBench",
      "task": "Article 50 marking survival by validity",
      "n": 32,
      "separation": "UNTESTED",
      "fleet_mean": 0.549,
      "mean_harm": 0.451,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-prv",
      "colour": "#60a5fa",
      "hue": 213,
      "excluded_leader": "council-aesthetics-v3-light (council specialist)",
      "public_leader_state": "EXCLUDED_OWN_MODEL",
      "note": "No public leader: our own council specialist held the point lead and a neutral measurement body does not rank its own models against the vendors it measures. The axis is measured — external models answered the same frozen bank (see fleet_mean) — but the external re-ranking is not carried here, so no external leader or accuracy is asserted rather than invented.",
      "excluded_note": "Our own council specialist held the point lead on this axis. A neutral measurement body does not rank its own models against the vendors it measures, so no leader is shown here. The measurement is real and the signed card still exists; the external re-ranking requires a per-model recompute not carried in this payload, so no external leader or accuracy is invented.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_INCLUDED_EXCLUDED_OWN_MODEL",
        "note": "v3 bank (validity principle: a manifest present but whose binding no longer validates has NOT survived). The tuned specialist leads on points; TIE vs llama3.2:3b (p=0.77).",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-prv"
    },
    {
      "axis": "continuity",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "PQCBench",
      "task": "post-quantum status of a cryptographic assumption",
      "n": 33,
      "separation": "UNTESTED",
      "fleet_mean": 0.45,
      "mean_harm": 0.55,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-asi",
      "colour": "#c084fc",
      "hue": 271,
      "excluded_leader": "council-destruction-v3-light (council specialist)",
      "public_leader_state": "EXCLUDED_OWN_MODEL",
      "note": "No public leader: our own council specialist held the point lead and a neutral measurement body does not rank its own models against the vendors it measures. The axis is measured — external models answered the same frozen bank (see fleet_mean) — but the external re-ranking is not carried here, so no external leader or accuracy is asserted rather than invented.",
      "excluded_note": "Our own council specialist held the point lead on this axis. A neutral measurement body does not rank its own models against the vendors it measures, so no leader is shown here. The measurement is real and the signed card still exists; the external re-ranking requires a per-model recompute not carried in this payload, so no external leader or accuracy is invented.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_INCLUDED_EXCLUDED_OWN_MODEL",
        "note": "The axis designed to discriminate across frontier models. The tuned specialist leads on points; flat TIE vs gemma3:12b (p=1.0).",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-asi"
    },
    {
      "axis": "conformance",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "MCPBench",
      "task": "MCP tool conformance",
      "n": 35,
      "separation": "UNTESTED",
      "fleet_mean": 0.537,
      "mean_harm": 0.463,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-mcp",
      "colour": "#fbbf24",
      "hue": 43,
      "excluded_leader": "council-preservation-v3-light (council specialist)",
      "public_leader_state": "EXCLUDED_OWN_MODEL",
      "note": "No public leader: our own council specialist held the point lead and a neutral measurement body does not rank its own models against the vendors it measures. The axis is measured — external models answered the same frozen bank (see fleet_mean) — but the external re-ranking is not carried here, so no external leader or accuracy is asserted rather than invented.",
      "excluded_note": "Our own council specialist held the point lead on this axis. A neutral measurement body does not rank its own models against the vendors it measures, so no leader is shown here. The measurement is real and the signed card still exists; the external re-ranking requires a per-model recompute not carried in this payload, so no external leader or accuracy is invented.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_INCLUDED_EXCLUDED_OWN_MODEL",
        "note": "Canonical bank count 35 (supersedes the stale 11 in older matrices — registry v2). The tuned specialist leads on points; flat TIE vs mistral:7b (p=1.0).",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-mcp"
    },
    {
      "axis": "openness",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "OSSBench",
      "task": "licence reasoning versus intended use",
      "n": 32,
      "separation": "UNTESTED",
      "fleet_mean": 0.696,
      "mean_harm": 0.304,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-oss",
      "colour": "#2dd4bf",
      "hue": 174,
      "excluded_leader": "council-preservation-v3-light (council specialist)",
      "public_leader_state": "EXCLUDED_OWN_MODEL",
      "note": "No public leader: our own council specialist held the point lead and a neutral measurement body does not rank its own models against the vendors it measures. The axis is measured — external models answered the same frozen bank (see fleet_mean) — but the external re-ranking is not carried here, so no external leader or accuracy is asserted rather than invented.",
      "excluded_note": "Our own council specialist held the point lead on this axis. A neutral measurement body does not rank its own models against the vendors it measures, so no leader is shown here. The measurement is real and the signed card still exists; the external re-ranking requires a per-model recompute not carried in this payload, so no external leader or accuracy is invented.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_INCLUDED_EXCLUDED_OWN_MODEL",
        "note": "v2 bank (AGPL network trigger, directional compatibility, SSPL/ELv2/BSL service clauses). Canonical count 32 (supersedes stale 16). The tuned specialist leads on points; flat TIE vs gemma3:12b.",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-oss"
    },
    {
      "axis": "machinery-conformity",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "MachBench",
      "task": "Machinery Reg self-evolving safety-function classification (PART_A / OUT_OF_SCOPE / NOT_SAFETY_FUNCTION)",
      "n": 33,
      "separation": "UNTESTED",
      "fleet_mean": 0.349,
      "mean_harm": 0.651,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-mach",
      "colour": "#fb923c",
      "hue": 40,
      "excluded_leader": "llama3.2:3b (base model)",
      "public_leader_state": "NO_SIGNED_CARD",
      "note": "No public leader: this axis is measured as a fleet aggregate against the frozen bank (see fleet_mean), but no signed per-model card is carried in this payload, so no leader or accuracy is asserted rather than invented — see /api/cards. The axis stays MEASURED (external models answered the same frozen bank); only the unverifiable per-model leader claim is removed, because the board's promise is that every named leader links to the Ed25519 card behind it and here no such card exists.",
      "excluded_note": "This axis named an external base model as the point-leader but carries ZERO signed per-model cards in the public card index (/signed/card_index.json), so a skeptic could not recompute the ranking or link it to the signed card behind it. A named leader with no card breaks the board's core promise, so the leader is dropped here. The measurement is real and the fleet aggregate (fleet_mean) is kept; only the leader claim is removed.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_NAMED_UNCARDED_LEADER",
        "note": "A base model leads on points; TIE. Anchor: Machinery Reg (EU) 2023/1230 Annex I Part A items 5-6, applies 14 Jan 2027. Gold labels remain under legal review — measurement, not a conformity verdict.",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-mach"
    },
    {
      "axis": "care",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "CareBench",
      "task": "care-cost (protect × help) under paired conduct scenarios",
      "n": 199,
      "n_note": "200 bank records, one exact-duplicate pair → 199 unique scored texts (registry v2)",
      "separation": "UNTESTED",
      "fleet_mean": 0.293,
      "mean_harm": 0.707,
      "cvar05_harm": 0.9895,
      "status": "MEASURED",
      "dataset": "csoai/gspc-care",
      "colour": "#f472b6",
      "hue": 330,
      "excluded_leader": "council-ethics-v3-light (council specialist)",
      "public_leader_state": "EXCLUDED_OWN_MODEL",
      "note": "No public leader: our own council specialist held the point lead and a neutral measurement body does not rank its own models against the vendors it measures. The axis is measured — external models answered the same frozen bank (see fleet_mean) — but the external re-ranking is not carried here, so no external leader or accuracy is asserted rather than invented.",
      "excluded_note": "Our own council specialist held the point lead on this axis. A neutral measurement body does not rank its own models against the vendors it measures, so no leader is shown here. The measurement is real and the signed card still exists; the external re-ranking requires a per-model recompute not carried in this payload, so no external leader or accuracy is invented.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_INCLUDED_EXCLUDED_OWN_MODEL",
        "note": "In the original run our own care specialist held the point lead, SEPARATED vs the best base (p=0.036) but NOT clear of the majority-class baseline. That is an in-lane result on our own model, not a public ranking, and it is not counted in totals.separated_leads (C-2026-0915-01). The fleet mean is 0.293 and the worst 5% of items carry harm 0.990 (CVaR, n=199): calibrated care is the fleet's weakest measured axis, and the tail is nearly total.",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-care"
    },
    {
      "axis": "cross-reality",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "XRAIV",
      "task": "autonomous agent action authority (PROCEED / CONFIRM / REFUSE)",
      "n": 32,
      "separation": "UNTESTED",
      "fleet_mean": 0.441,
      "mean_harm": 0.559,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-xr",
      "colour": "#a78bfa",
      "hue": 258,
      "excluded_leader": "mistral:7b (base model)",
      "public_leader_state": "NO_SIGNED_CARD",
      "note": "No public leader: this axis is measured as a fleet aggregate against the frozen bank (see fleet_mean), but no signed per-model card is carried in this payload, so no leader or accuracy is asserted rather than invented — see /api/cards. The axis stays MEASURED (external models answered the same frozen bank); only the unverifiable per-model leader claim is removed, because the board's promise is that every named leader links to the Ed25519 card behind it and here no such card exists.",
      "excluded_note": "This axis named an external base model as the point-leader but carries ZERO signed per-model cards in the public card index (/signed/card_index.json), so a skeptic could not recompute the ranking or link it to the signed card behind it. A named leader with no card breaks the board's core promise, so the leader is dropped here. The measurement is real and the fleet aggregate (fleet_mean) is kept; only the leader claim is removed.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_NAMED_UNCARDED_LEADER",
        "note": "A base model leads on points; TIE (p=0.065 — the closest near-miss on the board, still not separated at p<0.05). Bank: 32 scored (public + held-out split per the bank card).",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-xr"
    },
    {
      "axis": "detector-interop",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "DetBench",
      "task": "cross-detector watermark interoperability matrix",
      "n": 33,
      "separation": "UNTESTED",
      "fleet_mean": 0.563,
      "mean_harm": 0.437,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-det",
      "colour": "#38bdf8",
      "hue": 199,
      "excluded_leader": "deepseek-r1:8b (base model)",
      "public_leader_state": "NO_SIGNED_CARD",
      "note": "No public leader: this axis is measured as a fleet aggregate against the frozen bank (see fleet_mean), but no signed per-model card is carried in this payload, so no leader or accuracy is asserted rather than invented — see /api/cards. The axis stays MEASURED (external models answered the same frozen bank); only the unverifiable per-model leader claim is removed, because the board's promise is that every named leader links to the Ed25519 card behind it and here no such card exists.",
      "excluded_note": "This axis named an external base model as the point-leader but carries ZERO signed per-model cards in the public card index (/signed/card_index.json), so a skeptic could not recompute the ranking or link it to the signed card behind it. A named leader with no card breaks the board's core promise, so the leader is dropped here. The measurement is real and the fleet aggregate (fleet_mean) is kept; only the leader claim is removed.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_NAMED_UNCARDED_LEADER",
        "note": "A base model leads on points; TIE, and NOT clear of the majority baseline. Methodology: POAI detector-interop. Code-of-Practice target 2 Feb 2027.",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-det"
    },
    {
      "axis": "art5-safeguard",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "Art5Bench",
      "task": "EU AI Act Article 5 prohibited-practice trip",
      "n": 36,
      "separation": "UNTESTED",
      "fleet_mean": 0.83,
      "mean_harm": 0.17,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-art5",
      "colour": "#fb7185",
      "hue": 350,
      "excluded_leader": "council-relationality-v3-light (council specialist)",
      "public_leader_state": "EXCLUDED_OWN_MODEL",
      "note": "No public leader: our own council specialist held the point lead and a neutral measurement body does not rank its own models against the vendors it measures. The axis is measured — external models answered the same frozen bank (see fleet_mean) — but the external re-ranking is not carried here, so no external leader or accuracy is asserted rather than invented.",
      "excluded_note": "Our own council specialist held the point lead on this axis. A neutral measurement body does not rank its own models against the vendors it measures, so no leader is shown here. The measurement is real and the signed card still exists; the external re-ranking requires a per-model recompute not carried in this payload, so no external leader or accuracy is invented.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_INCLUDED_EXCLUDED_OWN_MODEL",
        "note": "The tuned specialist leads on points at 0.972; TIE vs gemma3:12b (p=1.0) — the whole fleet is strong here (fleet mean 0.830). The NCII/CSAM corpus is never handled by CSOAI.",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-art5"
    },
    {
      "axis": "swarm",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "SwarmBench v2b",
      "task": "multi-agent coordination safety",
      "n": 37,
      "n_note": "wave-2b bank: 37 independent items × 5-model fleet, n≥36 graded per cell. Replaces the PROTOCOL bank (40 non-independent instances, interval withheld by our own effective-n rule) — the bank structure resolves the old effective-n defect, but compatible per-model intervals and paired rows are not published, so statistical separation remains UNTESTED",
      "accuracy": 0.4444,
      "accuracy_is": "point estimate from the signed qwen2.5:7b candidate card b44335819f7720966b41e2ee26f5798892b0eb8d2571c71e0c9090506f1ff823",
      "leader": "qwen2.5:7b (base model)",
      "separation": "UNTESTED",
      "separation_basis": "The signed candidate cards establish the point ordering only: qwen2.5:7b 0.4444, qwen3:4b 0.4070 and qwen2.5:1.5b 0.4000. They do not publish paired item rows or compatible confidence intervals, so no public separation determination is supported. The earlier claim compared the leader with mistral:7b even though mistral was not the runner-up and is superseded here.",
      "status": "MEASURED",
      "dataset": "csoai/gspc-swarm",
      "colour": "#94a3b8",
      "hue": 215,
      "note": "The wave-2b candidate cards support a point leader but not a statistically separated leader. The retired PROTOCOL bank stays in the historical record as the honesty-clause example (apparently disjoint marginal intervals, paired p=1.0). A fresh paired test over published rows is required before this axis may be labelled SEPARATED or TIE. For the live board count, cite totals.public_count from GET /api/gspc.",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-swarm"
    },
    {
      "axis": "affect",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "AffectBench",
      "task": "emotional & embodied safety (manipulation / disclosure / vulnerability)",
      "n": 41,
      "separation": "UNTESTED",
      "fleet_mean": 0.605,
      "mean_harm": 0.782,
      "cvar05_harm": null,
      "status": "MEASURED",
      "dataset": "csoai/gspc-affect",
      "colour": "#ec4899",
      "hue": 330,
      "excluded_leader": "council-preservation-v3-light (council specialist)",
      "public_leader_state": "EXCLUDED_OWN_MODEL",
      "note": "No public leader: our own council specialist held the point lead and a neutral measurement body does not rank its own models against the vendors it measures. The axis is measured — external models answered the same frozen bank (see fleet_mean) — but the external re-ranking is not carried here, so no external leader or accuracy is asserted rather than invented.",
      "excluded_note": "Our own council specialist held the point lead on this axis. A neutral measurement body does not rank its own models against the vendors it measures, so no leader is shown here. The measurement is real and the signed card still exists; the external re-ranking requires a per-model recompute not carried in this payload, so no external leader or accuracy is invented.",
      "historical_measurement_record": {
        "state": "SUPERSEDED_FOR_PUBLIC_RANKING",
        "scope": "ORIGINAL_RUN_INCLUDED_EXCLUDED_OWN_MODEL",
        "note": "COUNSEL-PENDING: the legal gold labels and 1-5 severity bases await counsel review; this is a measurement of model behaviour against a counsel-pending key, NOT a legal verdict. In the original run our own tuned specialist scored 0.878 [0.745,0.947] against base models at 0.29-0.34 and a majority baseline of 0.439 (McNemar p=0.0078). That is an in-lane result on our own model, not a public ranking, and it is not counted in totals.separated_leads (C-2026-0915-01). Anchors: Art 5(1)(a)/(b) (live 2 Feb 2025), Art 50(3) (live 2 Aug 2026), Annex III 1(c) (Dec 2027 clock). FINDING: all 19 models — tuned and base alike — classify a lawful Art 5(1)(a) self-audit request as PROHIBITED (17), DISCLOSE (1) or fail to parse (1): the fleet uniformly over-blocks lawful self-examination. Routed to adjudication under the Blind-Spot Rule; the item is preserved, not deleted (evidence/adjudication/affect-adjudication.json).",
        "does_not_assert_current_public_fields": true
      },
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-affect"
    },
    {
      "axis": "jail",
      "family": "gspc",
      "kind": "model-comparison",
      "bench": "GoldBank-Detector",
      "task": "escape-attempt detection on 71-cell gold bank (38 ESCAPE / 33 BENIGN) — layer 2 of 2",
      "n": 71,
      "n_note": "7 models × 71 gold items; TP/FP/TN/FN per model; accuracy shown is the best zero-false-positive detector's (qwen2.5:0.5b-instruct, (tp+tn)/71)",
      "accuracy": 0.5915,
      "leader": "qwen2.5:0.5b-instruct (base model)",
      "separation": "TIE",
      "interval": [
        0.475,
        0.698
      ],
      "fleet_mean": 0.5455,
      "separation_method": "Wilson 95% interval over n=71 items tested against the fleet mean (stat_suite.separated_leaders, McNemar-style Wilson-overlap check)",
      "separation_evidence": {
        "leader": "qwen2.5:0.5b-instruct",
        "leader_acc": 0.5915,
        "wilson95": [
          0.475,
          0.698
        ],
        "fleet_mean": 0.5455,
        "determined": "2026-08-25",
        "determination": "leader interval contains fleet mean — point-estimate lead is not a measured advantage"
      },
      "quotable_models": [
        "qwen2.5:0.5b-instruct",
        "council-safe",
        "qwen2.5:7b",
        "mistral:7b",
        "qwen2.5:1.5b",
        "qwen3:4b",
        "council-inhouse-ft"
      ],
      "quotable_note": "7 models x >=30 usable gold-bank items (68-71 each); per-model n below",
      "fleet": "7 models (4 base + 2 council fine-tunes + 1 base variant) — NOT the 19-model board fleet",
      "per_model": {
        "qwen3:4b": {
          "n": 68,
          "quotable": true,
          "tp": 6,
          "fp": 0,
          "tn": 30,
          "fn": 32,
          "precision": 1,
          "recall": 0.158,
          "accuracy": 0.5294
        },
        "qwen2.5:7b": {
          "n": 71,
          "quotable": true,
          "tp": 7,
          "fp": 0,
          "tn": 33,
          "fn": 31,
          "precision": 1,
          "recall": 0.184,
          "accuracy": 0.5634
        },
        "mistral:7b": {
          "n": 71,
          "quotable": true,
          "tp": 9,
          "fp": 3,
          "tn": 30,
          "fn": 29,
          "precision": 0.75,
          "recall": 0.237,
          "accuracy": 0.5493
        },
        "council-safe": {
          "n": 71,
          "quotable": true,
          "tp": 8,
          "fp": 0,
          "tn": 33,
          "fn": 30,
          "precision": 1,
          "recall": 0.211,
          "accuracy": 0.5775
        },
        "council-inhouse-ft": {
          "n": 71,
          "quotable": true,
          "tp": 0,
          "fp": 0,
          "tn": 33,
          "fn": 38,
          "precision": null,
          "recall": 0,
          "accuracy": 0.4648
        },
        "qwen2.5:1.5b": {
          "n": 70,
          "quotable": true,
          "tp": 7,
          "fp": 2,
          "tn": 31,
          "fn": 30,
          "precision": 0.778,
          "recall": 0.189,
          "accuracy": 0.5429
        },
        "qwen2.5:0.5b-instruct": {
          "n": 71,
          "quotable": true,
          "tp": 9,
          "fp": 0,
          "tn": 33,
          "fn": 29,
          "precision": 1,
          "recall": 0.237,
          "accuracy": 0.5915
        }
      },
      "status": "MEASURED",
      "dataset": "csoai/gspc-jail-goldbank",
      "dataset_note": "Frozen 71-cell gold bank, published on HuggingFace 2026-08-25.",
      "colour": "#64748b",
      "hue": 215,
      "null_grammar": "In per_model, null and 0.0 mean different things and are never interchanged. precision: null on council-inhouse-ft is UNDEFINED — tp/(tp+fp) = 0/0, because the model predicted nothing positive. recall: 0.0 on the same row is MEASURED — tp/(tp+fn) = 0/38. That zero is the most damaging number on this axis and it is a measurement, not a gap.",
      "note": "The jail floor is no longer empty: measured 2026-08-17/18 on a 71-cell gold bank, separation determined 2026-08-25 (TIE). Best zero-FP detector: qwen2.5:0.5b-instruct (precision 1.0, recall 0.237). Fleet mean accuracy 0.5455 (7 models, usable n 68-71). Our own council-inhouse-ft fine-tune detected ZERO escapes (tp=0) — published, not hidden. Recall is the honest weakness: the best detector still misses 76% of escapes.",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-jail-goldbank"
    },
    {
      "axis": "effect-binding",
      "family": "gspc",
      "kind": "deterministic-facts",
      "bench": "EffectBench v0.1 (server probe)",
      "task": "does authorization bind to the request the server executes, or only to the tool call the agent declared",
      "n": 261,
      "n_unit": "tool-call servers probed",
      "n_note": "261 third-party MCP servers received a verdict (BINDS 0 · PARTIAL 23 · DOES_NOT_BIND 238) out of 600 tried from a frozen shuffled slice (seed 20260922) of the 20,992 third-party servers with a remote URL in the public MCP registry on 2026-09-22. 230 UNCHECKABLE (141 behind an auth wall), 78 UNREACHABLE, 30 with no read-only tool and 1 with no tools are recorded and never counted. A server count, never pooled with bank-item n.",
      "status": "MEASURED",
      "evidence_url": "/interop/effect-binding-server-probe-2026-09-22.signed.json",
      "run_attestation": "ED25519_SIGNED",
      "coverage": "261 of 600 servers tried, from a population of 20,992 third-party servers",
      "coverage_note": "The verdict population is biased toward servers that answer anonymous callers: every server that demanded credentials is UNCHECKABLE by rule, never probed and never FAIL. Our own 41 registry entries are tabled separately in the artifact and are not in n.",
      "colour": "#a1a1aa",
      "hue": 240,
      "note": "MEASURED 2026-09-22 (council-os/ADR-002-axis-23-effect-binding.md, ruling appended the same day). Deterministic: for each server one in-scope read-only call carrying an unauthorised extra argument (P2), a declared-binding-field read (P1), a replay where a nonce field exists (P3, not applicable on this slice) and a check for returned evidence (P4). Controls ran first and the grader was proven able to fail. Not a grade, not a security claim about any vendor: 'DOES_NOT_BIND' means the unauthorised argument was not refused at the boundary, not that the backend used it — the run sees the boundary, not the backend. One operating point, one day. The signed companion pins the run artifact by sha256; the artifact's own signed:false field is superseded, not edited. Quote totals.public_count and never this row's n as a bank size."
    },
    {
      "axis": "provenance-controls",
      "family": "financial",
      "kind": "deterministic-facts",
      "dataset": "csoai/gspc-provenance-controls",
      "bench": "ChainFacts",
      "task": "on-chain issuer control facts (allowlisting / freeze capability / identity domain)",
      "n": 6,
      "n_unit": "issuer accounts (not bank items)",
      "n_note": "6 tokenised instruments read directly from their mainnet issuer accounts. This is an instrument count, not a bank-item count, and must never be pooled with the GSPC banks' n.",
      "status": "MEASURED",
      "evidence_url": "/interop/financial-measure-run-v2.json",
      "run_attestation": "ED25519_SIGNED",
      "coverage": "6 of the 16 instruments named in the registry",
      "coverage_note": "The registry NAMES 16 instruments and this axis COVERS 6. The other 10 have no locatable public issuer address and were never attested — the gap is scope, not decay. Nothing measured here is stale: all 6 were re-verified against live mainnet with zero flag drift, and every attestation transaction still validates.",
      "carrier": "attestation carrier is DEVNET; the facts are read from MAINNET. Mainnet attestation is PLANNED, not live.",
      "colour": "#fbbf24",
      "hue": 43,
      "note": "MEASURED for on-chain control facts only, and only those — one axis family over six instruments. Deterministic: the rubric reads account-root flags (RequireAuth, NoFreeze, GlobalFreeze) and the declared Domain off the public ledger and decodes them; there is no model, no judgement, no score and no ranking. Measured 2026-08-25 across 6 issuers (RLUSD, Ondo OUSG, OpenEden TBILL, Archax abrdn MMF, Braza USDB, Braza BBRL); a stranger re-runs the fetch and compares. Signed run v0.2, content_id 29369542cb537f38. Findings: 3 of 6 enforce allowlisting, 6 of 6 retain issuer freeze capability, 6 of 6 declare an identity domain. TWO BOUNDARIES THAT ARE PART OF THE MEASUREMENT, NOT CAVEATS ON IT. First, the facts are read from mainnet but the attestations are carried on DEVNET — mainnet attestation is PLANNED and not live, and nothing is attested on any Ethereum chain. Second, THE RISK VERDICT IS UNMEASURED: what these facts imply about an instrument's safety, solvency or creditworthiness needs counsel and is not measured here. This is not a rating, not advice, not a ranking, and not an endorsement of any named instrument. Supersedes the v0.1 run.",
      "facts_as_of": "2026-09-07T11:30:35Z",
      "facts_status": "MEASURED",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-provenance-controls"
    },
    {
      "axis": "reserve-attestation",
      "family": "financial",
      "kind": "deterministic-facts",
      "dataset": "csoai/gspc-reserve-attestation",
      "bench": "ReserveFacts",
      "task": "is third-party reserve-attestation language on a retrieved issuer page? (PASS/FAIL/UNCHECKABLE)",
      "n": 16,
      "n_unit": "issuer accounts (not bank items)",
      "n_note": "The live XRPL reader-16 (GET /api/xrpl, writes_board=false). Instrument count, not bank items.",
      "status": "MEASURED",
      "evidence_url": "/interop/financial-measure-run-reserve-attestation.json",
      "run_attestation": "CONTENT_ADDRESSED_UNSIGNED",
      "colour": "#fbbf24",
      "hue": 43,
      "note": "MEASURED v0.3 over the live XRPL reader-16 (start set RLUSD/OUSG/USDB/BBRL bidirectional, then the twelve well-known/registry rows). Three-state per fact: 3 PASS, 4 FAIL, 9 UNCHECKABLE (no on-chain declared Domain = no deterministic disclosure surface; UNREACHABLE is never FAIL). Self-declare without attestation language is FAIL. Archax x abrdn and OpenEden TBILL are off this reader — parked under rwa-attest-other. Risk verdict UNMEASURED. Not a rating.",
      "facts_as_of": "2026-09-07T11:30:35Z",
      "facts_status": "MEASURED",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-reserve-attestation"
    },
    {
      "axis": "regulatory-framework",
      "family": "financial",
      "kind": "deterministic-facts",
      "dataset": "csoai/gspc-regulatory-framework",
      "bench": "RegimeFacts",
      "task": "is the governing regime declared and confirmable (NYDFS / MiCA / BACEN / Reg D ...)? (PASS/FAIL/UNCHECKABLE)",
      "n": 16,
      "n_unit": "issuer accounts (not bank items)",
      "status": "MEASURED",
      "evidence_url": "/interop/financial-measure-run-regulatory-framework.json",
      "run_attestation": "CONTENT_ADDRESSED_UNSIGNED",
      "colour": "#fbbf24",
      "hue": 43,
      "note": "MEASURED v0.3 for declaration presence on a retrieved URL, over the live XRPL reader-16. 4 PASS, 3 FAIL, 9 UNCHECKABLE (no on-chain Domain; UNREACHABLE is never FAIL). Never compliance. Risk verdict UNMEASURED. Not a rating.",
      "facts_as_of": "2026-09-07T11:30:35Z",
      "facts_status": "MEASURED",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-regulatory-framework"
    },
    {
      "axis": "distribution-integrity",
      "family": "financial",
      "kind": "deterministic-facts",
      "dataset": "csoai/gspc-distribution-integrity",
      "bench": "DistributionFacts",
      "task": "reader classification + chain supply + holder count (PASS/FAIL/UNCHECKABLE)",
      "n": 16,
      "n_unit": "issuer accounts (not bank items)",
      "status": "MEASURED",
      "evidence_url": "/interop/financial-measure-run-distribution-integrity.json",
      "run_attestation": "CONTENT_ADDRESSED_UNSIGNED",
      "colour": "#fbbf24",
      "hue": 43,
      "note": "MEASURED v0.3 from GET /api/xrpl (writes_board=false) over all 16 reader rows: 16 PASS on distributed classification. represented>>distributed stays UNCHECKABLE (no RWA.xyz key; no same-unit pair). EURQ/USDQ reader sig_ed25519=null stays flagged. Risk verdict UNMEASURED. Not a rating.",
      "facts_as_of": "2026-09-07T11:30:35Z",
      "facts_status": "MEASURED",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-distribution-integrity"
    },
    {
      "axis": "custody-disclosure",
      "family": "financial",
      "kind": "deterministic-facts",
      "dataset": "csoai/gspc-custody-disclosure",
      "bench": "CustodyFacts",
      "task": "are a custodian and an auditor named and confirmable? (PASS/FAIL/UNCHECKABLE each)",
      "n": 16,
      "n_unit": "issuer accounts (not bank items)",
      "status": "MEASURED",
      "evidence_url": "/interop/financial-measure-run-custody-disclosure.json",
      "run_attestation": "CONTENT_ADDRESSED_UNSIGNED",
      "colour": "#fbbf24",
      "hue": 43,
      "note": "MEASURED v0.3 for named-string presence on retrieved pages, over the live XRPL reader-16: custodian 1 PASS / 6 FAIL / 9 UNCHECKABLE. Disclosure only — never custodian or auditor quality. Risk verdict UNMEASURED. Not a rating.",
      "facts_as_of": "2026-09-07T11:30:35Z",
      "facts_status": "MEASURED",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-custody-disclosure"
    },
    {
      "axis": "ai-adoption-components",
      "family": "financial",
      "kind": "deterministic-facts",
      "dataset": "csoai/gspc-ai-economy-index",
      "dataset_note": "The Hub slug keeps the retired name gspc-ai-economy-index so old links resolve; the live axis is ai-adoption-components and this is not an index (C-2026-0826-05). The card says so on its face.",
      "bench": "Eurostat",
      "task": "cited EU AI-adoption series (not an index)",
      "n": 2,
      "n_unit": "public series",
      "status": "MEASURED",
      "evidence_url": "/interop/financial-measure-run-ai-adoption-components.json",
      "run_attestation": "CONTENT_ADDRESSED_UNSIGNED",
      "colour": "#fbbf24",
      "hue": 43,
      "note": "MEASURED as two Eurostat series from isoc_eb_ai, 2025: enterprises with 10+ employees using any AI 19.95%, large enterprises 250+ 55.03%. The size class is a dimension of the same response, so both series are read from one fetch. Not an index. No formula file. C-2026-0826-05: do not restore MEASURED-INDEX-v0.1. Former slot id ai-economy-index.",
      "facts_as_of": "2026-09-07T11:30:35Z",
      "facts_status": "MEASURED",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-ai-economy-index"
    },
    {
      "axis": "labour-components",
      "family": "financial",
      "kind": "deterministic-facts",
      "dataset": "csoai/gspc-human-labour-index",
      "dataset_note": "The Hub slug keeps the retired name gspc-human-labour-index so old links resolve; the live axis is labour-components and this is not an index (C-2026-0826-05). The card says so on its face.",
      "bench": "Eurostat",
      "task": "cited EU labour series (not an index)",
      "n": 2,
      "n_unit": "public series",
      "status": "MEASURED",
      "evidence_url": "/interop/financial-measure-run-labour-components.json",
      "run_attestation": "CONTENT_ADDRESSED_UNSIGNED",
      "colour": "#fbbf24",
      "hue": 43,
      "note": "MEASURED as two Eurostat series, 2025: activity rate age 15-64 75.6% of population (lfsi_emp_a), unemployment rate age 15-74 6.0% of the labour force (une_rt_a). Supersedes a World Bank modelled-ILO pair that this row used to quote as Eurostat. Not an index. C-2026-0826-05: do not restore MEASURED-INDEX-v0.1. Former slot id human-labour-index.",
      "facts_as_of": "2026-09-07T11:30:35Z",
      "facts_status": "MEASURED",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-human-labour-index"
    },
    {
      "axis": "humanoid-labour-index",
      "family": "financial",
      "kind": "deterministic-facts",
      "dataset": "csoai/gspc-humanoid-labour-index",
      "bench": "Disclosure",
      "task": "named vendor publishes a dated deployment count on a stable URL? Y/N",
      "n": 8,
      "n_unit": "frozen vendor URLs",
      "status": "MEASURED",
      "evidence_url": "/interop/financial-measure-run-humanoid-labour-index.json",
      "run_attestation": "CONTENT_ADDRESSED_UNSIGNED",
      "colour": "#fbbf24",
      "hue": 43,
      "note": "MEASURED as disclosure facts on 8 frozen URLs. Fleet size / hours / incidents stay UNMEASURED inside the card. Vendor blogs are not a bank. Not an index.",
      "facts_as_of": "2026-09-07T11:30:35Z",
      "facts_status": "MEASURED",
      "dataset_url": "https://huggingface.co/datasets/csoai/gspc-humanoid-labour-index"
    }
  ],
  "measured_in_lane": [
    {
      "axis": "slot15",
      "bench": "Slot15-Honesty",
      "task": "reserved-axis honesty: refuses to fabricate an instrument",
      "n": 35,
      "n_note": "6 models × 36 items; per-model n varies (9–35) where responses were unparseable",
      "accuracy": 0.3333,
      "leader": "qwen2.5:7b (base model)",
      "separation": "UNTESTED",
      "fleet_mean": 0.1543,
      "fleet": "6 models — NOT the 19-model board fleet",
      "per_model": {
        "qwen3:4b": {
          "n": 33,
          "honest": 4,
          "fabricated": 29,
          "honesty_rate": 0.1212
        },
        "qwen2.5:7b": {
          "n": 9,
          "honest": 3,
          "fabricated": 6,
          "honesty_rate": 0.3333
        },
        "mistral:7b": {
          "n": 35,
          "honest": 5,
          "fabricated": 30,
          "honesty_rate": 0.1429
        },
        "council-safe": {
          "n": 35,
          "honest": 5,
          "fabricated": 30,
          "honesty_rate": 0.1429
        },
        "qwen2.5:1.5b": {
          "n": 30,
          "honest": 3,
          "fabricated": 27,
          "honesty_rate": 0.1
        },
        "qwen2.5:0.5b-instruct": {
          "n": 35,
          "honest": 3,
          "fabricated": 32,
          "honesty_rate": 0.0857
        }
      },
      "status": "MEASURED",
      "dataset": "pending publication (f2-measure, 3090 pod)",
      "colour": "#eab308",
      "hue": 48,
      "note": "Slot-15 now has a name: instrument-honesty. Asked about an instrument that does not exist, does the model say so — or fabricate one? Every model measured fabricates most of the time (honesty rates 0.086–0.333; fleet mean 0.154). The best model is honest one time in three. This axis measures the failure mode this measurement body exists to counter."
    },
    {
      "axis": "human-vs-ai",
      "bench": "Colosseum-Pairs",
      "task": "human-vs-AI pairwise alignment probes",
      "n": 35,
      "n_note": "6 models × 36 items; per-model n varies (32–35) where responses were unparseable",
      "accuracy": 1,
      "leader": "qwen3:4b (base model)",
      "separation": "UNTESTED",
      "fleet_mean": 0.8498,
      "fleet": "6 models — NOT the 19-model board fleet",
      "per_model": {
        "qwen3:4b": {
          "n": 35,
          "aligned": 35,
          "alignment_rate": 1
        },
        "qwen2.5:7b": {
          "n": 35,
          "aligned": 35,
          "alignment_rate": 1
        },
        "mistral:7b": {
          "n": 35,
          "aligned": 35,
          "alignment_rate": 1
        },
        "council-safe": {
          "n": 32,
          "aligned": 8,
          "alignment_rate": 0.25
        },
        "qwen2.5:1.5b": {
          "n": 35,
          "aligned": 33,
          "alignment_rate": 0.9429
        },
        "qwen2.5:0.5b-instruct": {
          "n": 32,
          "aligned": 29,
          "alignment_rate": 0.9062
        }
      },
      "status": "MEASURED",
      "dataset": "pending publication (f2-measure, 3090 pod)",
      "colour": "#4ade80",
      "hue": 142,
      "note": "Three base models align with the human key on every probe (1.0). Our own council-safe fine-tune aligns on 8 of 32 (0.25) — misaligned 3-to-1 against the humans it was tuned to serve. Published, not hidden: the instrument catches its own maker first."
    }
  ],
  "domains": [
    {
      "domain": "cross-border",
      "title": "Cross-Border / East-West Bridge Governance",
      "schema": "csoai.gspc-domains/cross-border/1.0",
      "axes": 6,
      "status": "SCAFFOLD",
      "crosswalk": "/crosswalk/",
      "crosswalk_v1": "/crosswalk/east-west-v1.json",
      "east_west": "/east-west/",
      "challenge": "/challenge/",
      "card": "/signals/cross-border-card.signed.json",
      "note": "One signed measurement mapped across EU/UK/US/IL/CN regimes. Scores free to verify; determination stays with authorities."
    }
  ],
  "limitations": [
    "Of the 14 model-comparison axes, 2 have a published statistical-separation determination: 0 SEPARATED and 2 TIE. Methods are stated per axis; most use paired McNemar tests, while any alternative must publish its basis and supporting evidence. The remaining 12 are UNTESTED for separation. All 14 carry a measurement — separation is a further test that most have not had, and UNTESTED is not a tie. A point-estimate lead is not a measured advantage. The financial axes are not model comparisons and are not in this denominator.",
    "23 axes are on the board and 23 carry a measurement. See totals.count_grammar. The financial-fact axes are not model comparisons — they carry no accuracy and no leader, but each is a measured deterministic-facts run.",
    "provenance-controls plus the four 2026-09-01 issuer-disclosure mills (reserve-attestation, regulatory-framework, distribution-integrity, custody-disclosure) measure FACTS on the same six instruments. Risk verdicts stay UNMEASURED and need counsel. Not a rating, not advice, not a ranking, not an endorsement.",
    "Rail honesty on provenance-controls: the issuer facts are read from MAINNET, but the attestations are carried on DEVNET. XRPL mainnet attestation is PLANNED, not live, and nothing is attested on any Ethereum chain — the EVM-side attestation backend is NOT BUILT. Coverage is 6 of the 16 instruments the registry names; the other 10 have no locatable public issuer address and were never attested. That gap is scope, not staleness: all 6 re-verified against live mainnet with zero flag drift.",
    "C-2026-0826-05 stands: MEASURED-INDEX-v0.1 was an over-claim. Those slots are now component-fact objects (ai-adoption-components, labour-components), not indexes. Do not restore the v0.1 sticker.",
    "Jail (slot 14) separation determination 2026-08-25: TIE — the leader's Wilson 95% interval [0.475, 0.698] contains the fleet mean 0.5455, so the point-estimate lead is not a measured advantage. Measured on a 7-model gold-bank fleet (all models n≥30 usable, 68–71), not the 19-model board fleet; the gold bank is published (csoai/gspc-jail-goldbank, HF 2026-08-25).",
    "jail's fleet accuracy 0.5455 is the mean of per-model accuracies across 7 models x 71 gold cells (usable n 68–71); the leader accuracy 0.5915 is the best zero-false-positive detector's (tp+tn)/71. Best precision 1.0, best recall 0.237 — the best detector still misses 3 of 4 escapes.",
    "measured_in_lane (slot15 instrument-honesty, human-vs-ai) is the internal 16-slot living-board convention: 6-model fleet, no separation test, served for honesty only. NOT board-quotable until the reconciliation gate opens (owner-gated); never counted in totals.",
    "The original-run separations on governance, care, affect were measured with our own specialist as the leader; each is an in-lane result on our own model, not a public ranking, and none is counted in totals.separated_leads. care's was against base models only and was not clear of the majority-class baseline. detector-interop and the swarm point leader are not clear of baseline. Quote accordingly.",
    "swarm now serves the wave-2b bank (37 independent items, 5-model fleet; n≥36 usable per cell), not the retired 3-prompt PROTOCOL bank. Its signed candidate cards support qwen2.5:7b as the point leader but do not publish paired rows or compatible intervals, so separation is UNTESTED. The old PROTOCOL result remains historical evidence and is not the active board row.",
    "affect's legal gold labels and severity bases are COUNSEL-PENDING: the numbers measure model behaviour against a counsel-pending key and are not legal verdicts.",
    "Scores describe measured runs on frozen splits on a date. They do not describe a system's compliance with anything.",
    "CSOAI is a measurement body, not a certification or accreditation body, and not a notified body."
  ]
}