{
  "schema": "csoai.corrections/0.1",
  "policy": "Public, source-maintained corrections record. Each entry states what was wrong, how it was caught, and the fix. No append-only storage property is claimed.",
  "license": "CC-BY-4.0",
  "publisher": "Council of AI (CSOAI Ltd, UK Companies House 16939677)",
  "corrections": [
    {
      "date": "2026-09-23",
      "evidence": [
        "https://councilof.ai/api/gspc"
      ],
      "first_observed_at": "2026-09-23T03:55:06Z",
      "id": "C-2026-0923-02",
      "note": "Promoted from draft D-2026-09-23T03-02 by the owner. HAND-DRAFTED by the arena-separation lane (feat/arena-separation-2026-09-23), not by drift-draft.py's detector, and placed in the same approve-queue so promote-draft.sh D-2026-09-23T03-02 is the only step. kind summary_omits_measured_negative; fingerprint 236d1c31f796639e. No ledger id is assigned until promote-draft.sh runs. Nothing here is a grade or a mark.",
      "reached_the_public": true,
      "what_changed": "Fixed at the cause on branch fix/signed-surface-agreement-2026-09-23 (pushed to the pod bare repo, NOT merged at the time of this entry). functions/api/gspc.ts now derives the separation aggregate ONCE, beside the counts it already derived, so a reader meets the negative at the same moment as the measured count. totals.separation_public_count reads \"0 of 14 model-comparison axis separated a leader, 2 TIE, 12 UNTESTED\", with a note saying to read it WITH public_count and never instead of it. totals.count_grammar, the line the payload already points readers to, now carries the same sentence and the rule behind it: a measurement is not a separated leader, a point-estimate lead is not a measured advantage, and UNTESTED is not a tie. totals.lid, the one line the estate asks readers to quote verbatim and the line the home page renders verbatim, now states \"0 separated leaders\". separated_leads, ties and untested_separations read the same three constants instead of re-deriving them, so the headline, the grammar, the tallies and limitations[0] cannot drift apart - the defect this file already records at C-2026-0922 (G-3, two derivations of one quantity) is not reintroduced. functions/api/gspc.lid-truth.test.ts was extended to parse the new lid number against totals.separated_leads and to assert that the three separation states account for every model-comparison axis with none folded into another.",
      "what_was_wrong": "totals carries axes, measured_axes, unmeasured_axes, quotable_axes and the count line '23 axis · 23 measured', and no aggregate of the separation field at all. The same payload's limitations[0] states the measured position plainly: of the 14 model-comparison axes, 2 TIE, 12 UNTESTED. The count line is the line every other surface quotes, so the figure that travels is the one that cannot carry the negative.",
      "why_it_was_wrong": "A measured axis and an axis that can tell two models apart are different claims, and only the first is derived into totals. Separation is present per-axis in the array and stated in limitations, so the aggregate is derivable from bytes already served; it is simply not derived. This records the omission, not an intent."
    },
    {
      "date": "2026-09-23",
      "evidence": [
        "https://councilof.ai/api/gspc",
        "https://councilof.ai/signals/swarm.signed.json"
      ],
      "first_observed_at": "2026-09-23T03:55:06Z",
      "id": "C-2026-0923-01",
      "note": "Promoted from draft D-2026-09-23T03-01 by the owner. HAND-DRAFTED by the arena-separation lane (feat/arena-separation-2026-09-23), not by drift-draft.py's detector, and placed in the same approve-queue so promote-draft.sh D-2026-09-23T03-01 is the only step. kind measured_surfaces_disagree; fingerprint 7d98363c444e75d8. No ledger id is assigned until promote-draft.sh runs. Nothing here is a grade or a mark.",
      "reached_the_public": true,
      "what_changed": "Fixed at the cause on branch fix/signed-surface-agreement-2026-09-23 (pushed to the pod bare repo, NOT merged at the time of this entry). Establishing which surface was right came first, and the answer is that neither was wrong about its own bytes: the board grades a frozen 37-item SwarmBench v2b bank for per-item accuracy, the signal ranks recorded pairwise arena rounds by win-rate, and on this axis the two fleets share no model at all - the board's leader qwen2.5:7b is not among the three models ranked in the arena. Two determinations were wearing one word. Regenerating the signals under the 0.3 all-other-ranked-models rule that landed earlier the same day does NOT resolve it: nemotron-3-nano:30b's Wilson lower bound 0.796 clears both other ranked models' upper bounds (0.513, 0.435), so the arena verdict stays SEPARATED. The disagreement was never a rule-version artefact. The remedy is that one surface stops claiming the axis's separation. scripts/emit_signals.py (schema csoai.axis-signal/0.4) now joins every signal to its board row on the board's own dataset slug, with no typed crosswalk, and defers status and register to the board's separation verdict; it refuses to sign a signal whose board row it cannot find. The arena determination is not discarded: it stays in full in the elo_ fields, scoped by separation_of, by separation_authority (carrying the board's verdict, leader, bench and n) and by evidence_relation SEPARATE_EVIDENCE, which states in the signed bytes that the two are never added, reconciled or substituted. All 14 per-axis signals were regenerated through the producer and re-signed under did:web:csoai.org#board-attestation-1; no signed artifact was edited in place. swarm's published status moves MEASURED to UNTESTED and the superseded bytes are recorded in the new file's supersedes block. Separately, the signal now publishes register_board_drift on swarm rather than carrying the stale count silently: the axis register still describes the retired 40-item PROTOCOL bank while the board serves 37. That row is published as PUBLISHED_NOT_RECONCILED and deliberately not retyped, because reconciling it also requires majority_baseline re-derived on the current bank, which has not been done and is not invented. Four planted controls in scripts/arena/test_arena_controls.py hold the shape, including one that plants arena evidence that separates on an axis the board has not tested and asserts the chain cannot publish it as MEASURED.",
      "what_was_wrong": "Two surfaces this organisation publishes and signs give different answers to the same question about the same axis. GET /api/gspc reports swarm separation UNTESTED with leader 'qwen2.5:7b (base model)' over n=37. /signals/swarm.signed.json reports elo_separation SEPARATED with elo_leader 'nemotron-3-nano:30b' over 18 decided arena games. A reader asking whether we can tell two models apart on swarm gets two answers and two different model names, both carrying the board signature.",
      "why_it_was_wrong": "They are two different determinations over two different corpora — the board's is a paired McNemar test on the 2026-08-12 fleet run, the signal's is a Wilson interval over the hourly arena rounds — and neither surface says so where the other can be read. Nothing here establishes which is right. Recording the disagreement is the point; a surface that is silent about a second published answer is the defect."
    },
    {
      "id": "C-2026-0922-02",
      "date": "2026-09-22",
      "first_observed_at": "2026-09-22T12:48Z",
      "what_was_wrong": "32 signed measurement cards for the jail axis under /interop/mill-cards-signed/ published an accuracy for jailbreak-escape detection that was never measured. They were graded against a placeholder bank — the pod file /workspace/banks-all/gspc-jail.jsonl, sha256 f0f31f9a…, 41 rows whose prompts were the literal strings \"jail-000\", \"jail-001\" and so on, with no code cell in them. Each model was asked to classify a placeholder token and the exact-label grader scored the reply against the gold label; the published accuracies, 0.0 to 0.9487, are an artifact of which label a model happens to prefer. By 2026-09-22 15:30Z, 24 of the 32 were already superseded — 21 of those by another card from the same placeholder bank, which cured nothing — and 8 were still the live card for their cell: mistral:7b 0.9487 (n=39), phi3.5:3.8b 0.25 (n=40), qwen2.5:1.5b 0.0732 (n=41), mistral-nemo:12b and qwen2.5:0.5b-instruct 0.0488 (n=41), gemma3:4b 0.0256 (n=39), qwen2.5:7b 0.0244 (n=41), qwen3:4b 0 (n=41). All 32 carried status MEASURED and all 32 verify under did:web:csoai.org#board-attestation-1 — the signature was sound over a measurement that was not.",
      "why_it_was_wrong": "The intake allowlist pinned the bank bytes without anyone reading them. verify_runpod_gspc_intake.py proves that the bank used is the bank pinned; it cannot prove that the pinned bytes are a bank. The real bank existed and was public the whole time — HF csoai/gspc-jail-goldbank, samples.jsonl, sha256 0b45b620…, 71 real code cells, 38 ESCAPE and 33 BENIGN — and the board's jail row already cited it, so the number on the card and the bank named beside it were about different things.",
      "what_changed": "Nothing was deleted and no signed byte was edited. The eight cells were re-measured on the goldbank through the same pinned instrument (temperature 0, seed 0, 128 label tokens, the full 71 cells), and seven produced a board-signed MEASURED card that supersedes its placeholder: mistral:7b 0.9853 (n=68), gemma3:4b 0.8169 (n=71), phi3.5:3.8b 0.6308 (n=65), mistral-nemo:12b 0.6056 (n=71), qwen2.5:7b 0.5857 (n=70), qwen2.5:1.5b 0.4648 (n=71), qwen2.5:0.5b-instruct 0.4648 (n=71). The eighth, qwen3:4b, is UNMEASURED: on the real cells it emitted no parsable label on 71 of 71 items, every reply running to the 128-token cap, so n=0 and there is no card to point at — its placeholder is superseded with by_id null, because a card graded on stubs must not stand either way. All 32 placeholder cards remain on disk and keep resolving; the eight supersessions are recorded in /interop/mill-cards-signed/SUPERSEDED.jsonl naming this entry and the two bank digests. The producer is fixed at the source: the jail digest in scripts/runpod_gspc_bank_allowlist.current.json is now the goldbank, the worker and the playlist generator read the goldbank as its published Inspect-shaped rows rather than a rewritten copy, the pod bank file holds those bytes, and the hourly mill halts unless the digest matches — so no further placeholder run can be admitted. Card root re-stamped: 1423 live leaves, merkle_root 8add6156…, recomputed MATCH; its timestamp proof is PENDING at the calendar and not yet anchored to Bitcoin.",
      "reached_the_public": true,
      "note": "The two 0.4648 figures are the same number for the same reason and should not be read as detection: qwen2.5:0.5b-instruct and qwen2.5:1.5b answered BENIGN on all 71 cells, and 33 of the 71 cells are BENIGN. Found by this estate while restarting the mill on 2026-09-22; the first three cures landed the same day, these eight the same afternoon. The 12 hub-mill jail cards graded through provider APIs on the HF bank are a different corpus and are not covered here."
    },
    {
      "id": "C-2026-0922-01",
      "date": "2026-09-18",
      "first_observed_at": "2026-09-18T00:00Z",
      "what_was_wrong": "The 17 September note on the public root's odd-node duplication (corrections/merkle-count-binding-2026-09-17.md in the csoai/councilof-ai-mirror dataset) said that RFC 6962 domain-separation prefixes are the general fix for the padding collision it demonstrated, and that moving the public root to domain separation would close it by construction. Prefixing 0x00 before leaves and 0x01 before nodes while keeping odd-node duplication leaves the collision intact: the forged 306-leaf set and the honest 305-leaf set still hash to one root, because the substitution pairs a leaf with a leaf.",
      "why_it_was_wrong": "The prefixes stop a leaf digest from impersonating an interior digest — a different second-preimage class. What makes Certificate Transparency immune to the padding collision is that RFC 6962 never duplicates an odd node (it splits at the largest power of two below n) and signs tree_size. The note attributed CT's immunity to the wrong mechanism, so a reader building a v2 with prefixes alone would still ship a collidable root.",
      "what_changed": "A dated superseding note, corrections/merkle-count-binding-2026-09-18-SUPERSEDES.md, is published beside the original in the same dataset and under public/corrections/ in the councilof-ai repository, with a standard-library reproduction over the live root run three ways (unprefixed duplicating: collides; prefixed duplicating: still collides; RFC 6962 shape: does not). corrections/SUPERSESSIONS.md points both ways. The original file is not edited. The count-binding guidance in the original (bind card_count; reject index >= card_count) stands and is unaffected.",
      "reached_the_public": true,
      "note": "Caught 2026-09-18 while re-reading the note for the W3C Agent Conformance CG text, published 2026-09-22. Live root at publication: as_of 2026-09-22T08:54:02Z, card_count 305, merkle_root 40ce3833…"
    },
    {
      "id": "C-2026-0920-01",
      "date": "2026-09-20",
      "first_observed_at": "2026-09-20T01:52Z",
      "what_was_wrong": "The site_attestation on GET /api/gspc did not verify under its own published preimage rule. excludeOwnLeader() and dropUncardedLeader() in functions/api/gspc.ts returned leader: undefined as an own property on the 11 axes whose leader is excluded or uncarded; the edge signer's canonical() emits an own undefined property as the literal text \"leader\":undefined, while JSON.stringify — which produces the served bytes — drops the key entirely. The signed bytes were therefore unreconstructable from the served bytes by anyone. An outside reconciliation on 2026-09-20 tried 11 preimage variants across two independent implementations (Node with the signer's exact canonical(); Python ensure_ascii both ways); none verified, while the same payload's living_stamp verified under the same pinned key.",
      "why_it_was_wrong": "The signer and the serializer disagreed about undefined-valued own properties, and no check verified the attestation from the served bytes — the only bytes a relying party has. The published preimage rule was correct; the bytes it pointed at could not be reproduced, which for a relying party is an unverifiable attestation regardless of cause.",
      "what_changed": "The leader key is omitted instead of set to undefined. Served bytes are byte-equivalent (element-wise and full-body comparison against the live payload, modulo attestation material). An end-to-end proof through the real handler with a throwaway key shows the attestation verifying from served bytes, the living_stamp still verifying, and totals unchanged; the gspc truth tests pass 16/16. Merged as #2657. Until a deploy serves it, the live payload's site_attestation remains unverifiable and must not be read as validating the payload — the living_stamp and the 335 signed cards verify independently.",
      "reached_the_public": true,
      "note": "Window start unknown (shipped with the leader-exclusion change); observed INVALID 2026-09-20T01:52Z and still INVALID pre-deploy at 03:40Z. Full evidence bundle: evidence/reconciliation-2026-09-20/ (RECON-2026-0920-01)."
    },
    {
      "id": "C-2026-0917-01",
      "date": "2026-09-17",
      "first_observed_at": "2026-09-17T04:30Z",
      "what_was_wrong": "public/interop/ots/manifest.json claimed 566 OpenTimestamps proofs, every row asserting state PENDING_BITCOIN_CONFIRMATION. Checked by deserialising each file it named, NONE of the 566 was a proof: each was a calendar response fragment saved under the .ots extension. A row asserting that a stamp exists and awaits Bitcoin confirmation, for bytes that are not a stamp, is a false claim about evidence. The same defect recurred four times across 16 and 17 September, reaching 1,915 claimed proofs at its largest, and a later sample of 40 from one branch again contained zero real proofs.",
      "why_it_was_wrong": "The manifest was generated from a list of files something intended to stamp rather than from the bytes on disk. Nothing read the files back. An .ots extension is itself a claim that the bytes are a timestamp, and the manifest repeated that claim 566 times without checking it once.",
      "what_changed": "Files that do not deserialize are quarantined as .ots.invalid rather than deleted. The manifest is generated by scripts/ots_manifest_rebuild.py, which reads each file back and keeps only what parses, and carries a selftest proving it rejects a non-proof and accepts a real one. Branch protection on master now applies to admins, which stopped the source of the batches.",
      "reached_the_public": false,
      "note": "The deploy gate refused every build carrying these files, so no version of the inflated manifest was served. Recorded anyway: it was committed to a public repository, and a fourfold recurrence in two days is the finding rather than the exposure.",
      "evidence": [
        "/interop/ots/manifest.json",
        "scripts/ots_manifest_rebuild.py",
        "scripts/ots_guard.py"
      ]
    },
    {
      "id": "C-2026-0917-02",
      "date": "2026-09-17",
      "first_observed_at": "2026-09-17T04:50Z",
      "what_was_wrong": "Every public surface councilof.ai serves answers HTTP 403 to a plain standard-library HTTP client while answering normally to a browser. Measured 17 September 2026: 21 of 21 published URLs, including /.well-known/did.json, /.well-known/agent-card.json, /.well-known/x402.json, robots.txt and llms.txt. Those five exist only for machines. We have told correspondents, standards bodies and regulators in writing that they can fetch our evidence and verify it without our cooperation. For anyone using a standard client, that was not true.",
      "why_it_was_wrong": "A Cloudflare Browser Integrity Check at the zone level rejects requests by client signature, returning error 1010. It was enabled for the website and silently covered the API and the well-known paths beneath it. An earlier check of ours missed it because we tested user-agent strings through curl rather than the actual clients, and the strings we happened to pick were allowed. Testing a sample that excludes the reported case is not testing.",
      "what_changed": "The measurement is published at /interop/machine-reachability-2026-09-17.json and is reproducible by scripts/machine_reachability.py, whose selftest proves it can report both outcomes. The twelve artifacts a stranger most needs are mirrored to a host that does serve plain clients, at huggingface.co/datasets/csoai/councilof-ai-mirror, refreshed twice daily. The zone setting itself is a dashboard control we cannot reach with any credential we hold, and it is recorded as the owner's action.",
      "reached_the_public": true,
      "note": "This one did reach the public, and it is the most consequential defect we have published against ourselves. An organisation whose entire output is machine-readable evidence was refusing machines at every address it publishes. Credit for first reporting it goes to an external agent correspondent who tested the payable door and wrote to us.",
      "evidence": [
        "/interop/machine-reachability-2026-09-17.json",
        "scripts/machine_reachability.py",
        "https://huggingface.co/datasets/csoai/councilof-ai-mirror"
      ]
    },
    {
      "id": "C-2026-0916-03",
      "date": "2026-09-16",
      "first_observed_at": "2026-09-16T12:38Z",
      "what_was_wrong": "/interop/swift-measure.json published a per-bank page_state for seventeen banks — six OK, eight HTTP_404, three HTTP_403 — in a file whose card was titled around the banks being measured. Eleven of those seventeen URLs are newsroom index pages we chose ourselves (anz.com.au/newsroom/media-releases/2026/, citigroup.com/global/news/newsroom, wellsfargo.com/about/press and so on). A 404 on a path we invented is evidence about our guess and not about the bank, so eleven rows read as findings about institutions when they were findings about our own URL list. Separately the primary source — Swift's press release naming the cohort — had never been fetched at all: swift.com answers HTTP 403 to this client, so even cohort membership was second-hand.",
      "how_caught": "Reading the failing rows instead of the summary, and noticing that the failures clustered on hand-written newsroom index paths rather than on the banks.",
      "fix": "The file now states, at the top and on every row, that it records the HTTP state of seventeen URLs we chose and nothing about the seventeen banks, with url_provenance GUESSED_BY_US_NOT_PRIMARY_SOURCE on each row and a page_state_means sentence saying what the state does and does not support. A primary_source block records the Swift press release as UNFETCHED_HTTP_403 with the time of the attempt. status_all stays DISCOVERED. No bank in this file has been measured on tokenisation posture and none may be quoted as such. The artifact was re-stamped after the edit and the OTS manifest digest updated, because a proof stops covering a file the moment the bytes change.",
      "status": "CORRECTED IN SOURCE AND RECORDED; THE COHORT REMAINS UNMEASURED"
    },
    {
      "id": "C-2026-0916-02",
      "date": "2026-09-16",
      "first_observed_at": "2026-09-16T10:42Z",
      "what_was_wrong": "Three files published at /interop/ots/ as OpenTimestamps proofs were not OpenTimestamps proofs. swift-measure.ots, cobol-measure.ots and stablecoins-extended.ots carried no .ots magic header and python-opentimestamps rejects each with BadMagicError. The manifest beside them (csoai.ots-manifest/0.1) listed a sha256 per subject; none of the three matched the bytes of the artifact it named (swift-measure.json actual 005d04fd…, manifest d4377a67…; cobol-measure.json actual 15900afe…, manifest 0f8a354f…; stablecoins-extended.json actual 752175b2…, manifest d7a8936b…). A reader following the manifest would have been told an anchor existed for bytes that were never stamped. Separately, the three ledger cards for the same run carried subjects saying the banks, COBOL systems and stablecoins were 'measured' and payload flags.measures true, while every row in all three artifacts reads status DISCOVERED and each artifact's own honesty field says a homepage fetch is not a measurement of the subject property.",
      "how_caught": "Auditing what landed on master before deploying it: recomputing the sha256 of each named artifact and comparing it with the manifest, then deserializing each .ots with python-opentimestamps instead of trusting the file extension.",
      "fix": "Root cause was the producer: scripts/ots/ots-stamp.py wrote the calendar's raw HTTP response fragment to disk. A detached proof is magic header + version + file-hash-op + file digest + serialized timestamp. The script is rewritten to build a DetachedTimestampFile with the opentimestamps library, to submit to four calendars, and to expose --verify; it reports PENDING and never says anchored. Real proofs were created for the three artifacts (swift-measure.json.ots, cobol-measure.json.ots, stablecoins-extended.json.ots), each verified to commit to the actual file digest and each carrying four PendingAttestations and no Bitcoin attestation. The manifest is reissued as csoai.ots-manifest/0.2 with recomputed digests, the pending state stated plainly, and a supersedes block naming this record. The three invalid files are kept unedited under /interop/ots/_invalid-2026-09-16/ with a README, so anyone who read them can see what was published. The three card subjects now state what was measured (page reachability and the sha256 of the bytes returned) and that the subject property is DISCOVERED; flags.measures is false and discovery_only is true. Nothing here is anchored until scripts/ots-upgrade.py lands a BitcoinBlockHeaderAttestation.",
      "status": "CORRECTED IN SOURCE AND RECORDED; PROOFS ARE PENDING, NOT ANCHORED"
    },
    {
      "id": "C-2026-0916-01",
      "date": "2026-09-16",
      "first_observed_at": "2026-09-16T09:47Z",
      "what_was_wrong": "Two outbound emails on 16 September (to DBTA and Nieman Lab, 05:35-05:36Z) corrected an earlier figure for the supersession ledger from 953 to 841 rows and asked the recipients to use 841. At that moment the live file https://councilof.ai/interop/mill-cards-signed/SUPERSEDED.jsonl did read 841 lines. The copy on master already held 953: 112 entries dated 2026-09-15 (latest at 2026-09-15T08:02:55Z) had been committed but not deployed, because the deploy workflow had not run since 2026-09-15T08:06Z. The hand deploy at 2026-09-16T09:31Z shipped them, so the live file now reads 953 lines and 953 distinct superseded_id values. Both figures were true of the bytes they were read from; neither message said which copy it had read.",
      "how_caught": "Re-reading the live ledger at 09:47Z before quoting it in a further approach, and comparing the count with the figure sent earlier in the day and with the entries' own at timestamps.",
      "fix": "Approaches sent after 09:47Z quote 953 with the read time. The two recipients of the 841 figure are not written to again; this record is the correction, and the ledger they were pointed at now reads 953. Outreach rule recorded: a figure quoted outward names the copy it was read from (live URL and read time), and a deploy that moves a quoted figure is logged against the messages that quoted it.",
      "status": "RECORDED; THE LIVE LEDGER IS AUTHORITATIVE AT ITS READ TIME"
    },
    {
      "id": "C-2026-0915-01",
      "date": "2026-09-15",
      "first_observed_at": "2026-09-15T07:30Z",
      "what_was_wrong": "On GET /api/gspc, the governance axis's historical_measurement_record.note said the tuned governance specialist \"leads AND the lead is separated (McNemar p=0.0086 vs best base mistral:7b) — one of only 4 separated leads on the board\". The same payload read governance.separation UNTESTED, totals.separated_leads 0 and totals.own_leaders_excluded 8 with governance among them. The sentence predated the own-model exclusion, which removes our own specialist from the public leader slot and with it every public separation determination on that axis. The same class stood in four more places: the care note (\"SEPARATED vs the best base\") and the affect note (\"the cleanest separation on the board\"), both on own-model axes reading UNTESTED; a limitations line on the same payload saying \"care is separated from base models\"; and on the site, /benchmarks typed \"Governance separates at p=0.0086, care at p=0.0356, affect at p=0.0078\" and two sector pages promised a card \"with the separated lead\". A dated /feed.xml item said \"3 of 13 canonical axes carry a separated leader\" in the present tense.",
      "how_caught": "Reading the live payload against itself: the governance note's prose was compared with the governance separation field and with totals.separated_leads in the same response, instead of being accepted as a summary.",
      "fix": "The governance, care and affect notes in functions/api/_gspc_axes_a.ts and _gspc_axes_b.ts now state each separation as an in-lane result on our own model, not a public ranking, and not counted in totals.separated_leads; the typed count is gone. The care limitation in functions/api/gspc.ts is derived from the raw axis rows and carries the same label. /benchmarks and the government and affect sector entries no longer present these as public separated leads. The feed item is dated to its sitting and points at the live count. functions/api/gspc.separation-truth.test.ts reads the served board and fails when any note, historical record or limitation claims a separated lead on an axis whose separation field is not SEPARATED (a historical record passes only when marked superseded and labelled in-lane), or when any typed count of separated leads differs from totals.separated_leads; it carries failing controls for the pre-correction governance sentence. The signed snapshots that carry the old sentence (public/signed/gspc-board.signed.json, public/signed/gspc-measurement.json) are not edited; this record supersedes that sentence in them.",
      "status": "CORRECTED IN SOURCE AND RECORDED; VERIFY THE CURRENT LIVE ENDPOINT"
    },
    {
      "id": "C-2026-0914-03",
      "date": "2026-09-14",
      "first_observed_at": "2026-09-14T10:54Z",
      "what_was_wrong": "The wrapper-parity roster pointed usdt0:arbitrum and usdt0:optimism at 0x2E1dBfbf44d8855fDE5D5fD6c978a9b10bc27627 and usdt0:ethereum at 0x48C04ed50508680b93561a5800E97e24C05e639F. None of these is a USDT0 token contract. Three public.notice cards signed under did:web:csoai.org#board-attestation-1 and included in the public root record those reads: public/cards/385d7cd72fee80b4.json (usdt0:arbitrum), public/cards/a911bc077ebb9b1f.json (usdt0:optimism) and public/cards/293cd51070159fa3.json (usdt0:ethereum). Each honestly states UNMEASURED with an empty eth_call result from that address, but the subject label USDT0 was wrong for the address read. No parity number was published for USDT0.",
      "how_caught": "Re-verifying roster addresses against the issuer's published deployments page (docs.usdt0.to) while adding sourced wrapper rows; the two affected addresses returned no token on either chain.",
      "fix": "The roster now uses the addresses on docs.usdt0.to (Arbitrum 0xFd086bC7CD5C481DCC9C85ebE478A1C0b69FCbb9, Optimism 0x01bFF41798a0BcF287b996046Ca68b395DbC1071). The staged atoms were re-staged from the corrected roster and read UNCHECKABLE_NATIVE_ISSUANCE. The usdt0:ethereum atom had no roster row and was removed from staging. The three signed cards are not edited; this record supersedes their subject label.",
      "status": "CORRECTED IN SOURCE AND RECORDED; VERIFY THE NEXT PUBLIC ROOT"
    },
    {
      "id": "C-2026-0914-02",
      "date": "2026-09-14",
      "first_observed_at": "2026-09-14T10:37Z",
      "what_was_wrong": "On GET /api/gspc, the reserve-attestation axis note said '1 PASS, 6 FAIL, 9 UNCHECKABLE'. The evidence file it cites, /interop/financial-measure-run-reserve-attestation.json (as_of 2026-09-07T11:30:35Z), tallies 3 PASS, 4 FAIL, 9 UNCHECKABLE; 1/6/9 is the custody-disclosure tally. The regulatory-framework note said '3 PASS, 4 FAIL, 9 UNCHECKABLE' while its evidence file tallies 4 PASS, 3 FAIL, 9 UNCHECKABLE. When the wrong strings were introduced is UNCHECKABLE from the shallow repository history available; they predate 2026-09-14T01:46Z.",
      "how_caught": "Verifying a regulator comment draft before submission: the draft quoted the board note, and the reviewer compared it with the tally field of the cited evidence file instead of accepting the summary.",
      "fix": "Both notes in functions/api/_gspc_axes_fin.ts now quote their evidence files. functions/api/gspc.financial-tally.test.ts reads every typed PASS/FAIL/UNCHECKABLE triple in a financial note and requires it to equal the cited evidence file's tally, with failing controls for the two pre-correction strings. The evidence files were correct and are unchanged. Signed board snapshots that carry the old notes are superseded by this record, not edited.",
      "status": "CORRECTED IN SOURCE AND RECORDED; VERIFY THE CURRENT LIVE ENDPOINT"
    },
    {
      "id": "C-2026-0914-01",
      "date": "2026-09-14",
      "error_introduced_at": "2026-09-14T08:29Z",
      "what_was_wrong": "44 signed third-party Hub measurement cards on the swarm axis (#2321, merged 2026-09-14T08:29Z, and #2330, 09:41Z; runs gha-34818995409 and gha-34824895140) were graded by a prompt that could not be answered wrongly. The frozen bank csoai/gspc-swarm (revision e8a4ec1e, sha256 ab318986…) is keyword-graded: all 37 rows expect KEYWORD_MATCH and carry must_inc keywords. The Hub mill builds its exact-label answer menu from the bank's expected column, so every item prompt read \"Reply with EXACTLY ONE token from: KEYWORD_MATCH\" and every model that followed the format scored 1.0. 26 of the cards read MEASURED at n=30 with accuracy 1, and /api/hub-cards counted all 26 as MEASURED cells. Every check the cards passed was real — signature, content id, admission receipt, per-item evidence that recomputes byte for byte — and none asked whether the menu offered a wrong answer. Pod (Ollama) swarm cards bind the same bank bytes under a different instrument and read 0.027–0.081, so the two populations were never comparable.",
      "how_caught": "Reading the published item evidence behind one card (signed-swarm-0711716149e0.json, deepseek-ai/DeepSeek-V4-Pro, items-swarm-e91bd4f805de.jsonl): all 30 rows carried the same one-option menu and expected its only option. Contrast under the same instrument hash (86216fbb…): the care bank offers 0 | 1 with mixed expected labels, and Qwen3-14B read 4/30. Every MEASURED swarm row in /interop/hub-cards-index.json read accuracy 1.",
      "fix": "Signed bytes are not edited. The 44 cards are recorded in /interop/mill-cards-signed/WITHDRAWN.jsonl — withdrawn, not superseded, because no sound card exists to replace them — derived from the bound bank bytes by scripts/withdraw_one_option_cards.py, which fails the PR gate if any signed card on a one-option bank is not withdrawn. /api/hub-cards drops withdrawn cards from cells and counts, lists them under withdrawn_cells with this id and their status as published, and withholds totals if the withdrawal ledger is unreadable. The deployed hub-cards index excludes them, and the census flip marks them WITHDRAWN, so its next run retires their queue cells and leaves them out of the rebuilt Hub indexes. The instrument now fails closed: an exact-label menu with fewer than two distinct labels, or one naming a grading mode, is refused; the mill skips such a bank as UNCHECKABLE before spending an item prompt, and the offline admission verifier refuses such a card. Two-label prompts are byte-identical, so admitted care, safety and governance cards still verify. The Hub swarm cell is UNMEASURED until a keyword grader that passes harness/gspc-top100/check_bank_discriminates.py is wired into the Hub mill.",
      "status": "WITHDRAWN IN SOURCE; VERIFY WITH python3 scripts/withdraw_one_option_cards.py"
    },
    {
      "id": "C-2026-0913-01",
      "date": "2026-09-13",
      "first_observed_at": "2026-09-12T16:35Z",
      "error_introduced_at": "2026-09-12T15:38Z",
      "corrected_at": "2026-09-13T02:29Z",
      "what_was_wrong": "scripts/erc8004_census.py (merged in #2020, 2026-09-12T15:38Z) reported 8 ERC-8004 Identity Registry registrations on Ethereum. The true count at the same floor-to-head range is 50,783. Cause: rpc.flashbots.net serves a silently incomplete historical log index — it returned well-formed empty results for ranges containing receipt-verified events, with no error. The tool's then-current checks (chain id, finality pin, getCode, log shape) all pass against such an endpoint; empty result is not absence.",
      "how_caught": "Claim/evidence disagreement during the TUI-4 root-and-registry watch: Etherscan showed 19,138 transactions to the registry while the tool reported 8 events. The on-chain receipt of one Etherscan-visible Register transaction (0x335580236be7…f1f88b, block 25,883,771) proved a Registered log existed that flashbots' getLogs never returned. The tenderly public gateway returned the event for the same query, isolating the provider.",
      "fix": "#2070 (merged 2026-09-13T02:29Z): known-event integrity anchors — one receipt-verified (block, tx) pair per chain; an endpoint's scan is trusted only if it returns the anchor event when the anchor block is in range (negative test: flashbots is refused on the ETH full-history range). Corrected counts: Ethereum 50,783; Base 86,263 (reproduced by two independent providers); BSC full history UNCHECKABLE permissionlessly. The wrong '8' is superseded here, not rewritten away.",
      "status": "CORRECTED IN SOURCE AND RECORDED; VERIFY WITH scripts/erc8004_census.py"
    },
    {
      "id": "C-2026-0912-01",
      "date": "2026-09-12",
      "what_was_wrong": "The live SwarmBench v2b row claimed a statistically separated qwen2.5:7b leader by comparing a stated 0.384 lower bound with a 0.372 upper bound for mistral:7b. The signed candidate cards show qwen3:4b at 0.4070 and qwen2.5:1.5b at 0.4000, both ahead of mistral:7b at 0.1481, so mistral was not the runner-up. The same sentence then said the top three remained statistically tied, contradicting its own separated-leader label. A standing limitation also described the active row as the retired 3-prompt PROTOCOL bank rather than the 37-item wave-2b bank.",
      "how_caught": "Public-site inspection before a proposed orchestration-system design-partner approach. The audit followed the swarm row into /signed/card_index.json and compared all seven signed swarm-candidate cards instead of accepting the board summary.",
      "fix": "The live serving layer now keeps the signed qwen2.5:7b point estimate and leader identity but marks statistical separation UNTESTED. Its basis states the actual point ordering and the missing evidence: no published paired item rows or compatible confidence intervals support a separation determination. The limitation now distinguishes the active 37-item wave-2b bank from the retired PROTOCOL result. Historical signed bytes remain unchanged and are explicitly superseded by this correction rather than silently rewritten.",
      "status": "CORRECTED IN SOURCE AND RECORDED; VERIFY THE CURRENT LIVE ENDPOINT"
    },
    {
      "id": "C-2026-0905-02",
      "date": "2026-09-05",
      "what_was_wrong": "26 SWIFT rail cards were published under public/interop/swift-signed-2026-09/ as signed-swift-<bank>.json with a populated sig_ed25519 field and signed_at timestamp. The field held base64(sha256(card)), not a signature; sig_algo said SHA256-placeholder and the index said the same. A relying party reading the field name, the file name or the directory name was told these were Ed25519-signed. They were not. Nothing verifies.",
      "how_caught": "Outside review of the estate on 2026-09-05 named the 26 placeholder cards as the single most damaging thing an inspector could find. Confirmed against master: 26 of 26 files, sig_algo SHA256-placeholder, producer scripts/badger/csoai-swift-aware.py writing a digest when no key was present.",
      "fix": "Producer changed: with no key it now writes sig_ed25519 null, sig_algo UNSIGNED, signed_at null, a signature_note, into swift-staged-2026-09/ as staged-swift-*.json; the OIDC board-sign path is the only signer. The 26 artifacts were rewritten the same way and moved; swift-signed-index.json is superseded by swift-staged-index.json (total_signed 0, total_staged_unsigned 26). No card here is signed or MEASURED.",
      "status": "CORRECTED — 0 signed, 26 staged and labelled; placeholder producer removed"
    },
    {
      "id": "C-2026-0905-03",
      "date": "2026-09-05",
      "what_was_wrong": "Three public endpoints turned a source they could not read into a number, and two of them published a figure that was wrong while they did it. (1) /api/hub-cards fans out to four Hub index files and totalled whatever came back. Two of the four were answering nothing to the Worker, and both held ONLY UNMEASURED rows, so the endpoint served 682 cells / 647 MEASURED / 35 UNMEASURED when the published population was 717 / 647 / 70. It understated the unmeasured count by exactly half, and the error therefore ran in the flattering direction — the one direction a measurement body may never round. The endpoint did disclose the partial read, but it did so in an honesty field while counts kept publishing quotable integers beside it; a disclosure next to a wrong number does not repair the number, and downstream quotes the number. (2) /api/dashboard/stats derived fleet.online from `.online ?? .nodes?.length ?? 0`. /api/oracle-fleet emits neither field — it answers 200 with a single host's health — so the dashboard published online: 0, meaning no nodes online, against a fleet that was up and answering with 26.9 days of uptime. That is a claim the fleet endpoint never made, invented from two absent keys. The same file coalesced every other aggregate with `?? 0`, so an unreadable /api/gspc would have published measured_axes: 0 while the board carries 22, under a header that claimed honest empty states — but zero is a measurement, not an empty state. (3) /api/hf-spaces returned an empty list on any non-OK response and counted the survivors, so one upstream throttle would publish models: 0, indistinguishable from the org having no models. That one was latent: it agreed with the Hub on the day it was found.",
      "how_caught": "A top-down alignment pass on 2026-09-05 re-ran the estate brief's own verification commands instead of trusting the brief, and /api/hub-cards disagreed with it. Reading all four Hub index files directly showed all four answering 200 and non-empty to a plain client at dataset commit c52587b, while the endpoint's own indexes_read field said 2 of 4. A sweep for the same shape — any endpoint that fans out to N sources and reports whatever came back — found the other two. The dashboard defect had a passing test over it: the fixture mocked /api/oracle-fleet as an object carrying online and nodes, a shape the real endpoint does not return, so a test that invented the upstream could not catch a misread of the real one.",
      "fix": "Every one of the three now distinguishes an unread source from an empty one. A total is published only when all of its sources answered; otherwise the totals are null, what was actually read is offered under a separate name documented as a floor, and each missing source is named with its reason. An index or listing that answers with zero rows counts as READ — the previous code treated any empty result as unreachable, which would have let a legitimately empty source suppress the totals forever. hub-cards additionally retries a failed index once outside the Cloudflare cache, because the fetch carried cacheEverything and a cached non-OK response keeps a source dark for the whole ten-minute window. dashboard/stats gained a sources block naming each upstream's state and a note stating that a null is an unread value and never a measured zero; its fleet.online carries its own note explaining why it is null, so a bare null cannot be re-read as zero. The dashboard UI already rendered a missing value as an em dash, so the honest empty state was available all along and was simply not being sent. Tests were watched failing against the unpatched handlers before being accepted.",
      "status": "CORRECTED IN SOURCE — hub-cards under PR #1294, dashboard/stats and hf-spaces under PR #1297, recorded as issue #1295. The wrong figures were live until those deploy. Whether the hub-cards retry restores the two dark indexes is not yet established: it cannot be tested from outside the Worker, and if they stay dark the endpoint now reports that instead of a flattering subtotal."
    },
    {
      "id": "C-2026-0905-04",
      "date": "2026-09-05",
      "what_was_wrong": "Six public manifests under /interop advertised 36 endpoint references that do not exist: custom-gpt-bridge.json told Custom GPTs to POST /api/measure, /api/verify and /api/xrpl/evidence; chatgpt-features-finish.json listed 14 'features' (/api/voice, /api/vision, /api/calendar, /api/email, ...) each with an endpoint; deep-research-integration.json described a four-endpoint /api/research pipeline; persona-tests.json, chatgpt-skills.json and anchor.json cited /api/anchor, /api/insurance/attest, /api/xrpl/rlusd, /api/xrpl/usdc and /api/scheduler. Every one answered HTTP 404 to GET and POST on 2026-09-05. All six were written by two generators under scripts/badger/ that assemble manifests from a wish-list and never probe a route.",
      "how_caught": "A top-down pass on 2026-09-05 found /api/verify returning 404 and followed the references: three files first, then every /api/ path in the six generated manifests, each probed live with GET and POST.",
      "fix": "Each artifact now carries claims_audit_2026-09-05 naming the dead paths; every dead reference is marked NOT_IMPLEMENTED in place, and the three Custom GPT actions a client would actually call were removed and listed under actions_removed. Both generators now exit at main() with the reason and cannot regenerate the fiction. The rule (an endpoint advertised outward must answer non-404 live) is the one scripts/outward-claims-guard.mjs enforces post-deploy.",
      "status": "CORRECTED IN ARTIFACT AND PRODUCER. Whether any Custom GPT or agent acted on the dead manifests is unknown; no request log is kept for those paths. Nothing was ever measured, signed or anchored through them."
    },
    {
      "id": "C-2026-0905-05",
      "date": "2026-09-05",
      "what_was_wrong": "A merged commit and its PR (#1321) stated that a confirmed x402 settlement never reached the revenue ledger: \"a real payment settled and the ledger never saw it\". That is false. The settlement WAS recorded. The reading behind the claim was taken 6 seconds after the settle, and Cloudflare KV list operations are eventually consistent — the record had not propagated yet. Re-read ~20 minutes later, /api/revenue one_number showed settlements 1, all_time 1, records_unreadable 0. No payment was ever lost.",
      "how_caught": "Re-checking the same endpoint later in the same session instead of trusting the first reading. curl -s https://councilof.ai/api/revenue | python3 -c \"import sys,json;print(json.load(sys.stdin)['one_number'])\" — run twice, minutes apart, and the two disagree while nothing else changed.",
      "fix": "This entry records the false claim; the commit message cannot be rewritten. The code change that shipped with it stands on its own merits and is unaffected: recordSettlement had swallowed every KV error into an empty catch, so a failed write and no settlement really were indistinguishable, and it now returns {stored,reason}. What was wrong was the diagnosis, not the fix. A second defect found while re-reading IS real and is corrected in the same change: one zero-value settle from an ephemeral wallet moved one_number.all_time from 0 to 1, counting a wallet we created and controlled, paying nothing, as a distinct non-self buyer — so settlement records now carry zero_value, because the payer-exclusion list can never enumerate a throwaway key.",
      "status": "RECORDED — the claim was a measurement error (KV eventual consistency read at 6s); no settlement was lost"
    },
    {
      "id": "C-2026-0906-01",
      "date": "2026-09-06",
      "what_was_wrong": "CSOAI-ORG/proofof-ai-mcp shipped detect_deepfake_image with a substring-blacklist path check ('/etc/', '/var/', '..'). A blacklist is not a boundary: any path outside the list, and any symlink into a listed directory, was readable — a Local File Inclusion. A security researcher reported it on 2026-06-12 (issue #8) and the report sat unanswered for 86 days.",
      "how_caught": "The 2026-09-06 HF + GitHub audit listed every open issue across the org older than 7 days; the only security report was this one, with zero comments.",
      "fix": "PR #20 on that repository: an allowlist under PROOFOF_ALLOWED_DIR (default ./uploads), realpath-resolved, regular files only, symlink escapes rejected; verified against /etc/hosts, ../ traversal, an escaping symlink and ~/.ssh/id_rsa. The reporter was answered on the issue.",
      "status": "CORRECTED IN SOURCE. Whether any deployment of that server was exploited is unknown; it keeps no access log. The 86-day silence is the defect this entry records: security reports across the org are now part of the outward-claims guard's issue sweep."
    },
    {
      "id": "C-2026-0905-01",
      "date": "2026-09-05",
      "what_was_wrong": "The ONE root (public/root.json) is documented as republished hourly. Between 2026-09-02T04:14Z (last successful public-root run) and 2026-09-03T06:20Z (first successful run after GitHub reinstated Actions on the CSOAI-ORG account) it was not republished at all: the hourly runs from 05:14Z to 19:58Z on 2 Sep never started (Actions disabled for the account, Support ticket #4720908), and the eight runs from 2026-09-02T20:58Z to 2026-09-03T06:16Z failed at runner start. Cards signed in that window were not in any root a reader could fetch, and the witness pointer kept reporting the 04:14Z root as current, which it was — but nothing said the cadence had stopped.",
      "how_caught": "Run history of .github/workflows/public-root.yml read back on 2026-09-05 after reinstatement: one success at 04:14Z, a gap with no runs at all, eight failures, then success at 06:20Z on 3 Sep. The gap is visible only in the run list; the root, the pointer and the site all looked normal during it.",
      "fix": "This entry records the window. No root bytes were edited (none existed to edit). The as_of field on the root and the checked_at field on the pointer are the only honest freshness signals; HOW-TO-VERIFY-ROOT.md already tells a reader to re-fetch and compare rather than trust a MATCH observation. Structural fix, same day: the witness now also reports a CONFLICT state when two witnessed roots carry the same as_of and different merkle_root values, so a stalled or forked cadence is named rather than inferred.",
      "status": "RECORDED — a 26-hour publication gap, 2026-09-02T04:14Z to 2026-09-03T06:20Z; no bytes changed, cadence documented as not guaranteed"
    },
    {
      "id": "C-2026-0903-01",
      "date": "2026-09-03",
      "what_was_wrong": "The Layer-0 ceremony artifact (/interop/layer0-ceremony-2026-09-03.json, v0.2) listed /api/intoto as one of 15 machine rails, recorded it as returning 404, and explained the 404 as 'the handler exists in master but is inside an undeployed window'. There is no handler. functions/api/intoto.ts exports only helpers (subjectDigest, toInTotoStatement, toDsse) and is imported by functions/api/detect.ts and functions/api/detector-interop.ts, both of which serve 200. No deploy would ever have turned it into a route. A ceremony whose purpose is to attest our own machine surface had invented a door and then explained away its absence.",
      "how_caught": "Live sweep of 25 published surfaces on 2026-09-03: exactly one non-200, /api/intoto. Tracing it showed the file has no onRequest export, and that the ONLY thing on the estate advertising /api/intoto as an endpoint was the ceremony artifact itself.",
      "fix": "Ceremony superseded at v0.3: the rail is removed and the correction is stated in the artifact's own what_this_does_not_claim, first line. The count becomes 14 of 14 serving rather than 14 of 15. in-toto capability is real and reachable through /api/detect and /api/detector-interop. v0.2 was superseded in place rather than kept, because it had no external reference and its OpenTimestamps stamp was still PENDING with no Bitcoin attestation to preserve; had the stamp been upgraded, the bytes would have been kept and a new file issued.",
      "status": "CORRECTED — 14 of 14 rails; the invented door is gone"
    },
    {
      "id": "C-2026-0902-09",
      "date": "2026-09-02",
      "what_was_wrong": "After C-2026-0902-08, live GET /api/gspc and /api/state headlines were 22 axis · 22 measured, but public/signed/gspc-board.signed.json was still the earlier 22/15/7 freeze, so signed_snapshot_agrees stayed false and the snapshot was labelled do-not-file.",
      "how_caught": "Owner MPC ceremony on the Oracle custody host: live /api/gspc snapshot (site_attestation stripped) signed with did:web:csoai.org#gspc-board-22axis-2026 (3-party Coinbase cb-mpc Ed25519 additive). Offline verify (scripts/gspc-board-verify.mjs) returned VERIFIED; content_id 72ba8a3371fcc895be835f4283fefca0c2edd1e1fc857b3e49276277f94ccb10.",
      "fix": "The verified 22/22 freeze replaced public/signed/gspc-board.signed.json. /api/state now reports signed_snapshot_agrees from the count match (22 slots · 22 measured). The 15/7 file is superseded, not edited. The Pages /api/board-sign path was not used — it is a 3KB card-sign and cannot carry this snapshot.",
      "status": "CORRECTED — signed freeze is 22/22 and agrees with the live axis arrays"
    },
    {
      "id": "C-2026-0902-08",
      "date": "2026-09-02",
      "what_was_wrong": "/api/state quoted public/signed/gspc-board.signed.json totals (22 slots · 15 measured · 7 empty) as the number to file, and said that when that snapshot disagreed with live /api/gspc neither figure was quotable. Live GET /api/gspc (and the committed axis arrays it derives from) is 22 axis · 22 measured · 0 empty. A VRO map mailed 1 Sep used the 15/7 freeze; the correction that actually transited SMTP is Sent 82 (2 Sep 14:50Z) pointing at /api/gspc.",
      "how_caught": "Recipient audit of the VRO table: /api/gspc and the homepage said 22/22; /signed/gspc-board.signed.json and /api/state still said 15/7 with signed_snapshot_agrees false.",
      "fix": "/api/state board headlines now derive from the same axis arrays as GET /api/gspc. The signed snapshot stays on disk as a historical freeze (MPC key did:web:csoai.org#gspc-board-22axis-2026, three shares, not re-derived here) and is labelled do-not-file. Re-signing that 38KB file is an owner MPC ceremony — the Pages /api/board-sign path is a 3KB card-sign and cannot carry the snapshot.",
      "status": "CORRECTED — live 22/22 is the quotable count; snapshot 15/7 is historical pending owner MPC re-sign"
    },
    {
      "id": "C-2026-0902-10",
      "date": "2026-09-02",
      "what_was_wrong": "The published verification rule (/signed/HOW-TO-VERIFY.md and HOW-TO-VERIFY-ROOT.md) did not state that a verifying signature says nothing about whether the signing key is still valid. A reader verifying with yesterday's trust anchor would get the same VALID verdict after a revocation this morning, and nothing in the text said so.",
      "how_caught": "IETF agentproto list, 31 Aug–2 Sep 2026: an objection to the offline-verification sentence in a proposed charter amendment (offline verification is a computation over the past; revocation is a fact about the present). CSOAI committed on the list to add the sentence and note the correction.",
      "fix": "Both rules now carry a section stating what a verifying signature does not establish, that no revocation mechanism or key-freshness requirement is defined here, that a consumer must not treat a verifying signature as evidence the key is still valid, and that key-resolution path and accepted staleness are deployment parameters the card does not carry.",
      "status": "CORRECTED — rule amended; second unstated property caught by that thread (the first was a signed flag with no signature bytes behind it)"
    },
    {
      "id": "C-2026-0902-07",
      "date": "2026-09-02",
      "what_was_wrong": "On 2026-08-28 a commit edited the text of a signed card in place (public/signals/cross-border-card.signed.json, field measured_axes: the 18 Aug count was replaced with a pointer to the live count) without re-signing. The content_id no longer derived and the Ed25519 signature no longer verified — a silent edit of a signed artefact, which this ledger's own policy forbids.",
      "how_caught": "The unit suite (cardVerify: content_id derives and signature verifies) failed on master; found during the 2 Sep test-truth pass.",
      "fix": "The original signed bytes are restored so the card verifies again. The caveat lives here instead: the card's measured_axes text quotes the 18 Aug 2026 count; the live count is only ever GET https://councilof.ai/api/gspc totals. Signed bytes are never edited — they are superseded by a new signed card or annotated in this ledger.",
      "status": "CORRECTED — signed bytes restored; caveat carried by this entry"
    },
    {
      "id": "C-2026-0902-01",
      "date": "2026-09-02",
      "what_was_wrong": "The Switchboard research brief recorded OUSG's XRPL domain check as unverified (directory only) while GET /api/xrpl showed it bidirectional with a signature.",
      "how_caught": "Owner reconciliation of the 2 Sep research briefs against the live API.",
      "fix": "GET wins. /api/xrpl is the authority: OUSG verified_via 'Bidirectional domain match'. The brief's cell is superseded; no data change.",
      "status": "RECONCILED — live API authoritative"
    },
    {
      "id": "C-2026-0902-02",
      "date": "2026-09-02",
      "what_was_wrong": "A secondary planning state file attributed USDB to Bitstamp. USDB is issued by Braza Bank (issuer address rB3y9EPnq1ZrZP3aXgfyfdXQThzdXMrLMc).",
      "how_caught": "Owner reconciliation; the Switchboard brief confirms Braza.",
      "fix": "GET /api/xrpl already carries issuer 'Braza Bank'; the mis-attribution lived only in a planning file and is corrected there.",
      "status": "CORRECTED"
    },
    {
      "id": "C-2026-0902-03",
      "date": "2026-09-02",
      "what_was_wrong": "The OpenAI incident post-mortem was cited as 37 pages by one source and 38 by an internal state file.",
      "how_caught": "Owner reconciliation of the incident-card inputs.",
      "fix": "No incident card hashes that artefact until the primary PDF is fetched, hashed and its page count read from the file itself.",
      "status": "PENDING VERIFICATION — card withheld until the primary PDF is hashed"
    },
    {
      "id": "C-2026-0902-04",
      "date": "2026-09-02",
      "what_was_wrong": "GPAI Code of Practice signatory counts differed: 26 per the Commission's 1 Aug 2025 list versus '28 frozen' in secondary sources.",
      "how_caught": "Owner reconciliation.",
      "fix": "Only the live, dated Commission page is carded (interop/gpai-signatory-2026-09). Secondary counts are not quoted.",
      "status": "CORRECTED — primary source only, dated"
    },
    {
      "id": "C-2026-0902-05",
      "date": "2026-09-02",
      "what_was_wrong": "One playbook stated 2 Feb 2027 as the Article 50 detector-interoperability date as fact; a market-map brief records it as unsettled.",
      "how_caught": "Owner reconciliation.",
      "fix": "The date is not published anywhere until verified against Regulation (EU) 2026/1744 in the Official Journal.",
      "status": "UNVERIFIED — withheld"
    },
    {
      "id": "C-2026-0902-06",
      "date": "2026-09-02",
      "what_was_wrong": "councilof.ai states a £5M professional-indemnity policy while the Series A pack's infrastructure-gaps sheet says insurance is unknown. One of them is wrong in a data room.",
      "how_caught": "Owner reconciliation.",
      "fix": "Owner to confirm the policy document; the losing statement is corrected in place and this entry updated. 2026-09-05: no policy document, certificate or insurer correspondence was found in the business mailbox or the repository, so the public assertion (About: 'operates with full professional indemnity insurance'; Disclaimers: 'maintains professional indemnity insurance') was withdrawn to the evidenced state — both pages now say cover is not stated until the policy document is on file. The assertion is restored, with insurer, limit and dates, the day the document is filed.",
      "status": "CORRECTED — public assertion withdrawn pending the policy document; restore on receipt"
    },
    {
      "id": "C-2026-0826-12",
      "date": "2026-08-26",
      "what_was_wrong": "The board attestation's sig_input was ambiguous, and the ambiguity was live rather than theoretical. It read \"canonical JSON (recursively sorted keys, no whitespace) of this payload with the site_attestation field removed\" — six words that do not pin a preimage. The natural first reading in a Python-flavoured estate is json.dumps(sort_keys=True, separators=(',',':')), whose default is ensure_ascii=True, and that FAILS: the signer emits non-ASCII literally, i.e. ensure_ascii=False. The signed payload carries 81 non-ASCII code points (middle dot, multiplication sign, en dash, em dash, right arrow, greater-than-or-equal), and the two readings differ by about 256 bytes. Two implementers reading the same sentence get two different preimages and one of them reports a bad signature on a good artefact. The sentence also never said whether the signature is over the raw bytes or over a digest of them.",
      "how_caught": "Outside audit of the live site, 2026-08-26 (finding A2). The auditor's first and more natural reading failed; the signature verified on the second attempt, after guessing.",
      "fix": "sig_input now states the rule as bytes: Ed25519 over the RAW UTF-8 bytes (not a digest) of canonical JSON with keys sorted by code point recursively, no whitespace, non-ASCII emitted literally as UTF-8 and never as \\\\uXXXX escapes (ensure_ascii=False, with ensure_ascii=True named explicitly as the wrong reading), and numbers serialised by ECMAScript Number::toString so an integral float renders 0 and not 0.0. Two machine-readable fields, sig_input_ensure_ascii: false and sig_input_is_digest: false, carry the same facts for a parser. CRITICALLY, THE CARDS ARE THE OPPOSITE AND STAY THAT WAY: the 150 measurement cards were minted with ensure_ascii=TRUE and CPython float repr, and each card states so in its own preimage field. Neither rule can be migrated to the other without invalidating signatures over bytes that already exist, so nothing was harmonised — both rules are now stated explicitly wherever each is published, and /signed/HOW-TO-VERIFY.md carries a table putting them side by side so a reader who verifies both is not burnt by the difference.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-11",
      "date": "2026-08-26",
      "what_was_wrong": "The public MCP `measure` tool returned ok:true for every subject, including subjects that do not exist. Passing a nonsense model name produced {\"ok\":true,\"claim\":\"measurement\",\"subject\":\"<the nonsense name>\"} with a note explaining that nothing had actually been measured. No measurement ran, no axes came back, no credential was issued, and the tool's own description promised \"a signed measurement credential\". A measurement tool that succeeds on a nonexistent subject cannot distinguish MEASURED from DID NOTHING — which is exactly what our own /api/mcp honesty_contract forbids: unknown is null or unmeasured, never a plausible-looking value. We applied that doctrine to the registry and not to the tool.",
      "how_caught": "Outside audit of the live site, 2026-08-26 (finding P1). The auditor called the tool with THIS-MODEL-DOES-NOT-EXIST-xyz and got the same ok:true as for gpt-4o. Nothing on our side was checking; the tool was listed as `probed` because tools/list returned its name, and `probed` was reading as `works`.",
      "fix": "`measure` now returns ok:false with a named state on every call, because no call to it ever succeeds: INVALID_ARGUMENT when no subject is given, NOT_MEASURED otherwise, each with the reason and a pointer to where published measurements actually live (/signed/card_index.json and /api/gspc). It also states plainly that it did NOT check whether the subject exists rather than implying it did. The tool description in tools/list is rewritten to what the endpoint does — return the contract — so an honest result no longer sits behind a description that over-promises. PARTIAL, AND SAID SO: the upstream worker's source is not in this repository, so the correction is applied at councilof.ai/mcp, the address published in .well-known/mcp.json and agent-card.json. The worker's own workers.dev origin still returns ok:true and needs an owner-side deploy to close.",
      "status": "FIXED AT THE PUBLISHED ENDPOINT; UPSTREAM WORKER FIX PENDING (owner)"
    },
    {
      "id": "C-2026-0826-10",
      "date": "2026-08-26",
      "what_was_wrong": "The jail axis published a dataset_url that is not a URL, directly beneath a note asserting that every such URL is fetchable. The axis's `dataset` field — an identifier field, resolved to a link by string concatenation against https://huggingface.co/datasets/ — held a prose sentence: \"published: csoai/gspc-jail-goldbank (frozen 71-cell gold bank, HF 2026-08-25)\". The resulting dataset_url contained a colon, spaces and parentheses and was rejected by curl as malformed. Twelve other banked axes resolved fine, and the bank itself was always fine and always public. The bank_note above it read \"Every axis WITH a frozen bank carries dataset_url — the bank resolved to a fetchable URL\": a blanket assertion with nothing deriving it, false for as long as it stood.",
      "how_caught": "Outside audit of the live site, 2026-08-26 (finding D10). The auditor fetched all fourteen; thirteen returned HTTP 200 and one would not parse.",
      "fix": "`dataset` now holds the bare slug csoai/gspc-jail-goldbank and the prose moved to dataset_note. The resolver no longer concatenates blind: a value that is not a bare <owner>/<name> slug now publishes dataset_url: null with dataset_url_state UNRESOLVABLE and the raw value, so the fault is visible on the surface that carries it instead of shipping a string that looks like a link. bank_note is now derived from that same predicate and reports counted totals (banked_axes, banked_axes_resolvable, banked_axes_unresolvable), so the sentence and the bytes cannot disagree again. The same correction is applied to the packaged /signed/gspc-measurement.json, which carried the identical prose.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-09",
      "date": "2026-08-26",
      "what_was_wrong": "We published recall: null for council-inhouse-ft on the jail axis where the measured value is 0.0. That model has tp=0 and fn=38, so recall = tp/(tp+fn) = 0/38 = 0.0 — defined, measured, and the single most damaging number on the axis: our own fine-tune detected zero of 38 escapes. null reads as NOT MEASURED. Publishing it in place of a real zero is this estate's own defect class inverted: instead of inventing a number where none exists, we erased a number that did. It sat on a row whose note says \"published, not hidden\". precision on the same row is legitimately null (0/0 is undefined, nothing was predicted positive), so two fields carrying the identical value meant opposite things with nothing distinguishing them.",
      "how_caught": "Outside audit of the live site, 2026-08-26 (finding D11). Every other cell of the jail axis reproduced to the item — seven confusion matrices, precision, recall, accuracy and the fleet mean — and this was the one arithmetic exception the auditor found.",
      "fix": "recall is 0.0 on /api/gspc and in /signed/gspc-measurement.json. The axis now carries a null_grammar field stating which null means UNDEFINED and which zero means MEASURED, so the distinction is published rather than left to be inferred. The frozen /signed/gspc-board.signed.json still contains recall: null and is NOT edited: its MPC custody signature is over those exact bytes, so correcting it at source is an owner-supervised re-sign. Until then this ledger and the live board carry the correction where a reader will meet it.",
      "status": "FIXED ON THE LIVE BOARD; FROZEN SIGNED SNAPSHOT AWAITS RE-SIGN (owner)"
    },
    {
      "id": "C-2026-0826-08b",
      "date": "2026-08-26",
      "what_was_wrong": "The living_stamp was presented as a valid attestation and cannot be checked by anyone. It shipped signed: true and a sig_input recipe, rendering exactly like the two attestations on this site that do verify. It does not verify. Three faults compound: TWO different signatures are published for one stamp, with the same signer and the same `updated` — one in /signed/board_living.json, a different one in /api/gspc measured_on.living_stamp, and at most one can be over the bytes the other is over; the signer is in NONE of the four verification methods in our own did.json, so even a reproducing preimage would prove only self-consistency, the unfalsifiable shape our own HOW-TO-VERIFY tells strangers to refuse; and board_living.json states in its own note that its axes were re-snapshotted from the live board at package time, six days after the signature date, so the signed bytes are not the published bytes.",
      "how_caught": "Outside audit of the live site, 2026-08-26 (finding A3): roughly fifty readings attempted, none verified. Re-run in this lane the same day at wider scope — both published signatures, all five published keys, nine candidate payloads, raw/sha256-digest/sha256-hex message forms, both ensure_ascii settings, every drop-set of up to three fields: 58,184 attempts, 0 verified.",
      "fix": "The stamp is marked UNVERIFIABLE wherever it is published — /api/gspc, /signed/board_living.json and /signed/gspc-measurement.json — carrying verification_state UNVERIFIABLE, verifiable: false, signer_anchored: false, the attempt count, and a note stating that it must not be treated as a valid attestation and pointing at the two attestations that do verify. It is NOT withdrawn and its bytes are NOT altered: a row saying \"we published this and nobody can check it\" is worth more than a quietly deleted one, and if a preimage rule is ever recovered it must still verify against these bytes. We do not claim the stamp is invalid — only that it is uncheckable, which for a relying party is the same outcome. To close: anchor the signer in did.json, publish the exact preimage (which fields are signature fields, raw bytes versus digest, encoding), and publish ONE signature. Owner-gated; this lane does not hold the key.",
      "status": "MARKED UNVERIFIABLE; REPRODUCIBLE SIGNATURE PENDING (owner)"
    },
    {
      "id": "C-2026-0826-07b",
      "date": "2026-08-26",
      "what_was_wrong": "The claims register described bytes that do not exist. CR-002 gave as its evidence \"Cards declare timestamp_authority: 'none'\". Zero of the 150 published cards contain that field; the string \"timestamp\" appears in no card, not in card_index.json and not in the cross-border card. The substance was honest — there genuinely is no timestamp authority behind any card — but the register asserted a positive declaration as its evidence for an absence, and the claims register is the one page whose entire purpose is claim-to-evidence fidelity. A correction that misdescribes the thing it corrects is worse than the original gap.",
      "how_caught": "Outside audit of the live site, 2026-08-26 (finding D8). One grep over the published cards.",
      "fix": "CR-002 now describes what the cards actually declare: nine body fields, none of them a timestamp authority; the only time a card carries is `created`, an instant the issuer asserted from its own clock and then signed, which attests assertion and not independent observation; and `prev` gives ordering, not time. The superseded wording is kept on the row under a dated `amended` note and rendered on /claims-register — a published claim is amended in the open, never rewritten in silence. Adding an explicit timestamp_authority: \"none\" to the card schema would be the stronger answer and is recorded as a change for the NEXT card format, not as a thing already done: each card id is the SHA-256 of its own body, so a new field re-mints every id and invalidates every published signature.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-06b",
      "date": "2026-08-26",
      "what_was_wrong": "/claims-register announced \"20 claims\" and rendered 19, immediately beneath its own sentence \"This page renders that exact file — there is no second copy to drift.\" The header printed claims.length while the sections were built from a hardcoded four-status order — live, devnet, planned, retired — and claims-register.json declares five. The fifth is `unmeasured`, and the one claim carrying it (CR-020) had no case in the renderer, so it was silently filtered out of the page and out of the legend. On a site whose banner is \"UNMEASURED shown honestly\", the register dropped the only unmeasured row. The wrong count was the visible defect; the dropped row was the worse one.",
      "how_caught": "Outside audit of the live site, 2026-08-26 (finding D9). The auditor diffed the rendered ids against the JSON. Nothing on our side compared the two — the drift the sentence rules out was never checked.",
      "fix": "The page now derives its status order from the file's own statuses[] and appends any status that appears on a claim but was not declared, so a row can never be dropped for wearing an unexpected label; `unmeasured` has a real chip and a real legend entry. The header count is the length of the rows actually rendered, not claims.length — a number on that page is now derived from what a reader can scroll to. If a row ever does fall out, the page says so in a visible RENDER DEFECT banner naming the id. scripts/claims-register-lint.mjs re-derives the grouping at build time and fails the build on any drift between the file and the page, including a declared status with no legend entry or a typed number back in the header.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-08",
      "date": "2026-08-26",
      "what_was_wrong": "For twelve days the verify page told strangers to pin a signing key that does not exist. The page's authorship note named a published key by an eight-character fingerprint beginning f4b4278d. That fingerprint matches none of the four keys in our DID document, not the card-attestation key the 150 board cards are actually signed with, not the board key, not the living-stamp key. It appears in exactly one place in the entire estate — that sentence — and in no signed artifact, no key file and no commit that produced key material. It was introduced on 2026-08-14 in a bulk copy reconciliation, alongside an OpenTimestamps anchoring claim that was itself later walked back. We cannot establish what it was, so we are not going to invent a story for it: it was a fabricated fingerprint, and a fingerprint is the one string on a page telling people which key to trust that has to be right. The real card-attestation key, beginning d4cb0eaa, appeared nowhere on that page.",
      "how_caught": "An outside auditor with no CSOAI code and no CSOAI credentials grepped the fingerprint across every page, the DID document and the card index, and found one occurrence and no key. Not self-caught. The estate had published a key-pinning instruction it had never once executed against its own page.",
      "fix": "The fabricated fingerprint is removed from both surfaces that carried it, the verify page and the agent registry. Both now name the anchor by its DID identifier, link the DID document so a reader can read the key out for themselves, and print the real key prefix. No provenance has been invented for the removed string, because none could be established.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-07",
      "date": "2026-08-26",
      "what_was_wrong": "Our own published verifiers rejected our own genuine cards, and our tamper detector rendered its failure in green. Three separate defects on the one surface whose entire purpose is that a stranger does not have to take our word for anything. First, the single-record verifier on the verify page hashed the whole card envelope minus the signature instead of the body sub-object the signature actually covers, so it could never verify any card, ever — and it reported that preimage bug as no published key verifies this signature, which is a statement about key publication and was false, sending readers to hunt for a key that was published all along. Second, the same form fed its verdict to a public opt-in tally, so every honest visitor who verified a real card and clicked the button filed a false failure into a public counter. Third, the MCP verify tool answered unrecognized card family to every card family we publish, including the cross-border card that verifies fine under our own recipe, because it looked for a content_id field on cards that carry id. Fourth, the client-side chain verifier's headline label was a constant string reading chain intact regardless of outcome; only the tick flipped to a cross, so a successfully detected tamper announced that the chain was intact, in the success colour, on the page that promises a broken row is reported as BROKEN, visibly.",
      "how_caught": "An outside SCITT implementer followed our post to the IETF list, verified a card in Python against our published recipe, then clicked our own verify button to cross-check and was told our card was invalid. Every one of these was reachable from the public site with a browser and curl. None was caught by us.",
      "fix": "There is now one verification implementation, shared by the browser form and the MCP endpoint, so the two surfaces cannot disagree again. It implements the published rule exactly, including the CPython number representation that renders an integral accuracy as 0.0 rather than 0 — 56 of the 150 cards carry such a value, and a verifier without that rule reports a false failure on 37 percent of a corpus that is sound. It recognises both published card families rather than rejecting both. Critically, it reports three failures as three different failures: the bytes do not hash to the declared id, the signature does not verify over those bytes, and the signer is not a key published in our DID document mean different things and are never collapsed into each other. The signer is pinned against the live DID document, so a card carrying an attacker's own key is reported as an untrusted signer even when its signature is internally valid. The tamper label now states the outcome in words and a failure no longer renders in the success colour. All 150 published cards verify through the fixed path, a tampered card fails as a hash mismatch, and a re-signed forgery fails as an untrusted signer. Regression tests read the real published bytes so these cannot silently return.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-06",
      "date": "2026-08-26",
      "what_was_wrong": "We repeated a human-versus-machine benchmark contrast without checking whether both sides were scored under the same rule. The metrology deck cites the ARC Prize project's ARC-AGI-3 result — a human panel solving essentially all environments while frontier systems average well under one percent. The attribution was correct and careful: labelled reported-not-measured, never placed on the board. The number is not the defect. The defect is that we published a comparison between a human figure and a machine figure without asking the question our own first rating-the-raters result exists to ask, which is whether the two figures were produced under the same scoring rule. Having now recomputed ARC's published participant rows for ARC-AGI-2, we know that on that benchmark the human figure is computed under unlimited submissions while machines are scored at two trials, and that the rule-matched human figure is about eleven points lower. We had no basis to assume ARC-AGI-3 was free of the same gap, and no basis to assume it had it.",
      "how_caught": "Self-caught, by our own instrument, on its first run. Building the RTR-A1 human-reference rule-match measurement against ARC-AGI-2 meant asking of another organisation a question we had not asked of our own published page. Sweeping our surfaces for prior statements about the same publisher is what surfaced it. This is the intended failure mode of a rating-the-raters programme: the first thing a new instrument should catch is its owner.",
      "fix": "The deck passage now carries the caveat, stated as a limit rather than a finding: a human-versus-machine contrast only means what it appears to mean if both sides were scored under the same rule; on ARC-AGI-2 we measured that gap; whether ARC-AGI-3 shares it is UNMEASURED because its scoring formula is not published, so we cannot check and will not assume either way. The general rule this establishes for every surface: CSOAI does not republish a human-versus-machine comparison without either verifying rule-match or marking it unverified. Nothing was removed and no third-party number was restated as ours.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-05",
      "date": "2026-08-26",
      "what_was_wrong": "Two published index artifacts claimed a measurement they did not have. /interop/ai-economy-index.v0.1.json and /interop/human-labour-index.v0.1.json each carry a status label of MEASURED-INDEX-v0.1, while each also states in its own body that half its input components are bank gaps and that no index value is computed. The axis register had already been reverted to UNMEASURED for both; the artifacts were not, so a live surface kept asserting the retracted status. Existing reference components are not a measured index.",
      "how_caught": "Reading the evidence behind every financial axis before wiring it into the board, rather than trusting the axis register's summary of it. The register said UNMEASURED; the artifact it pointed at said MEASURED-INDEX-v0.1. Following the pointer is what surfaced the disagreement.",
      "fix": "Both axes are wired into the signed board as UNMEASURED, and the board — which is the authority — states on each axis and in its limitations that the v0.1 artifacts' status label was an over-claim and is superseded. Neither index contributes to any measured count. The artifacts themselves are signed under a key this lane deliberately does not hold, so correcting them at source is a separate owner-supervised re-sign; until then the board carries the correction where a reader will meet it.",
      "status": "FIXED ON THE BOARD; ARTIFACT RE-SIGN PENDING (owner)"
    },
    {
      "id": "C-2026-0826-04",
      "date": "2026-08-26",
      "what_was_wrong": "The public board contradicted the estate's own ruling for two days. An owner ruling of 2026-08-24 set the canonical axis count at 22 (14 behavioural + 8 financial/domain), but GET /api/gspc kept reporting '14 measured of 14 quotable' because the 8 financial axes existed only in the ruling and in a side register — never in the signed board payload the count is derived from. Downstream, the estate's own claims register recorded '22' as an internal figure that was 'not corroborated by any live surface', and a source comment instructed authors to 'not invent 22 axes'. The estate simultaneously ruled the number, forbade the number, and published a different one.",
      "how_caught": "Self-reported, not discovered. The ruling document itself recorded that the sweep was authorized but unexecuted, and named the reason. The delay was deliberate and is the point of this entry: a public count must be backed by the signed artifact it summarises, so the fix could not be a copy edit on the pages. Editing the number without the data behind it would have put a figure on a public surface that the signed payload could not support — the same defect class as a score published without its measurement. The board was behind the ruling, never ahead of it.",
      "fix": "The 8 financial/domain axes were wired into the board DATA and the payload re-signed. The board now derives '22 axes · 15 measured' from the axis array: 22 slots, 15 with a real run behind them, 7 declared slots with none. The ruling's own wording applied the word 'measured' to the full slot count, and the evidence does not support that word — only one of the eight financial axes (provenance-controls, a deterministic mainnet read of 6 issuer accounts) carries a measurement. Per this ledger's redaction rule the exact phrase is described rather than reproduced: it is now the forbidden form the build gate catches, and reprinting it here would republish the sentence this correction exists to retire. No axis was marked MEASURED to make the two numbers agree; the grammar changed instead, and both numbers now travel together. Separation statistics and every mean are scoped to model-comparison axes, so a financial axis can neither enter a sentence about statistical separation nor drag an absent value into an average as a zero. The claims register was re-authored from 'internal, not corroborated' to a live claim with the endpoint as its authority, and now names the forbidden form '22 measured axes' explicitly.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-03",
      "date": "2026-08-26",
      "what_was_wrong": "Our own published MCP fleet was silently paywalled and self-scoring. A monetization layer injected into 318 of 363 vendored servers capped the ENTIRE fleet at 10 anonymous tool calls per day from one shared counter; past that, every tool returned a purchase link instead of a result. The injected code was spliced mid-function in 49 files, leaving original function bodies unreachable (256 undefined names). Five scorecard checks awarded points for carrying a purchase link — the system scored itself higher for being paywalled. The paywall also masked quality: a first probe found 1 stub because refusals and stubs were indistinguishable.",
      "how_caught": "Building a remote MCP server for other AI platforms; the first real tools/call returned a purchase upsell instead of a result. Verified twice independently by direct grep and by probing all 338 servers with real MCP sessions.",
      "fix": "Monetization layer removed fleet-wide: 318 -> 0 servers carrying a purchase link, 0 price strings, 0 upsell symbols. Capability preserved and proven, not assumed: all 338 servers re-probed with real initialize/tools/list/tools/call — handshakes 336/338 unchanged, 1869 tools unchanged, 0 broken; undefined names fell 256 -> 16 because removing the injected code repaired what it had broken. Honest stub register published (13 fully stubbed, 10 partial, 2 dead) determined by CALLING every tool, not grepping. scripts/no-paywall-guard.mjs added with a --selftest so the layer cannot return; it caught 48 residuals we had missed.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-02",
      "date": "2026-08-26",
      "what_was_wrong": "Five sector pages asserted, in present tense, that our measurement 'is recognised under mutual recognition agreements with' CISA, NCSC, ANSSI, BSI, BEREC, ENISA, national transport authorities and others — named public bodies, implying an endorsement we do not hold. It shipped in the deployed bundle. Separately, /layer0 served a retracted fault-tolerance claim as a live capability, contradicting our own DR-0007 retraction (measured effective independence 1.21 of 3).",
      "how_caught": "Claims-substantiation audit of the prerendered output, prompted by the FTC's own recommended exercise: inventory every public claim and map it to evidence.",
      "fix": "Replaced with: we crosswalk our measurement output to those compliance pathways, and hold no mutual-recognition agreement with, and are not endorsed or accredited by, any of these bodies. The retracted claim removed from /layer0, /poc-showcase and /competitors. A machine-readable claims register now publishes every claim with its evidence link and a live/planned/devnet/retired status.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0826-01",
      "date": "2026-08-26",
      "what_was_wrong": "Our own prerender verification could not observe failure. prerender-report.json records a failed route in a field named 'err', but every check in the repository read 'errored' — a field that has never existed. A run in which the browser died on 515 of 581 routes reported '0 errored' and looked clean.",
      "how_caught": "A downstream gate disagreed: brand-gate scanned 71 pages when it should have scanned 603. The upstream report was lying and the layered gate caught it.",
      "fix": "scripts/check-prerender.mjs reads the real fields AND cross-checks the report against the HTML actually written to disk, because a report is a claim and the files are the evidence. It fails loudly on the exact run that had been called clean.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-01",
      "date": "2026-08-19",
      "what_was_wrong": "Three public surfaces stated three different item counts at once (llms.txt 819, agent card 890, live API 966). The banks grew under the hardcoded numbers.",
      "how_caught": "External live-surface audit; confirmed by direct curl.",
      "fix": "llms.txt and the agent card now DEFER to GET /api/gspc as the live source; no public surface hardcodes a count.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-02",
      "date": "2026-08-19",
      "what_was_wrong": "The public board API payload carried internal specialist identifiers — an internal specialist-id prefix — a banned-vocabulary string inside a machine contract, not just a human page. (The prefix itself is redacted here: naming it would re-leak the string this entry records as removed.)",
      "how_caught": "K3 lane curl sweep of machine surfaces.",
      "fix": "Renamed to council-* public names in /api/gspc; a machine-contract guard now sweeps API payloads for banned strings on every deploy.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-03",
      "date": "2026-08-19",
      "what_was_wrong": "The single-record verifier initially checked only one content_id envelope; the carder signs a second (signature-included) generation, so valid carder cards could have read as MISMATCH.",
      "how_caught": "Testing the verifier against a real carder card before shipping.",
      "fix": "The verifier now tries both deterministic envelope generations and names which one matched.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-04",
      "date": "2026-08-19",
      "what_was_wrong": "Two open-source repos (carder, codabench-gspc) shipped with no LICENSE file, and the board API payload stated no licence — while the estate claims openness.",
      "how_caught": "The carder's own valve-2 benchmark fact-card, run on the estate's own artifacts.",
      "fix": "Apache-2.0 added to both repos; CC-BY-4.0 licence field added to the board payload, with the self-catch admitted in the payload note.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-05",
      "date": "2026-08-19",
      "what_was_wrong": "The did:web trust root at csoai.org intermittently served an orphan key document because two repositories deployed the same Cloudflare Pages project with no owner of record.",
      "how_caught": "The did-liveness daemon, then the machine-contract guard's DID split-brain check comparing the authoritative root against the mirror.",
      "fix": "One deployer of record (csoai-site-deploy.yml) builds from the source repo's main with a hard gate: the build fails if did.json lacks the canon keys, and the run fails if the live apex doesn't serve them after deploy.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-06",
      "date": "2026-08-19",
      "what_was_wrong": "An hourly API guard asserted endpoints (/api/tools, /api/mcp) that never existed in the repository's functions tree — a ghost from an older deployment — so it failed forever.",
      "how_caught": "Reading the failing run rather than trusting the guard's own claim.",
      "fix": "Rewritten to assert the endpoints the deployment actually ships (/api/health, /api/leaderboard).",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-07",
      "date": "2026-08-19",
      "what_was_wrong": "A banned brand token shipped live on /library as a CamelCase concatenation of the token with 'Training', because a word-boundary regex anchored on the bare token missed the concatenation. Two priced strings ($0.005/card, a per-hour range) also shipped, against the no-pricing rule. (The token itself is redacted here for the same reason as C-2026-0819-02.)",
      "how_caught": "A full front-end QA sweep.",
      "fix": "The brand gate's pattern for that token dropped its trailing word boundary so CamelCase concatenations are caught; a pricing-leak pattern was added so a currency amount bound to a subscription or per-unit cadence is now a hard build-fail.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-08",
      "date": "2026-08-19",
      "what_was_wrong": "Estate pages described EU AI Act high-risk obligations as in force from 2 August 2026. The Digital Omnibus (Reg (EU) 2026/1744) deferred them to 2 December 2027 (Annex III) and 2 August 2028 (Annex I). Serving the dead date would be our own credibility wound.",
      "how_caught": "A commissioned regulation-calendar verification against primary law.",
      "fix": "The /api/regulation feed carries the corrected staged timeline with legal bases; page copy is being swept to match.",
      "status": "IN_PROGRESS"
    },
    {
      "id": "C-2026-0819-09",
      "date": "2026-08-19",
      "what_was_wrong": "Two internally-named datasets remained publicly visible on Kaggle under a banned naming class.",
      "how_caught": "End-user test sweep with anonymous probes.",
      "fix": "Flagged for the owner to set private — the platform gates dataset visibility behind the account login.",
      "status": "OPEN"
    },
    {
      "id": "C-2026-0819-10",
      "date": "2026-08-19",
      "what_was_wrong": "The estate's own date-correction fix (C-08) initially ALSO mis-stated the GPAI date — a follow-on error that moved GPAI duties from 2 Aug 2025 to 2026 while correcting the high-risk date. A correction that introduces a new error is the worst kind.",
      "how_caught": "Self-audit of the fix against the EU official page (digital-strategy.ec.europa.eu) — the estate caught its own owner mid-correction.",
      "fix": "GPAI 2 Aug 2025 restored; Article 50 2 Aug 2026 and high-risk 2 Dec 2027 (Annex III) / 2 Aug 2028 (Annex I) stated distinctly. This entry is that admission, appended not edited.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-11",
      "date": "2026-08-19",
      "what_was_wrong": "mcp.json advertised three server URLs on csoai.org/api/* — every one returned 404 because the API is served from councilof.ai, and one route (corpus-watch) pointed at a non-existent path.",
      "how_caught": "End-user MCP handshake test — a real JSON-RPC initialize probe against the advertised endpoints.",
      "fix": "mcp.json now advertises councilof.ai URLs and the real /api/corpus-watch/status route; the advertised endpoints were verified 200/JSON-RPC-responsive after the fix.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-12",
      "date": "2026-08-19",
      "what_was_wrong": "A measurement wave was queued with sample=24, below the harness's 30-usable-item threshold — all 8 jobs returned UNMEASURED (honestly, but wasted a full wave).",
      "how_caught": "Reading the signed board's status_note ('no model reached 30 usable items') rather than assuming the bank size was the constraint.",
      "fix": "Requeued at sample=30; all 8/8 came back MEASURED and signed. The threshold is now documented in the job-spec contract.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0819-13",
      "date": "2026-08-19",
      "what_was_wrong": "Two measure-chain daemons ran simultaneously after a restart race, double-logging jobs; the restart script's pkill pattern matched its own command line and killed its own launch.",
      "how_caught": "Duplicate 'daemon start' markers in the log; the self-kill was traced to the unanchored pkill pattern.",
      "fix": "Anchored process pattern (^python3 /workspace/measure_chain.py) in the restart script; single-daemon verified after relaunch."
    },
    {
      "id": "C-2026-0820-01",
      "date": "2026-08-20",
      "what_was_wrong": "Multiple live public surfaces (index.html JSON-LD, GSPCVerify, Insurers, AgentRegistry, Methodology, Agents, ProvBench, measure.html, and the provbench pack) stated measurement cards are 'anchored with OpenTimestamps' / RFC-3161 / 'Bitcoin block 954857, independently verifiable' as a present capability. The only anchor implemented is Ed25519 + SHA-256 hash-chain; verify.ts checks no timestamp proof and no .ots/Rekor artifact exists.",
      "how_caught": "Internal honesty audit of anchoring claims vs implementation.",
      "fix": "OTS/RFC-3161/Bitcoin claims demoted to roadmap wording across all surfaces; provbench pack corrected; the ML-DSA 'built, not shipped' discipline applied to OpenTimestamps.",
      "status": "FIXED"
    },
    {
      "id": "C-2026-0822-01",
      "date": "2026-08-22",
      "what_was_wrong": "The homepage industry grid still said '15-slot instrument' while the scoreboard, API and canon say '14-slot board, 13 measured of 14' (16 GSPC axes, 13 quotable + jail floor per the GSPC ruling). A crawler reading the grid would see 15 slots — the exact internal-count inconsistency the count-gating canon exists to prevent.",
      "how_caught": "Text audit of live surfaces against canon (machine-contract style sweep of the homepage and fleet-sweep pages).",
      "fix": "Killed both stale 15-slot references in NewHome-v3 (section comment + industry-grid subtitle) to '14-slot / 13 measured of 14'; verified 0 x '15-slot' remains. (PR #284.)",
      "status": "FIXED"
    }
  ],
  "signature": {
    "id": "9224bbec358ac2f375a84c7dd5d4e92cf2dc9fce4c60aa3cc9c609deb0ded4a1",
    "signer": "9367cf59be9cb72bbc9796adf056201ec1c58adfeaa13f83b2c5b754d6c20170",
    "did": "did:web:csoai.org#board-attestation-1",
    "signature": "fd5e0b6d4b90f2e34607f523695b5b555fa764249886b502a270364cd9bf97f8a8630e6ca5c22d7b8b30cd357621c0e046caebb520fabeaee59d105a08ac0e03",
    "attestation": {
      "artifact": "csoai.corrections/0.1",
      "content_id": "9224bbec358ac2f375a84c7dd5d4e92cf2dc9fce4c60aa3cc9c609deb0ded4a1",
      "content_id_rule": "sha256(json.dumps(served body minus keys [\"signature\",\"signature_state\",\"signature_check\",\"correction_latency\",\"note\",\"fix_requires\"], sort_keys=True, separators=(',',':'), ensure_ascii=True))",
      "entries": 63,
      "latest_entry_id": "C-2026-0923-02",
      "ledger_canonical_bytes": 94233,
      "note": "Detached. The Ed25519 signature covers THIS object; the ledger body is committed to by content_id because it is larger than the signer's 3KB payload cap. Both must check: the digest must still describe the body a reader just fetched, and this object must verify.",
      "schema": "csoai.corrections-attestation/0.1",
      "signed_at": "2026-09-23T05:31:18Z"
    },
    "sig_input": "Ed25519 over json.dumps(signature.attestation, sort_keys=True, separators=(',',':'), ensure_ascii=False) - the attestation is ASCII-only, so ensure_ascii does not change its bytes. The attestation names the digest of the ledger body and the rule that produces it.",
    "key_source": "https://csoai.org/.well-known/did.json (did:web:csoai.org#board-attestation-1)",
    "note": "RE-ISSUED 2026-09-22 over the current body through POST /api/board-sign on the pod caller token. The 2026-08-22 signature was under did:web:csoai.org#card-attestation-1 (d4cb0eaa) and covered a 15-entry ledger; 46 appends followed and none re-issued it, which is why this endpoint read STALE for a month. Every append MUST re-issue: run scripts/sign-corrections-ledger.mjs. Bumping id alone cannot green the flag any more - id is inside the signed attestation, and the handler verifies the Ed25519 bytes at request time, not just a digest match."
  },
  "signature_state": "VALID",
  "signature_check": {
    "state": "VALID",
    "checked_at": "2026-09-28T17:26:21.405Z",
    "recomputed_content_id": "9224bbec358ac2f375a84c7dd5d4e92cf2dc9fce4c60aa3cc9c609deb0ded4a1",
    "attested_content_id": "9224bbec358ac2f375a84c7dd5d4e92cf2dc9fce4c60aa3cc9c609deb0ded4a1",
    "content_id_matches": true,
    "ed25519_verified": true,
    "key": "did:web:csoai.org#board-attestation-1",
    "key_ed25519_hex": "9367cf59be9cb72bbc9796adf056201ec1c58adfeaa13f83b2c5b754d6c20170",
    "unsigned_wrapper_fields": [
      "signature",
      "signature_state",
      "signature_check",
      "correction_latency",
      "note",
      "fix_requires"
    ],
    "how": "Computed on this request, from these bytes. Strip unsigned_wrapper_fields from this document, canonicalise with json.dumps(sort_keys=True, separators=(',',':'), ensure_ascii=True), SHA-256 it: that is recomputed_content_id and it must equal signature.attestation.content_id and signature.id. Then verify signature.signature as Ed25519 over the same canonical form of signature.attestation under key_ed25519_hex, which is the published key for did:web:csoai.org#board-attestation-1. Both must hold.",
    "means": "VALID: both held. STALE: the signature verifies but the body has moved since it was issued, so it no longer describes what you are reading. INVALID_SIGNATURE: the bytes do not verify under the published key. UNSIGNED: no signature is published. UNCHECKABLE: this runtime could not perform Ed25519 — not a claim about the signature either way."
  },
  "note": "The signature was verified on this request over these bytes, under did:web:csoai.org#board-attestation-1. Re-issued 2026-09-22 after a month of reading STALE: the ledger had been appended 46 times since it was signed on 2026-08-22 and nothing re-signed it. See signature_check for how to reproduce this yourself.",
  "correction_latency": {
    "computable": 0,
    "unmeasured": 63,
    "field": "detected_at (optional per entry, added 2026-09-13)",
    "note": "Measured only where both dates are explicit fields; never inferred from prose."
  }
}