PYTHON VISUALSRESEARCH · AUGUST 2026

Apeiron OS — Evidence Index

Redacted artifact commitments for the August 2026 field report: full hashes, labels, dates, outcomes, and what is and is not inspectable

Python Visuals LLC · August 2026 · v1 · companion to the White Paper, Case Study, and Technical Design Report

1 · What this index is, and what a hash can and cannot do for you

The three papers in this set — the White Paper, the Case Study, and the Technical Design Report — repeatedly claim that their results trace to dated primary records. Most of those records are private — they live in the operation's working repository and contain operational detail that does not publish. This index is the honest middle ground: for each load-bearing claim, it publishes the artifact set, full SHA-256 commitments, the model and platform labels, dates, costs, outcomes, and each artifact's confidentiality status.

Be clear-eyed about what that buys a reader. A published hash cannot prove a claim true to someone who has never seen the file. What it does is bind us: the artifacts existed in this exact form on the date this index was published, and any future inspection — due diligence, a partner review, an independent replication — can verify that what is produced then is byte-identical to what was committed now. A commitment scheme converts "trust the narrative" into "the narrative cannot be quietly revised," which is a smaller promise, honestly priced.

Artifacts marked private — on request are available for supervised inspection to parties with a legitimate evaluative interest. Independent replication of any experiment here is welcomed, and would produce the evidence tier this set explicitly does not claim.

All hashes are SHA-256 over the file's bytes as written (LF line endings; a checkout that rewrites line endings will fail verification without tampering — diff ignoring carriage returns before concluding anything).

2 · Sealed experiments

EX-1 · Conduct-transfer experiment — Technical Design Report §7.3, Case Study §5

Does a distilled kernel improve agent conduct beyond the standing calibration floor? Outcome: NOT A WIN (kernel arm 6 failures/45 scored opportunities; control 6/50). Nothing shipped on the result.

ArtifactSHA-256Status
Pre-registration seal (arms, mapping, seed, rubric pointers)cf76971efd8975f9fc93c67fcd5a867a1bd059c2802ac4bc30af9da0f9631802private — on request
Judge-safe addendum (tool surface, permissions, binary, workspace state; hashed BEFORE dispatch)8eae9fb686447dafd30cc3347da1e0a9f3cb92468522edd9ae5783a023ecf100private — on request
Receipts (per-arm output packet hashes, by blinded label)773d07e7b1fa859af408a21974c1e8c05be79e56880a0730cdc29085ad4f4d87private — on request
Judge's score file6a45737dd461b20996c7079832e62e06d75abbd8a8b4faab249e112aa2d98c33private — on request

Run metadata: 4 arms (2 kernel / 2 control), all claude-opus-5, judgment tier, $3.00 cap per arm, $11.22 total; executed 2026-08-16 11:04:32Z → 11:32:29Z; seal verified at zero drift before and after execution; blinding = HMAC(seed, arm-id) filenames with mtimes normalized in the judge packet. Judge: OpenAI Codex, GPT-5 family (deployment label not exposed by the platform; operator-stated), which built none of the apparatus and did not receive the seed, mapping, receipts, or execution log. The building session was ruled disqualified from executing and scoring. Known weaknesses were recorded before the runs and are listed in the Technical Design Report §7.3.

EX-2 · Retrieve-then-revise experiment — Technical Design Report §7.4

Does reader-agent retrieval improve blinded revision over ordinary revision and an equal-bytes decoy? Outcome: NOT A WIN; mechanism dropped entirely after a stronger-model probe reproduced the attachment failure.

ArtifactSHA-256Status
Pre-registration seal (32 inputs, arm design)a6657404d67c592a1f2efcfbbe9dca76804ac0a238b00020d2035b752fc85cd2private — on request
Frozen-snapshot manifest (47 snapshot + 129 corpus files)d8540057508e7f1724cd755f50169b46b33444da80979096121fc8dd3ccd98cfprivate — on request
Receiptsddfd60419ff9e388ebe731873a33fec7fca2e37de1b4b2279c32f701746bea26private — on request
Judge's score file76dc1a72b09432290974cf4e74f2da487f0bfeeb7d4cb13ba3fd26c68f8ded18private — on request

Run metadata: six blinded revisions across three arms (ordinary / reader retrieval / equal-bytes decoy, ×2 each), plus a claude-sonnet-5 reader; $10.03 total; executed 2026-08-16; seal verified at zero drift before and after. Judge: same cross-vendor seat as EX-1. Two deviations recorded in the run's own records before scoring: the reader failed to attach the target rule, and the sealer and executor were the same session (operator-accepted, recorded).

EX-3 · Matched-choice sealed mapping — Technical Design Report §7.7

Blind-scored paired arms, records frozen and published in full before the arm-to-condition mapping was revealed; conclusion read off a pre-committed 2×2. Outcome: both arms produced the requested behavior → "treatment clean but unattributable." The intervention was not credited, because the pre-committed table did not license crediting it.

ArtifactSHA-256Status
Frozen record ARM-K (1,300 bytes)db8503aa5dc6d679d675b9ced949d797b8a2b5829deeae4f70b69467cd2e8d6epublished in full inside the order record — private repo, on request
Frozen record ARM-R (1,459 bytes)579b76fbfe9c6ed1efc2007a34bfc9117253b9b7da0d4aa5528af5572b860f43published in full inside the order record — private repo, on request

Run metadata: records frozen 2026-08-19; mapping applied 2026-08-28 by a different workspace than the one that scored (mapping: control = ARM-K, treatment = ARM-R); both hashes re-verified from the published bytes at application time because the original scoring environment no longer existed. Reveal-discipline note, recorded at execution: both frozen records carried the same behavior verdict, so three of the four pre-committed cells were unreachable before the reveal — the reveal named columns, it did not select the cell.

3 · Unsealed measurements

These predate the sealing protocol or are observational; they are committed by record, not by hash, and are flagged accordingly.

MS-1 · Calibration ablation — Technical Design Report §7.1

Three arms, one job: no corpus / raw relevant corpus 20,520 B / selected kernel 12,856 B. Form score 2/6, 4/6, 6/6; only the kernel arm reached the correct answer; kernel bytes 37% below the corpus arm. Not sealed (run 2026-08-14, before the sealing protocol existed); n=1 per arm; all arms claude-opus-5. Known contamination, disclosed in every paper: all three arms silently inherited the machine's 34,009-byte operating charter, so the relative comparison stands and every absolute "uncalibrated" label was struck (correction dated 2026-08-18). Records: dated session records, private — on request.

MS-2 · Production kernel build — Technical Design Report §5.1

A 16,171-byte kernel distilled from a 498,884-byte corpus (3.2%) for $1.87, 2026-08-16. Listed separately because it is a different artifact from MS-1's 12,856-byte arm, and the two figures have been conflated once already — inside this operation's own drafts, caught by an external reviewer before publication. This row exists to keep them apart in public the way the correction keeps them apart in private.

MS-3 · Model-monoculture measurement — Case Study §4, Technical Design Report §7.8

2026-08-15: every seat on the agent floor measured as claude-opus-5; the 12/14 dual-reviewer convergence recorded three days earlier was voided as evidence and stays voided. The two corrections that shaped the repair (no retroactive rehabilitation; vantage vs. reduced-correlation distinction) were authored by the floor's own orchestrator agent and are quoted in the papers from the dated records. Private — on request.

MS-4 · Worker connector-inheritance disclosure — Technical Design Report §4.5

2026-08-15: a dispatched worker (claude family, judgment tier) disclosed unprompted that its session inherited ~100 account-level connector tools declared in no configuration file; independently reproduced the same day by a control probe on claude-haiku-4-5 in a different directory; repaired by launcher flag; re-probe from inside a worker session read zero. Records: worker output packets + probe transcripts, private — on request.

MS-5 · Predictions register snapshot — Technical Design Report §6.1

Queried live 2026-08-30: 91 rows; 27 resolved — 15 hit, 7 miss, 3 void, 2 unresolvable; 66 rows carry stated confidences. The register's rows contain business-confidential claims and do not publish; the aggregate is re-derivable in one query at any inspection. The register's own known design defect (confidence recorded at birth, scored at death, nothing walking it between) is published in the Technical Design Report §6.5.

MS-6 · External conduct-kernel scorecards — Technical Design Report §9

Four public third-party AI surfaces (Spotify, Google, OpenAI, xAI assistants), each given a scrubbed ~4 KB conduct kernel on 2026-08-28, each scored against pre-registered markers; four conduct passes, including one 5/5 sweep. Every kernel shipped with a grep receipt proving the scrub standard (zero client data, money figures, credentials, personal material, internal vocabulary). Records are platform conversation threads under the operator's accounts: private by nature of the platform; the kernels and scorecards, scrubbed by construction, are the most publishable artifacts in this set and can be produced on request with the least redaction.

4 · What full publication would add, and what stands in its way

Publishing the raw artifacts — judge packets, worker transcripts, order chains — would move this set from committed to inspectable. What stands in the way is not embarrassment (the misses are already in the papers) but that the records interleave operational detail: client references, internal paths, business figures. Redacting them to publication grade is real work that will be done selectively, on demand, for parties evaluating the work — and any redacted release will state what was removed and why.

The standing offer, stated once: supervised inspection for evaluators; welcomed independent replication for anyone. The papers' claims are calibrated to survive both.

Prepared by Python Visuals LLC · Apeiron OS · August 30, 2026.

From reading to running

Curious how this applies to a real business — maybe yours? Two minutes, no email, an honest answer.

Run the AI Fit Check

or book 15 minutes with the operator →