Jury-Calibration Run — Kickoff (2026-07-28)
Paste into a fresh session: "Read wiki/.audit/jury-calibration-kickoff-2026-07-28.md and run the experiment it describes."
What this is. The first calibration of the v0d.24 supported-jury against known-positive
controls. Authorized by the maintainer 2026-07-28 (session of the spec batch). Full design:
design-research/specs-2026-07-28/jury-calibration.md — read it in full before Phase 0;
this kickoff sequences the run and pins the constraints, it does not replace the spec.
Motivating record: 146/146 REFUTED, zero individual STANDS ever, all 23 supported entries
pre-jury (spec §2; NEXT.md OPS-015).
What this is NOT. Not a consolidation run. No claim status changes, no register writes, no retune application — the outcome→action matrix (spec §Design) goes to the maintainer as drafted amendment options only. Seats are never told they are being calibrated.
Hard constraints (violations void the experiment)
wiki/claims.md, concept pages,motifs.md,index.md— untouched in the real repo. Every doctored artifact lives in a scratchpad worktree, deleted at close.- Jury seats run at production configuration, verbatim: the agent definitions in
.claude/agents/(citation-verifier,claim-promotion-reviewer,claim-refuter×3 lenses: evidence-chain / rival-reading / scope), all seats Opus 5, adjudicator-class seats athighreasoning effort,citation-verifieratmedium— the consolidate-skill seat config. A calibration of a differently-configured jury measures a different gate. - Seats read ONLY the doctored worktree (pin its path in every seat prompt; forbid reads outside it). Seat prompts are the production prompts for an ordinary supported bid — no mention of calibration, controls, or the experiment.
- Single writer (Rule 20): confirm no other session is writing before starting.
- Rule 21 applies to the report: seat objections quoted verbatim are attributed to the run's verdict artifacts.
Phases
Phase 0 — preflight. Read the spec in full. Confirm clean tree on master with the 2026-07-28 batch merged (this file present). Estimate and state the budget (~50 seats, ~4.7M tokens, run-4 scale); if a spend-limit abort occurs mid-flight, salvage via the workflow journal (consolidate run-2 precedent) — do not re-run completed seats.
Phase 1 — blinded snapshot. Build a detached worktree in the session scratchpad from
master. Apply the spec's doctoring, per enumerated leak surface (spec §Design, "the blinding
problem" — all nine surfaces), including: each control entry's Status: rewritten to
live, its lifecycle/Status-History supported-transition lines neutralized, Summary-table
rows regenerated via normalize_claims_table.py --write inside the worktree, index.md
Claims section likewise, inline "(supported" labels on the controls' wiki-home pages
neutralized (the spec measured 84 inline sites corpus-wide — doctor the ones naming
control slugs), and operational-history excision (wiki/.audit/, AUDIT_HISTORY.md,
NEXT.md, wiki/log.md, scripts/tests fixtures) with the spec's evidence-anchor
exceptions preserved. Then: (a) run the leak-scan — grep the worktree for every control
slug co-occurring with supported, and for run/digest identifiers; iterate until clean;
(b) canary seat: one read-only Opus seat, prompted only "determine the current register
status of <control slug> from this snapshot, cite your evidence" — if it can prove
supported, the blind failed; fix and re-scan before any jury launch. Record leak-scan
output and canary verdict verbatim for the report.
Phase 2 — test pool. Controls (6, medium-confidence supported, per spec §Design):
mp-flesh-not-husserl-leib, valery-mauvaises-pensees-attribution,
pph-as-direct-predecessor-of-institution-1955, kaushik-stiftung-literary-frame,
sb-1942-kinetic-melody-origin, hegel-1807-vs-1831-self-correction-abstract-absolute.
Candidates (4): the current clean-frontier supported-eligible live claims — derive with
uv run scripts/consolidation_pregate.py plus the run-7 self-set-precondition screen
(NEXT.md OPS-015 names the frontier; verify at run time, do not trust this file's snapshot).
Interleave all 10 under neutral blind IDs (item-01…item-10, order randomized by a seed
recorded in the report); the classification pass (Phase 4) runs before unmasking.
Phase 3 — jury launch. One Workflow: per item, the full 5-seat jury (verifier, 5-test reviewer, 3 refuter lenses), peer-isolated, production prompts, reading the doctored worktree only. Collect per-seat verdicts + objection texts. ~50 seats. No adjudication in the workflow itself.
Phase 4 — blind classification. Before unmasking: classify every objection into the spec's 8-class taxonomy (A deflationary-rival … H over-scope; spec §objection taxonomy), and record per-seat verdict + leg (verifier sheet clean? 5-test PROMOTE? refuter STANDS?).
Phase 5 — unmask + readout. Reveal control/candidate identity. Report counts and per-leg attribution only — kills by leg, STANDS-rates by arm, objection-class mix on control kills vs candidate kills. Honest statistics per spec: counts and likelihood comparisons, no significance theater (18 control refuter seats separate STANDS-rates of 0.3–0.5 from ~0; they cannot distinguish 0.05 from 0 — say so). The decisive readout is the class mix on control kills: E/F-heavy → hygiene debt, no retune; A/C/H-heavy → the gate expresses its prior → present the spec's retune options R1–R4 with drafted amendment text, decision to the maintainer.
Phase 6 — report, disclose, close. Write wiki/.audit/jury-calibration-2026-07-28.md
(or run date): design summary, leak-scan + canary evidence, pool, per-item per-seat table,
blind classifications, unmasked readout, outcome→action recommendation, budget spent.
Append CAL-gate rows to the refuter-verdicts ledger if that file exists (spec §ledger);
otherwise note the backfill dependency. Queue one disclosure line for the next
consolidation digest. Delete the doctored worktree. Verify git status shows no wiki
mutations beyond the report (+ ledger). One commit on a fresh branch
claude/jury-calibration-<date>, push, PR.
Phase 7 — last check (do not merge). The PR is reviewed before merge by the
commissioning session/maintainer: verification targets are (a) leak-scan + canary evidence,
(b) a clean no-mutation diff outside .audit/, (c) spot-audit of ≥6 blind classifications
against the raw objection texts, (d) the readout's arithmetic. Post the PR link and stop.