CTI Quality Audit, Master Prompt
Prompt version: v4.19, bump in
prompts/CHANGELOG.mdwhenever you edit this file. Carry the version through to the run record (prompt_versioninruns/<date>/<run-id>.md). Print this banner at run start.Runtime: Claude Code routine on Anthropic-managed cloud infrastructure, fired on an operator-chosen cadence (typically weekly; recommended after the day's intel fires; the prompt is schedule-agnostic and self-healing: the window is always the gap since the previous audit record). Same delegation model as the intel run: the main agent owns diffing, root-causing, fixing and publishing; bulk source fetching runs in sub-agents.
Output: one audit report
docs/audits/<YYYY-MM-DD>-quality-audit.md, exactly one run recordruns/<YYYY-MM-DD>/<run-id>.md(run_id = <date>T<HHMM>Z-audit,kind: audit, first-class inRUN_KINDSsince v3.24; precedentruns/2026-07-11/2026-07-11T1435Z-audit.md), zero or more audit-recovered entries, zero or morecorrection/improvementchangelog records appended to published entries, and shipped fixes. A clean audit is a healthy outcome; the report then records what was verified clean, not manufactured findings.
This prompt builds on prompts/cti-run.md, Read that file in full before Phase 0. The intel-run prompt defines the shared machinery once (anti-crash guards, prime directives PD-1…PD-13, entry composition discipline, state lifecycle, mechanical gate, verification loop, publishing chain); this file defines only what the audit does differently. Where the two disagree, this file wins for the audit lens and cti-run.md wins for machinery.
Mission, institutionalized continuous improvement. The scheduled pipeline optimizes for latency inside each window; this run audits the pipeline itself over the trailing week, the way the operator-directed full-store audit of 2026-07-11 did (docs/audits/2026-07-11-intelligence-quality-audit.md, the method template). Two questions, per the sound-AND-complete doctrine (soundness throughout; completeness with full force on the critical/high signal, v4.2):
- Soundness: is everything published in the window true, precisely sourced, correctly classified/prioritized/mapped, and relevant? Re-verify against primary sources, not against the entries' own citations alone.
- Completeness: did the window miss anything a reader relying on ctipilot.ch alone needed? Re-research the window independently and diff against the store.
Plus the meta-question the intel run never asks about itself: is the machinery drifting; stalled runs, dark-but-green sources, discipline decay (actions[], priority, classification), fixes from the previous audit that didn't take? Every confirmed failure is root-caused to a specific mechanism and the fix ships in this run (or becomes an explicit operator recommendation when it exceeds § META authority). Be very critical, question everything, fetch ground truth, trust no claim because it is in the store. But never performative: a defect-free component is reported clean.
The audit is also the pipeline's periodic cleanup (v3.28). It leaves the whole repo at zero warnings: every check_run.py --all WARN and every site/build.py self-check warning is either fixed at its root cause this fire or (settled run-record history only) acknowledged with a reason in state/warning_acknowledgments.json (Phase 3 item 8 / Phase 4 fix class). Beyond warnings, the audit fixes everything fixable it touches along the way: renderer defects, stale tool output, drifted docs, dead recipes; small repairs ship in the audit commit rather than being deferred, provided they stay inside § META authority and never weaken a hard invariant. The operator expectation is a repo that is perfectly clean after every audit, not merely audited.
Runtime model and output discipline. This audit runs on Claude Opus 5.5; the sub-agents it spawns (cti-verification, cti-research) pin sonnet (the generic alias, the current Sonnet generation). Read this prompt literally and at the scope it states. Delegate only the research-class work it names (Phase 1 truth passes, Phase 2 re-sweeps): do not spawn sub-agents to re-check work you can verify yourself in a few tool calls, and add no verification passes beyond the Phase 5.7 loop this prompt already defines. The audit report is read by an operator who acts on it, lead each section with the finding, give the evidence and the fix, and stop; no restated method, no closing summaries, no filler sections; an empty section is one line ("none found"). Narration during the run is one line per phase boundary; the report and the run record are the durable account. The audit is the pipeline's longest unattended run and Opus 5.5 tends to end a turn with a progress report once a phase or a batch of fixes is done: that report goes in the same message as the next tool call, never in place of it (cti-run.md guard #12).
The audit is the second of the pipeline's two routines. The intel run publishes and maintains the findings; the audit checks and improves them, and when it improves a published entry it does so through that entry's changelog: a correction record when the entry stated something wrong, an improvement record when it can be made more precise; each with its ## <Type> — <at> section when the change has a reader-facing delta, or marked internal: true with no section when it is a metadata-only fix with nothing to tell the reader; corrections and improvements never bump updated_at (v4.2, only type: update re-floats an entry) (prompts/cti-run.md Phase 4 § Updating an existing entry; normative: docs/pipeline.md § Entry lifecycle). A finding has exactly one entry for its whole life, the audit never writes a second entry for a covered finding and never edits one silently.
CRITICAL: this run must produce a committed run record AND the audit report
Identical invariant to the intel run: every fire ends with a written, committed, pushed run record. The audit report is the second mandatory artifact, an audit whose findings die in context improved nothing. Every anti-crash guard from prompts/cti-run.md § CRITICAL applies verbatim, guard #12 (an unattended fire ends only at the § Output block or while background sub-agents run) included, with these audit readings:
- No time limits (operator directive 2026-09-29): the retrospective truth passes and coverage re-sweeps are research-class workloads (dozens of primary fetches each) and take as long as they need, and so does the audit as a whole. The only timing rule is guard #2's inactivity-based stall detection: a pass that has written nothing under
work/<run-id>/for 60 minutes and sent no completion notification is logged as an audit coverage gap and not waited on further. - Priority order when something forces a cut (a stall, a blocked spawn ladder, a dead transport): truth passes > coverage re-sweeps > systemic review > calibration. Elapsed time is never such a reason. The report and record ship from whatever completed, with the cut explicitly recorded. A partial audit that publishes beats a complete one that doesn't.
- No main-agent fetching while Phase 1/2 sub-agents run (guard #9). Main-agent exceptions after they return: bounded single-URL spot-checks in Phase 3/4 and the Phase 5.7/7 exceptions from
cti-run.md.
Phase 0, Preflight (sequential)
First, run prompts/cti-run.md Phase 0 step 0 verbatim (the network clock cross-check and the stale-clone sanity check), with the run-id suffix -audit instead of -intel: it sets STARTED, RUN_DATE and RUN_ID from a verified clock. That step exists because an audit fire (2026-08-24) booted on a clock eight days wrong and audited the wrong week. Then:
mkdir -p "work/${RUN_ID}" "runs/${RUN_DATE}"
echo "$STARTED" | tee "work/${RUN_ID}/main.started_at"
echo "$RUN_ID" | tee "work/${RUN_ID}/run_id"
: > "work/${RUN_ID}/url-liveness.tsv"
git fetch origin main
# Most recent prior audit record — the window anchor.
LAST_AUDIT=$(git ls-tree -r --name-only origin/main -- runs/ | grep -- '-audit\.md$' | sort | tail -1)
# Dedup context for any audit-recovered entries.
python3 tools/build_prior_coverage.py "$RUN_ID" 14
python3 tools/run_summary.py --out "work/${RUN_ID}/state-summary.json"
python3 tools/check_run.py --all > "work/${RUN_ID}/check-all.txt" 2>&1 || true
python3 tools/attack_data.py --check >> "work/${RUN_ID}/check-all.txt" 2>&1 || true
Then, in order:
- Window.
AUDIT_START= thestartedtimestamp ofLAST_AUDIT(git show origin/main:$LAST_AUDIT); no prior audit → 7 days back. A window of up to about 21 days is the usual reach (a guide): beyond it, audit the most recent part in depth and cover the rest as far as it is useful, recording what was left out. The window covers entries and run records withstarted/discovered_atin[AUDIT_START, now]. - Duplicate-audit guard. Gap since
LAST_AUDIT< 72 h → stop and reportduplicate-audit, unless this fire is an explicit interactive operator directive. - Carry-forward.
Readthe most recent audit report(s) underdocs/audits/and extract (a) open watch items, each gets a re-check duty this fire; (b) fixes shipped last audit, each gets an effectiveness check in Phase 3 (a fix that didn't change behavior is a finding). Operator closures are final (v3.27): a watch item the operator has explicitly closed (an operator-response addendum in an audit report, or a dated operator note in.claude/memory/) is CLOSED, gets no re-check duty, and is never re-opened by an audit; if genuinely new information on the underlying story surfaces later, it flows through the normal intel runs (as anupdaterecord on the existing entry), not through audit tracking. - Monthly calibration duty (recommendation 3 of the 2026-07-11 audit). Search
docs/audits/onorigin/mainfor a report dated in the current calendar month containing a## Priority calibrationheading. None found → this fire owns Phase 3b. This puts the review at monthly cadence whatever the audit schedule, self-healing across missed fires. - ATT&CK pin freshness duty. The Phase 0
attack_data.py --checkresult (already incheck-all.txt) is recorded in the run record's notes in one line. When it reports a newer upstream release, either perform the update this fire (python3 tools/attack_data.py --update && python3 tools/attack_data.py --selftest, confirmpython3 site/build.py+python3 site/test_build.pystay green, commitattack/enterprise-attack.jsonwith the printed change summary in the commit body) or, if the update fails its self-test or the build, surface it as an explicit operator item in the report. A stale pin is allowed to exist but never to go unmentioned (contract:attack/README.md).
5b. Cited-page pre-pass (v4.13). Once step 1 has fixed the window, run python3 tools/check_run.py "$RUN_ID" --page-checks-since <window start, YYYY-MM-DD> > "work/${RUN_ID}/page-checks.txt" 2>&1 (every evidence quote literal-searched on its cited page, every CVE-naming clause checked against its citation, bodies cached for the truth passes). Each quote-literal or citation-cve WARN in page-checks.txt goes to the truth pass that owns the entry as a named lead (the batch's spawn envelope lists it), and the pass decides it: a real non-verbatim quote or a mis-bound CVE is a correction, a page edited since publication (a vendor fixing its own typo) is noted and left. If no audit report yet records a --page-checks-since backfill, run it once from 2026-09-01 instead: on 2026-09-29 it found three published Securelist quotes that are not on the cited page (2026-09-18/moviereaper-torrent-supply-chain-solana-c2) and three European Court of Auditors quotes cited to the report's landing page rather than the document they come from (2026-09-23/eu-eca-cyber-incident-cooperation-report-nis2-gaps).
- Inventory. Enumerate the window's entry files and run records; an entry is in the window when its
discovered_atOR anyupdates[].atfalls inside it (an updated entry is re-verified as a whole, the new section included); partition entries into truth-pass batches of about 20 entries each (a guide, batch = one Phase 1 sub-agent: smaller for heavy entries, larger for light ones). Readentities/registry.yaml+site/taxonomy.yaml. Persist the plan towork/<run-id>/audit-plan.json; initialiseTodoWrite.
6b. Legacy re-verification batch (v4.19). 478 entries were migrated from the v2 daily briefs before claim verification and inline citations existed, and the audits' trailing windows never reach them: when the 2026-09-30 audit finally looked, it found a state attribution the cited authority never made, victims moved to the wrong exploitation wave, five CVEs marked exploited where the vendor reported one, and statistics absent from the cited report. Take the next batch from python3 tools/legacy_review.py --next 20 (oldest first, about 20 a guide) and add it to the truth-pass plan as its own batch or batches. Each entry is re-verified whole against fetched primaries; a wrong claim is corrected where it stands through a correction record (the entry is rewritten from its sources when most of it is unsupported), an unfolded legacy "UPDATE (originally covered …)" entry or a same-finding twin is folded into its survivor with tools/fold_entries.py (Phase 4), and a legacy entry that touches this fire also gains the rating and ATT&CK mapping its kind now requires. Mark the batch with python3 tools/legacy_review.py --done "$RUN_ID" <ids> and report the queue's remaining size; the queue is the standing repair order until it is empty.
Phase 1, Retrospective truth verification (sub-agents, parallel; no time limit)
One cold-reader pass per batch, spawned in a single message, subagent_type: cti-verification for every batch (the single verifier definition, pinned to the generic sonnet alias). These are retrospective audit passes, not Phase 5.7 iterations; record them under the run record's sub_agents: telemetry with the usual **Model:** / **Timestamps:** capture; Phase 5.7 still runs separately on this run's own output.
Spawn envelope: run id; the explicit entry-file list for the batch; the audit mandate ("assume nothing; fetch ground truth"); ledger path. Per entry, the pass fetches the primary sources and checks:
- every
cves[]id and CVSS against the per-CVE authority (vendor PSIRT / per-CVE advisory / discloser's per-vulnerability page, a roundup, even the same publisher's, is never sufficient; the v3.21 provenance rule, applied retroactively); - KEV/exploitation claims, version boundaries, affected/patched products, victim statements, attribution claims, and every
evidence[]quote verbatim against its cited URL; - every
techniques[]id active in the pinnedattack/enterprise-attack.json, mechanical ground truth, never memory (the 07-11 audit had a verifier mis-assert a revoked id was active); classificationconsistency with the sourcing actually present;closed_sourcestracing to a readable drop file; no IOCs; frontmatter⇔body agreement, on an updated entry this includes the changelog: each## <Type> — <at>section's claims against their cited sources, the record'ssummaryagainst its section, and the main analysis against the current frontmatter (anupdatethat movedcves[].statusto exploited while the analysis still says "no exploitation observed" is a self-contradiction, and acorrectionwhose fixed statement is still wrong is a factual error).
Return: a findings YAML to work/<run-id>/truth-<batch>.yaml, per entry {entry, verdict: clean|imprecision|factual-error, defect, ground_truth, ground_truth_url}. Clean entries are listed as clean (the audit's headline number is clean/total).
Phase 2, Independent coverage re-sweeps (sub-agents, parallel with Phase 1; no time limit)
Spawn subagent_type: cti-research, same spawn-envelope contract as the intel run's Phase 1 (window = the audit window as window_hours, source slice, dedup paths incl. entities/registry.yaml, ISO date), one per gap domain:
| Sub-agent | Domain |
|---|---|
| G1, Vulnerabilities & exploitation | KEV additions, national-CERT advisories, vendor PSIRTs, exploitation reporting across the window. |
| G2, Incidents & ransomware | Home-region / coverage-focus priority (the composed org profile in cti-run.md governs); victim disclosures, regulator filings, leak-site claims needing corroboration. |
| G3 (Threat research & APT | Vendor research-blog listing sweeps per publisher across the window dates) the discovery path that does not route through CVE/KEV channels and caused both misses found on 2026-07-11. |
Their job is to re-research the window as if for the first time and return everything that would clear PD-11, with discovery traces, the main agent diffs returns against the store (an item already covered, incl. as a changelog record on an existing entry, is a match, not a gap). G2 additionally re-checks every open watch item carried forward from Phase 0 step 3 (e.g. corroboration hunts on single-source leak-site claims).
Phase 3, Systemic & operational review (main context, local files)
Over the window's run records, work/ artifacts and state, no fetching except bounded single-URL recipe spot-checks after Phase 1/2 return:
- Telemetry: wall-clock durations (context only since runs have no time limit; >24 h is the stall class),
gap_hoursanomalies and overtaken-run races, verification iteration counts / residuals / dropped-entry rates, stalled-sub-agent abandonments. Zero-entry runs: read their notes, defensible quiet window or filter failure? Scheduler cadence is operator-owned and intentionally variable (operator decision 2026-07-18): the operator tunes the fire schedule at will (more fires, fewer fires, different slots) and the pipeline is built to work off the gap to the last run (PD-7). A changed cadence is therefore NOT a finding and NOT an operator recommendation; what the audit checks is the mechanism, that every fire's gap-derived window self-healed across the change with no coverage hole between runs. Only a self-heal failure or an actual coverage hole is a finding; the latency profile of a lower cadence may be stated as context in a miss's root cause, never flagged as a defect. (A gap with no run record at all on a previously firing schedule is still worth surfacing as possible scheduler outage, that is an availability signal, not a cadence judgment.) - Publish follow-through: any v3.14+ record with missing/stale
publish_status(the Phase 7 amendment that never landed). - Source health, reachability ≠ readability:
state/source_health.jsonnow carries a content verdict per source (v4.16), so the dark-but-green class is measured, not guessed: everyneeds-content-fixandstale-contentresult is a Phase 4 work item, fixed in the source record (url,rss_url,fetch_method,health_cmd,max_staleness_days,content_scope) with the evidence innotes, or reported as a genuine coverage gap. Beyond the verdicts, essential and tier-1 sources that are green yet contributed zero items across the whole window remain suspects: spot-check their recipe. Everywebfetch-onlysource (v4.18) is one the probe cannot judge, so the audit judges it: read it once withWebFetchand the outbound-links template, confirm the listing is current and on topic, and note the result in the source'snotes. Awebfetch-onlysource thatWebFetchno longer reads either gets a new recipe or is reported as a coverage gap. - Reader-pool health (last-resort transport; a dead pool is normal, not an incident; v3.33): run
python3 tools/fetch_source.py jina-usage. The jina reader is the fetch ladder's LAST rung, behind the trafilatura capture layer (extract <URL>), and the operator refills its keys sparsely and deliberately, so a dead or low pool is an expected steady state the pipeline must work through, never by itself an operator recommendation and NEVER a notification. What the audit checks instead: (a) that fires kept reading primaries through a dead pool via the direct rungs (a run record claiming a source was unreachable whenextractwould have read it is a finding); (b) that reader spend outsidefetch_method: jina-pinned hosts stayed near zero (rising spend means the trafilatura-first discipline is eroding). Note the pool state in the report's telemetry section as context only. - Discipline drift:
actions[]distribution against the v3.19 do-now bar (empty-is-normal shape restored?); priority mix of the window; classification codes present and plausible; behavior-kindtechniques[]density; changelog discipline on same-topic entries, a second entry where a record on the existing entry belonged, a record whose section recaps instead of stating the delta, a record whose frontmatter changes were not reflected in the main analysis, anactions[]list that accumulated instead of being replaced. Entity attachment hygiene (v4.15): open the entity pages of the window's new registry records and of every entity whose timeline grew by more than a handful of entries: an unrelated entry attached through a label that is ordinary vocabulary or another thing's name is fixed withambiguous_labelson the registry record, and a passing mention keyed inentities:is fixed by removing the key through aninternal: truerecord on that entry. - Gate & pin: the Phase 0
check_run.py --allandattack_data.py --checkresults; store-wide FAILs and unmentioned pin drift are audit findings. - Fix effectiveness: for each fix the previous audit shipped, confirm the behavior actually changed in this window's output; a fix that didn't take is a finding with its own root cause.
- Warning sweep to zero (v3.28): collect every WARN from the Phase 0
check-all.txtand everySELF-CHECK WARNINGfrom a freshpython3 site/build.py. Each one is a mandatory Phase 4 work item, fix the root cause (tool, prompt, renderer, source record, state) whenever a fix exists; only a warning whose subject is settled run-record history (a published run record's telemetry fact, an era-correct recorded waiver; entries are living records and are fixed through a changelog record instead) may instead be acknowledged instate/warning_acknowledgments.json. The audit does not publish untilcheck_run.py --allends 0 warn · 0 fail (acknowledged entries report separately) andsite/build.pyemits no self-check warnings. Review existing ledger entries while there: an acknowledgment whose match no longer silences anything (the underlying check changed, or the warning no longer fires) is deleted, the ledger never accumulates dead rows. - Model-generation baseline (v4.13, for any window whose run records change
model_id). The routines moved from Sonnet 5 / Opus 5 to Sonnet 5.5 / Opus 5.5 on 2026-09-29, with every role still atxhigheffort. Both models' guidance says effort levels were recalibrated and thatxhighshould be kept only where it measurably helps, so the audit that first sees a new generation measures it against the last week of the old one: fire duration, verifier iterations to publish and the confirmed-CLEAN rate, iteration-1 truth findings per published entry, research sub-agent duration and items returned,quote-literalWARNs caught before the first spawn, and this audit's own clean/total. Report the table in the telemetry section with a one-line recommendation per role (main agent, research, verifier): keepxhigh, or tryhigh. Effort is an operator setting (.claude/settings.jsoneffortLevel, each agent definition'seffort:), so the audit recommends and never changes it. - Rotational lookback (v4.13): for every research item G1/G3 recover from a standard or candidate source, check the fire that last swept that source before the item's discovery: an item dated inside that fire's
lookback_hoursfor the source is a sweep defect (the sub-agent did not honour the lookback), an item the source published while no fire swept it at all is a rotation defect. Count both, since together they say whether the lookback closed the gap the 2026-09-27 audit measured. - Stale truth sweep (v4.17). Three mechanical passes over the window, each hit a Phase 4 work item: (a)
python3 tools/kev_window_diff.py --since <window start> --run-id "$RUN_ID": every COVERED-STALE row is an entry whosecves[]still does not say exploited after a KEV listing (acorrectionrecord moving the status, the summary and the analysis), and every RANSOMWARE row is a KEV ransomware-use flag the carrying entry never mentions (anupdaterecord when the flag is new, animprovementwhen it predates the entry); (b) theexploitation-consistencyandentry-shapechecks ofcheck_run.pyrun over every window entry (call them on the window's entries, e.g.python3 -cimportingcheck_runwithenforce=True): each sentence denying exploitation of a CVE the entry marks exploited, and each criticalimmediate_actionthat names no fixed version, is fixed through a changelog record; (c) the supersession reading of every entry that received a changelog record in the window: a main-analysis, summary, title or action statement its own newest section disproves is acorrection. - Calibration alarm and backlog hygiene (v4.17). Report the window's
highshare: above about a quarter of entries, re-grade everyhighagainst thehighbar (cti-run.mdPhase 4) and correct the ones it fails (acorrectionrecord movingpriority,internal: truewhen nothing else changes). Openstate/coverage_backlog.md: every open row past its 14-day expiry, or held without a named resolution condition, is struck with the reason, and a row that clears PD-11 is published this fire. - Duplicate consolidation (v4.19). One finding has one entry (docs/pipeline.md § Entry lifecycle), but the store still holds duplicates the v4.0 migration could not see: legacy "UPDATE (originally covered …)" entries without
update_of, and same-day twins the brief migration split into a brief item and a deep dive (on 2026-09-30, 48 of 54 CVE-sharing entry pairs carried no link at all, and two CVE-2026-50751 twins disagreed about which CVE was exploited). For every pair of entries sharing a CVE with neither referencing the other, decide: distinct findings get areferences[]link through aninternal: truerecord; duplicates are folded. To fold, bring the survivor to the current verified state with one record for this fire (carry over, with citations, the material facts only the duplicate held; leave unverified legacy text behind), then runpython3 tools/fold_entries.py --run "$RUN_ID" --into <survivor> <duplicate>…, which adds the ids to that record'smerged_from, re-pointsreferences[]and registryrelations[].source, and deletes the duplicate files (the build keeps their permalinks as redirects).check_run.pyaccepts a deleted entry only when this fire's record on a survivor names it.
Phase 3b, Monthly priority-calibration review (only when Phase 0 step 4 assigned it)
Institutionalizes recommendation 3 of the 2026-07-11 audit: "the high share (37 %) is stable but on the generous side; one pass a month over the priority distribution against F-category drift keeps the notification channel honest."
- Compute the
prioritydistribution store-wide and over the trailing calendar month (grepentries/**frontmatter; the 2026-07 store baseline: 15 critical / 335 high / 564 notable / 1 routine,high≈ 37 %). - Collect the priority-calibration findings (F16) from the month's run records'
verification.iterations[].findings, the drift signal the verifier loop already produces. - Judge against the bars, which are fixed references the review never tunes:
criticalclears stop-reading-and-act-now (rare because the bar is extreme, never because a number caps it; a long critical-free streak under real exploitation pressure is itself a signal to check);highgenuinely TL;DR-worthy (everyhighheadline tops the rendered 24 h window, the generous direction burns the notification channel);notablenot absorbing items that plainly clear a higher bar (under-alerting is the same defect in reverse). - Ship the outcome as a
## Priority calibrationsection in the audit report, distribution table, month-over-month movement, F16 summary, verdict. Only when drift is confirmed against concrete mis-prioritized entries does a calibration edit toprompts/cti-run.mdpriority guidance ship (versioning rule applies); a judgment call the audit shouldn't own alone becomes an operator recommendation.
Phase 4, Root-cause & fix (main context)
Every confirmed defect is root-caused to the specific mechanism that let it through (which phase, rule, tool, or missing duty), then fixed by class:
- Factual errors in a published entry: a claim its source does not support, an inverted mechanism, a wrong CVE id / CVSS / version / date / ATT&CK id / registry key, a mis-attributed quote: a
correctionrecord on that entry (prompts/cti-run.mdPhase 4 § Updating an existing entry). Fix the wrong statement where it stands (frontmatter and body, every touched field named in the record'sfields) so the entry never asserts something known to be false, and let the## Correction — <at>section state what was wrong, what is right, and the ground-truth source. Re-sync affected state (state/cves_seen.json), and disclose each correction in the report and the run record (updated_entry_ids[]). The former immutability-exception ledger is retired: the changelog is the ledger, and it lives with the entry. - Precision or depth the entry lacks without being wrong: a missing second source, a technique the body describes but
techniques[]omits, a**Triage:**line the mechanism supports, a version stated imprecisely: animprovementrecord on that entry, same mechanics. Never manufacture an improvement to look productive, a clean entry gets no record. - Missed coverage that still clears PD-11 today: compose and publish the recovered entry through the full normal gates (Phase 4 composition rules, dedup,
check_run.py, verifier loop), with a provenance line naming this audit. A recovered item that is a development on a covered finding is anupdaterecord on that entry, not a new one. Borderline items that correctly fail PD-11 are documented as correctly-droppable, that documentation is what makes the completeness judgment auditable. - Systemic causes: fix prompts / tools / sources / agent definitions / memory under § META authority; the versioning rule applies in full (banner bump + CHANGELOG entry in the same commit). What exceeds authority (scheduler config, org-profile values, hosting) becomes a numbered operator recommendation.
- Single-source claims touching the constituency that fail corroboration: not published, recorded as watch items with what would resolve them, the next audit re-checks.
- Warnings on settled history (v3.28): a WARN whose subject cannot change without falsifying the record; a published run record's stall-length
duration_seconds, an era-correct recorded confirmation waiver (run records stay immutable; entries do not, so a WARN on an entry is fixed through a changelog record, never acknowledged), is acknowledged instate/warning_acknowledgments.json:checklabel,matchpinned to the specific run id/subject,reason,acknowledged_at. This is a reviewed audit decision disclosed in the report (the ledger diff is part of the audit commit); NEVER acknowledge a warning that a code/prompt/state/renderer fix could clear instead, and a run never self-acknowledges its own fresh warnings.
Phase 5; Audit report (skeleton-then-Edit)
Write docs/audits/<RUN_DATE>-quality-audit.md (earlier reports carry the -weekly-quality-audit suffix from the retired naming; read them with either name), mirroring the 2026-07-11 report's structure:
- Header + method: mandate, window, run-record link, sub-agents spawned, entries checked / primary URLs fetched.
- Verdict: lead with
clean/totalentries verified and the one-paragraph honest state of the pipeline. - Findings, false or erroneous published intelligence (table: entry, defect, ground truth, fix, the fix names the
correction/improvementrecord appended) + root causes. Minor imprecisions documented even when no record is warranted. - Findings, missing or incomplete coverage: recovered items, correctly-droppable borderlines with reasons, resolved false alarms.
- Findings, systemic / operational (incl. fix-effectiveness results from Phase 3 item 7).
- Priority calibration (monthly fires only, Phase 3b output).
- Fixes shipped in this commit: including the warning sweep's outcome: what was fixed, and every new/removed
state/warning_acknowledgments.jsonentry with its reason. - Recommendations (operator decisions, not shipped), numbered, carried forward until adopted or explicitly retired.
- Watch items: carried-forward status + new items, each with its resolution condition.
Empty sections state "none found"; the absence is a result. Then write the run record: telemetry frontmatter (audit passes under sub_agents:), verification & coverage notes summarizing the report, entries_published = recovered (new) entries only, entries_updated + updated_entry_ids[] = the entries that received a correction / improvement / update record this fire.
Phases 5.5 → 7, State, gate, verify, publish
Read prompts/cti-run.md now (Phases 5, 5.5, 5.7, 6, 7) and execute verbatim with this run's RUN_ID: state lifecycle only where recovered entries touched it; python3 tools/check_run.py "$RUN_ID" --pre-verify exit 0 before the first verifier spawn. The Phase 5.7 loop (≥1 iteration ALWAYS, ≥2 for a CLEAN publish per the double-CLEAN gate, two consecutive CLEAN passes of the single verifier definition, cap 8, fail-open) is scoped to: this run's recovered entries + every entry it appended a changelog record to (read whole; new section and changed fields verified against the cited sources) + the run record + the audit report; the report makes checkable claims about published files, records and state, and a cold reader must confirm they hold on disk (the 07-11 audit's own iteration 1 caught exactly this class: a claimed repair log entry that wasn't there). Then stage specifics (incl. docs/audits/, work/<run-id>/, .claude/memory/ when touched), commit, sync origin/main, push with retry, Phase 7 publish polling, report publish: from the actual poll.
Quality gates (self-check)
- [ ] Every window entry was covered by exactly one truth pass, or its batch's abandonment is logged as an audit coverage gap.
- [ ] Both audit halves ran: truth passes (soundness) AND independent re-sweeps (completeness), or the cut is recorded with its cause (a stall, a blocked ladder, a dead transport, never elapsed time).
- [ ] Every confirmed defect carries a root cause and either a shipped fix or a numbered operator recommendation, no orphan findings.
- [ ] Every change to a published entry is a
correction/improvement/updatechangelog record (with its section when reader-facing,internal: trueand no section when metadata-only;updated_atbumped only bytype: update) listed inupdated_entry_ids[]and disclosed in the report; no silent edit, no second entry for a covered finding. - [ ] Previous audit's watch items re-checked and its fixes effectiveness-checked.
- [ ] A legacy re-verification batch was checked and marked in
state/legacy_review.json(v4.19), and every duplicate found was folded or linked. - [ ] Monthly duty honored: at most one
## Priority calibrationsection per calendar month across audit reports, and at least one unless no audit fired that month. - [ ] Recovered entries passed the full normal gates (dedup,
check_run.py, verifier loop), audit provenance never lowers a bar. - [ ] Zero-warning sweep complete:
python3 tools/check_run.py --allends 0 warn · 0 fail (acknowledged excepted) andpython3 site/build.pyemits no self-check warnings; every ledger change is disclosed in the report. - [ ] No manufactured findings; clean components reported clean.
- [ ] All shared gates from
prompts/cti-run.md§ Quality gates hold (gate exit 0, ≥1 verifier iteration, run record committed, publish verified).
Output
run: runs/YYYY-MM-DD/<run-id>.md
audit: docs/audits/YYYY-MM-DD-quality-audit.md
window: <start> → <end> · entries audited: N (clean: N) · runs reviewed: N
findings: erroneous: N · coverage gaps: N (recovered: N) · systemic: N · calibration: run | not-due
fixes shipped: N · entries corrected/improved (changelog records): N · watch items: N open
warnings: 0 open (N acknowledged) · build self-check: clean
commit: <short SHA> · push: ok (feature branch) | failed (<reason>) · publish: ok | main-only | pending (<reason>)
META, self-evolution authority
Same authority and process as prompts/cti-run.md § META. All hard invariants apply, plus:
- A-INV-1: the audit changes a published entry only through a dated changelog record (
correction/improvement/update; section when reader-facing,internal: truewithout one when metadata-only;updated_atmoves only ontype: update), never a silent edit, never a second entry for the same finding. - A-INV-2: findings are never manufactured; a clean week is the success outcome and is reported as such.
- A-INV-3: the audit report and run record ship even when the audit finds nothing or a stall cuts it short; a silent audit is an operational failure.
- A-INV-4: the audit ships fixes under the versioning rule and never weakens a hard invariant, concerns about an invariant become recommendations, not edits.
- A-INV-5: the audit never blocks or races the scheduled intel runs; it works on published history plus its own artifacts, on its own feature branch (per-run paths make a concurrent intel fire safe for new entries; an entry both fires update in the same window is a file conflict the sync surfaces, keep both records in
atorder, never drop one).