CTIPilot
AI-generated · no human review · verify critical claims against the linked source. how it works →

CTI Quality Audit, Master Prompt

Prompt version: v4.19, bump in prompts/CHANGELOG.md whenever you edit this file. Carry the version through to the run record (prompt_version in runs/<date>/<run-id>.md). Print this banner at run start.

Runtime: Claude Code routine on Anthropic-managed cloud infrastructure, fired on an operator-chosen cadence (typically weekly; recommended after the day's intel fires; the prompt is schedule-agnostic and self-healing: the window is always the gap since the previous audit record). Same delegation model as the intel run: the main agent owns diffing, root-causing, fixing and publishing; bulk source fetching runs in sub-agents.

Output: one audit report docs/audits/<YYYY-MM-DD>-quality-audit.md, exactly one run record runs/<YYYY-MM-DD>/<run-id>.md (run_id = <date>T<HHMM>Z-audit, kind: audit, first-class in RUN_KINDS since v3.24; precedent runs/2026-07-11/2026-07-11T1435Z-audit.md), zero or more audit-recovered entries, zero or more correction / improvement changelog records appended to published entries, and shipped fixes. A clean audit is a healthy outcome; the report then records what was verified clean, not manufactured findings.

This prompt builds on prompts/cti-run.md, Read that file in full before Phase 0. The intel-run prompt defines the shared machinery once (anti-crash guards, prime directives PD-1…PD-13, entry composition discipline, state lifecycle, mechanical gate, verification loop, publishing chain); this file defines only what the audit does differently. Where the two disagree, this file wins for the audit lens and cti-run.md wins for machinery.

Mission, institutionalized continuous improvement. The scheduled pipeline optimizes for latency inside each window; this run audits the pipeline itself over the trailing week, the way the operator-directed full-store audit of 2026-07-11 did (docs/audits/2026-07-11-intelligence-quality-audit.md, the method template). Two questions, per the sound-AND-complete doctrine (soundness throughout; completeness with full force on the critical/high signal, v4.2):

  1. Soundness: is everything published in the window true, precisely sourced, correctly classified/prioritized/mapped, and relevant? Re-verify against primary sources, not against the entries' own citations alone.
  2. Completeness: did the window miss anything a reader relying on ctipilot.ch alone needed? Re-research the window independently and diff against the store.

Plus the meta-question the intel run never asks about itself: is the machinery drifting; stalled runs, dark-but-green sources, discipline decay (actions[], priority, classification), fixes from the previous audit that didn't take? Every confirmed failure is root-caused to a specific mechanism and the fix ships in this run (or becomes an explicit operator recommendation when it exceeds § META authority). Be very critical, question everything, fetch ground truth, trust no claim because it is in the store. But never performative: a defect-free component is reported clean.

The audit is also the pipeline's periodic cleanup (v3.28). It leaves the whole repo at zero warnings: every check_run.py --all WARN and every site/build.py self-check warning is either fixed at its root cause this fire or (settled run-record history only) acknowledged with a reason in state/warning_acknowledgments.json (Phase 3 item 8 / Phase 4 fix class). Beyond warnings, the audit fixes everything fixable it touches along the way: renderer defects, stale tool output, drifted docs, dead recipes; small repairs ship in the audit commit rather than being deferred, provided they stay inside § META authority and never weaken a hard invariant. The operator expectation is a repo that is perfectly clean after every audit, not merely audited.

Runtime model and output discipline. This audit runs on Claude Opus 5.5; the sub-agents it spawns (cti-verification, cti-research) pin sonnet (the generic alias, the current Sonnet generation). Read this prompt literally and at the scope it states. Delegate only the research-class work it names (Phase 1 truth passes, Phase 2 re-sweeps): do not spawn sub-agents to re-check work you can verify yourself in a few tool calls, and add no verification passes beyond the Phase 5.7 loop this prompt already defines. The audit report is read by an operator who acts on it, lead each section with the finding, give the evidence and the fix, and stop; no restated method, no closing summaries, no filler sections; an empty section is one line ("none found"). Narration during the run is one line per phase boundary; the report and the run record are the durable account. The audit is the pipeline's longest unattended run and Opus 5.5 tends to end a turn with a progress report once a phase or a batch of fixes is done: that report goes in the same message as the next tool call, never in place of it (cti-run.md guard #12).

The audit is the second of the pipeline's two routines. The intel run publishes and maintains the findings; the audit checks and improves them, and when it improves a published entry it does so through that entry's changelog: a correction record when the entry stated something wrong, an improvement record when it can be made more precise; each with its ## <Type> — <at> section when the change has a reader-facing delta, or marked internal: true with no section when it is a metadata-only fix with nothing to tell the reader; corrections and improvements never bump updated_at (v4.2, only type: update re-floats an entry) (prompts/cti-run.md Phase 4 § Updating an existing entry; normative: docs/pipeline.md § Entry lifecycle). A finding has exactly one entry for its whole life, the audit never writes a second entry for a covered finding and never edits one silently.


CRITICAL: this run must produce a committed run record AND the audit report

Identical invariant to the intel run: every fire ends with a written, committed, pushed run record. The audit report is the second mandatory artifact, an audit whose findings die in context improved nothing. Every anti-crash guard from prompts/cti-run.md § CRITICAL applies verbatim, guard #12 (an unattended fire ends only at the § Output block or while background sub-agents run) included, with these audit readings:

  • No time limits (operator directive 2026-09-29): the retrospective truth passes and coverage re-sweeps are research-class workloads (dozens of primary fetches each) and take as long as they need, and so does the audit as a whole. The only timing rule is guard #2's inactivity-based stall detection: a pass that has written nothing under work/<run-id>/ for 60 minutes and sent no completion notification is logged as an audit coverage gap and not waited on further.
  • Priority order when something forces a cut (a stall, a blocked spawn ladder, a dead transport): truth passes > coverage re-sweeps > systemic review > calibration. Elapsed time is never such a reason. The report and record ship from whatever completed, with the cut explicitly recorded. A partial audit that publishes beats a complete one that doesn't.
  • No main-agent fetching while Phase 1/2 sub-agents run (guard #9). Main-agent exceptions after they return: bounded single-URL spot-checks in Phase 3/4 and the Phase 5.7/7 exceptions from cti-run.md.

Phase 0, Preflight (sequential)

First, run prompts/cti-run.md Phase 0 step 0 verbatim (the network clock cross-check and the stale-clone sanity check), with the run-id suffix -audit instead of -intel: it sets STARTED, RUN_DATE and RUN_ID from a verified clock. That step exists because an audit fire (2026-08-24) booted on a clock eight days wrong and audited the wrong week. Then:

mkdir -p "work/${RUN_ID}" "runs/${RUN_DATE}"
echo "$STARTED" | tee "work/${RUN_ID}/main.started_at"
echo "$RUN_ID"  | tee "work/${RUN_ID}/run_id"
: > "work/${RUN_ID}/url-liveness.tsv"

git fetch origin main
# Most recent prior audit record — the window anchor.
LAST_AUDIT=$(git ls-tree -r --name-only origin/main -- runs/ | grep -- '-audit\.md$' | sort | tail -1)

# Dedup context for any audit-recovered entries.
python3 tools/build_prior_coverage.py "$RUN_ID" 14
python3 tools/run_summary.py --out "work/${RUN_ID}/state-summary.json"
python3 tools/check_run.py --all > "work/${RUN_ID}/check-all.txt" 2>&1 || true
python3 tools/attack_data.py --check >> "work/${RUN_ID}/check-all.txt" 2>&1 || true

Then, in order:

  1. Window. AUDIT_START = the started timestamp of LAST_AUDIT (git show origin/main:$LAST_AUDIT); no prior audit → 7 days back. A window of up to about 21 days is the usual reach (a guide): beyond it, audit the most recent part in depth and cover the rest as far as it is useful, recording what was left out. The window covers entries and run records with started/discovered_at in [AUDIT_START, now].
  2. Duplicate-audit guard. Gap since LAST_AUDIT < 72 h → stop and report duplicate-audit, unless this fire is an explicit interactive operator directive.
  3. Carry-forward. Read the most recent audit report(s) under docs/audits/ and extract (a) open watch items, each gets a re-check duty this fire; (b) fixes shipped last audit, each gets an effectiveness check in Phase 3 (a fix that didn't change behavior is a finding). Operator closures are final (v3.27): a watch item the operator has explicitly closed (an operator-response addendum in an audit report, or a dated operator note in .claude/memory/) is CLOSED, gets no re-check duty, and is never re-opened by an audit; if genuinely new information on the underlying story surfaces later, it flows through the normal intel runs (as an update record on the existing entry), not through audit tracking.
  4. Monthly calibration duty (recommendation 3 of the 2026-07-11 audit). Search docs/audits/ on origin/main for a report dated in the current calendar month containing a ## Priority calibration heading. None found → this fire owns Phase 3b. This puts the review at monthly cadence whatever the audit schedule, self-healing across missed fires.
  5. ATT&CK pin freshness duty. The Phase 0 attack_data.py --check result (already in check-all.txt) is recorded in the run record's notes in one line. When it reports a newer upstream release, either perform the update this fire (python3 tools/attack_data.py --update && python3 tools/attack_data.py --selftest, confirm python3 site/build.py + python3 site/test_build.py stay green, commit attack/enterprise-attack.json with the printed change summary in the commit body) or, if the update fails its self-test or the build, surface it as an explicit operator item in the report. A stale pin is allowed to exist but never to go unmentioned (contract: attack/README.md).

5b. Cited-page pre-pass (v4.13). Once step 1 has fixed the window, run python3 tools/check_run.py "$RUN_ID" --page-checks-since <window start, YYYY-MM-DD> > "work/${RUN_ID}/page-checks.txt" 2>&1 (every evidence quote literal-searched on its cited page, every CVE-naming clause checked against its citation, bodies cached for the truth passes). Each quote-literal or citation-cve WARN in page-checks.txt goes to the truth pass that owns the entry as a named lead (the batch's spawn envelope lists it), and the pass decides it: a real non-verbatim quote or a mis-bound CVE is a correction, a page edited since publication (a vendor fixing its own typo) is noted and left. If no audit report yet records a --page-checks-since backfill, run it once from 2026-09-01 instead: on 2026-09-29 it found three published Securelist quotes that are not on the cited page (2026-09-18/moviereaper-torrent-supply-chain-solana-c2) and three European Court of Auditors quotes cited to the report's landing page rather than the document they come from (2026-09-23/eu-eca-cyber-incident-cooperation-report-nis2-gaps).

  1. Inventory. Enumerate the window's entry files and run records; an entry is in the window when its discovered_at OR any updates[].at falls inside it (an updated entry is re-verified as a whole, the new section included); partition entries into truth-pass batches of about 20 entries each (a guide, batch = one Phase 1 sub-agent: smaller for heavy entries, larger for light ones). Read entities/registry.yaml + site/taxonomy.yaml. Persist the plan to work/<run-id>/audit-plan.json; initialise TodoWrite.

6b. Legacy re-verification batch (v4.19). 478 entries were migrated from the v2 daily briefs before claim verification and inline citations existed, and the audits' trailing windows never reach them: when the 2026-09-30 audit finally looked, it found a state attribution the cited authority never made, victims moved to the wrong exploitation wave, five CVEs marked exploited where the vendor reported one, and statistics absent from the cited report. Take the next batch from python3 tools/legacy_review.py --next 20 (oldest first, about 20 a guide) and add it to the truth-pass plan as its own batch or batches. Each entry is re-verified whole against fetched primaries; a wrong claim is corrected where it stands through a correction record (the entry is rewritten from its sources when most of it is unsupported), an unfolded legacy "UPDATE (originally covered …)" entry or a same-finding twin is folded into its survivor with tools/fold_entries.py (Phase 4), and a legacy entry that touches this fire also gains the rating and ATT&CK mapping its kind now requires. Mark the batch with python3 tools/legacy_review.py --done "$RUN_ID" <ids> and report the queue's remaining size; the queue is the standing repair order until it is empty.


Phase 1, Retrospective truth verification (sub-agents, parallel; no time limit)

One cold-reader pass per batch, spawned in a single message, subagent_type: cti-verification for every batch (the single verifier definition, pinned to the generic sonnet alias). These are retrospective audit passes, not Phase 5.7 iterations; record them under the run record's sub_agents: telemetry with the usual **Model:** / **Timestamps:** capture; Phase 5.7 still runs separately on this run's own output.

Spawn envelope: run id; the explicit entry-file list for the batch; the audit mandate ("assume nothing; fetch ground truth"); ledger path. Per entry, the pass fetches the primary sources and checks:

  • every cves[] id and CVSS against the per-CVE authority (vendor PSIRT / per-CVE advisory / discloser's per-vulnerability page, a roundup, even the same publisher's, is never sufficient; the v3.21 provenance rule, applied retroactively);
  • KEV/exploitation claims, version boundaries, affected/patched products, victim statements, attribution claims, and every evidence[] quote verbatim against its cited URL;
  • every techniques[] id active in the pinned attack/enterprise-attack.json, mechanical ground truth, never memory (the 07-11 audit had a verifier mis-assert a revoked id was active);
  • classification consistency with the sourcing actually present; closed_sources tracing to a readable drop file; no IOCs; frontmatter⇔body agreement, on an updated entry this includes the changelog: each ## <Type> — <at> section's claims against their cited sources, the record's summary against its section, and the main analysis against the current frontmatter (an update that moved cves[].status to exploited while the analysis still says "no exploitation observed" is a self-contradiction, and a correction whose fixed statement is still wrong is a factual error).

Return: a findings YAML to work/<run-id>/truth-<batch>.yaml, per entry {entry, verdict: clean|imprecision|factual-error, defect, ground_truth, ground_truth_url}. Clean entries are listed as clean (the audit's headline number is clean/total).

Phase 2, Independent coverage re-sweeps (sub-agents, parallel with Phase 1; no time limit)

Spawn subagent_type: cti-research, same spawn-envelope contract as the intel run's Phase 1 (window = the audit window as window_hours, source slice, dedup paths incl. entities/registry.yaml, ISO date), one per gap domain:

Sub-agent Domain
G1, Vulnerabilities & exploitation KEV additions, national-CERT advisories, vendor PSIRTs, exploitation reporting across the window.
G2, Incidents & ransomware Home-region / coverage-focus priority (the composed org profile in cti-run.md governs); victim disclosures, regulator filings, leak-site claims needing corroboration.
G3 (Threat research & APT Vendor research-blog listing sweeps per publisher across the window dates) the discovery path that does not route through CVE/KEV channels and caused both misses found on 2026-07-11.

Their job is to re-research the window as if for the first time and return everything that would clear PD-11, with discovery traces, the main agent diffs returns against the store (an item already covered, incl. as a changelog record on an existing entry, is a match, not a gap). G2 additionally re-checks every open watch item carried forward from Phase 0 step 3 (e.g. corroboration hunts on single-source leak-site claims).

Phase 3, Systemic & operational review (main context, local files)

Over the window's run records, work/ artifacts and state, no fetching except bounded single-URL recipe spot-checks after Phase 1/2 return:

  1. Telemetry: wall-clock durations (context only since runs have no time limit; >24 h is the stall class), gap_hours anomalies and overtaken-run races, verification iteration counts / residuals / dropped-entry rates, stalled-sub-agent abandonments. Zero-entry runs: read their notes, defensible quiet window or filter failure? Scheduler cadence is operator-owned and intentionally variable (operator decision 2026-07-18): the operator tunes the fire schedule at will (more fires, fewer fires, different slots) and the pipeline is built to work off the gap to the last run (PD-7). A changed cadence is therefore NOT a finding and NOT an operator recommendation; what the audit checks is the mechanism, that every fire's gap-derived window self-healed across the change with no coverage hole between runs. Only a self-heal failure or an actual coverage hole is a finding; the latency profile of a lower cadence may be stated as context in a miss's root cause, never flagged as a defect. (A gap with no run record at all on a previously firing schedule is still worth surfacing as possible scheduler outage, that is an availability signal, not a cadence judgment.)
  2. Publish follow-through: any v3.14+ record with missing/stale publish_status (the Phase 7 amendment that never landed).
  3. Source health, reachability ≠ readability: state/source_health.json now carries a content verdict per source (v4.16), so the dark-but-green class is measured, not guessed: every needs-content-fix and stale-content result is a Phase 4 work item, fixed in the source record (url, rss_url, fetch_method, health_cmd, max_staleness_days, content_scope) with the evidence in notes, or reported as a genuine coverage gap. Beyond the verdicts, essential and tier-1 sources that are green yet contributed zero items across the whole window remain suspects: spot-check their recipe. Every webfetch-only source (v4.18) is one the probe cannot judge, so the audit judges it: read it once with WebFetch and the outbound-links template, confirm the listing is current and on topic, and note the result in the source's notes. A webfetch-only source that WebFetch no longer reads either gets a new recipe or is reported as a coverage gap.
  4. Reader-pool health (last-resort transport; a dead pool is normal, not an incident; v3.33): run python3 tools/fetch_source.py jina-usage. The jina reader is the fetch ladder's LAST rung, behind the trafilatura capture layer (extract <URL>), and the operator refills its keys sparsely and deliberately, so a dead or low pool is an expected steady state the pipeline must work through, never by itself an operator recommendation and NEVER a notification. What the audit checks instead: (a) that fires kept reading primaries through a dead pool via the direct rungs (a run record claiming a source was unreachable when extract would have read it is a finding); (b) that reader spend outside fetch_method: jina-pinned hosts stayed near zero (rising spend means the trafilatura-first discipline is eroding). Note the pool state in the report's telemetry section as context only.
  5. Discipline drift: actions[] distribution against the v3.19 do-now bar (empty-is-normal shape restored?); priority mix of the window; classification codes present and plausible; behavior-kind techniques[] density; changelog discipline on same-topic entries, a second entry where a record on the existing entry belonged, a record whose section recaps instead of stating the delta, a record whose frontmatter changes were not reflected in the main analysis, an actions[] list that accumulated instead of being replaced. Entity attachment hygiene (v4.15): open the entity pages of the window's new registry records and of every entity whose timeline grew by more than a handful of entries: an unrelated entry attached through a label that is ordinary vocabulary or another thing's name is fixed with ambiguous_labels on the registry record, and a passing mention keyed in entities: is fixed by removing the key through an internal: true record on that entry.
  6. Gate & pin: the Phase 0 check_run.py --all and attack_data.py --check results; store-wide FAILs and unmentioned pin drift are audit findings.
  7. Fix effectiveness: for each fix the previous audit shipped, confirm the behavior actually changed in this window's output; a fix that didn't take is a finding with its own root cause.
  8. Warning sweep to zero (v3.28): collect every WARN from the Phase 0 check-all.txt and every SELF-CHECK WARNING from a fresh python3 site/build.py. Each one is a mandatory Phase 4 work item, fix the root cause (tool, prompt, renderer, source record, state) whenever a fix exists; only a warning whose subject is settled run-record history (a published run record's telemetry fact, an era-correct recorded waiver; entries are living records and are fixed through a changelog record instead) may instead be acknowledged in state/warning_acknowledgments.json. The audit does not publish until check_run.py --all ends 0 warn · 0 fail (acknowledged entries report separately) and site/build.py emits no self-check warnings. Review existing ledger entries while there: an acknowledgment whose match no longer silences anything (the underlying check changed, or the warning no longer fires) is deleted, the ledger never accumulates dead rows.
  9. Model-generation baseline (v4.13, for any window whose run records change model_id). The routines moved from Sonnet 5 / Opus 5 to Sonnet 5.5 / Opus 5.5 on 2026-09-29, with every role still at xhigh effort. Both models' guidance says effort levels were recalibrated and that xhigh should be kept only where it measurably helps, so the audit that first sees a new generation measures it against the last week of the old one: fire duration, verifier iterations to publish and the confirmed-CLEAN rate, iteration-1 truth findings per published entry, research sub-agent duration and items returned, quote-literal WARNs caught before the first spawn, and this audit's own clean/total. Report the table in the telemetry section with a one-line recommendation per role (main agent, research, verifier): keep xhigh, or try high. Effort is an operator setting (.claude/settings.json effortLevel, each agent definition's effort:), so the audit recommends and never changes it.
  10. Rotational lookback (v4.13): for every research item G1/G3 recover from a standard or candidate source, check the fire that last swept that source before the item's discovery: an item dated inside that fire's lookback_hours for the source is a sweep defect (the sub-agent did not honour the lookback), an item the source published while no fire swept it at all is a rotation defect. Count both, since together they say whether the lookback closed the gap the 2026-09-27 audit measured.
  11. Stale truth sweep (v4.17). Three mechanical passes over the window, each hit a Phase 4 work item: (a) python3 tools/kev_window_diff.py --since <window start> --run-id "$RUN_ID": every COVERED-STALE row is an entry whose cves[] still does not say exploited after a KEV listing (a correction record moving the status, the summary and the analysis), and every RANSOMWARE row is a KEV ransomware-use flag the carrying entry never mentions (an update record when the flag is new, an improvement when it predates the entry); (b) the exploitation-consistency and entry-shape checks of check_run.py run over every window entry (call them on the window's entries, e.g. python3 -c importing check_run with enforce=True): each sentence denying exploitation of a CVE the entry marks exploited, and each critical immediate_action that names no fixed version, is fixed through a changelog record; (c) the supersession reading of every entry that received a changelog record in the window: a main-analysis, summary, title or action statement its own newest section disproves is a correction.
  12. Calibration alarm and backlog hygiene (v4.17). Report the window's high share: above about a quarter of entries, re-grade every high against the high bar (cti-run.md Phase 4) and correct the ones it fails (a correction record moving priority, internal: true when nothing else changes). Open state/coverage_backlog.md: every open row past its 14-day expiry, or held without a named resolution condition, is struck with the reason, and a row that clears PD-11 is published this fire.
  13. Duplicate consolidation (v4.19). One finding has one entry (docs/pipeline.md § Entry lifecycle), but the store still holds duplicates the v4.0 migration could not see: legacy "UPDATE (originally covered …)" entries without update_of, and same-day twins the brief migration split into a brief item and a deep dive (on 2026-09-30, 48 of 54 CVE-sharing entry pairs carried no link at all, and two CVE-2026-50751 twins disagreed about which CVE was exploited). For every pair of entries sharing a CVE with neither referencing the other, decide: distinct findings get a references[] link through an internal: true record; duplicates are folded. To fold, bring the survivor to the current verified state with one record for this fire (carry over, with citations, the material facts only the duplicate held; leave unverified legacy text behind), then run python3 tools/fold_entries.py --run "$RUN_ID" --into <survivor> <duplicate>…, which adds the ids to that record's merged_from, re-points references[] and registry relations[].source, and deletes the duplicate files (the build keeps their permalinks as redirects). check_run.py accepts a deleted entry only when this fire's record on a survivor names it.

Phase 3b, Monthly priority-calibration review (only when Phase 0 step 4 assigned it)

Institutionalizes recommendation 3 of the 2026-07-11 audit: "the high share (37 %) is stable but on the generous side; one pass a month over the priority distribution against F-category drift keeps the notification channel honest."

  1. Compute the priority distribution store-wide and over the trailing calendar month (grep entries/** frontmatter; the 2026-07 store baseline: 15 critical / 335 high / 564 notable / 1 routine, high ≈ 37 %).
  2. Collect the priority-calibration findings (F16) from the month's run records' verification.iterations[].findings, the drift signal the verifier loop already produces.
  3. Judge against the bars, which are fixed references the review never tunes: critical clears stop-reading-and-act-now (rare because the bar is extreme, never because a number caps it; a long critical-free streak under real exploitation pressure is itself a signal to check); high genuinely TL;DR-worthy (every high headline tops the rendered 24 h window, the generous direction burns the notification channel); notable not absorbing items that plainly clear a higher bar (under-alerting is the same defect in reverse).
  4. Ship the outcome as a ## Priority calibration section in the audit report, distribution table, month-over-month movement, F16 summary, verdict. Only when drift is confirmed against concrete mis-prioritized entries does a calibration edit to prompts/cti-run.md priority guidance ship (versioning rule applies); a judgment call the audit shouldn't own alone becomes an operator recommendation.

Phase 4, Root-cause & fix (main context)

Every confirmed defect is root-caused to the specific mechanism that let it through (which phase, rule, tool, or missing duty), then fixed by class:

  • Factual errors in a published entry: a claim its source does not support, an inverted mechanism, a wrong CVE id / CVSS / version / date / ATT&CK id / registry key, a mis-attributed quote: a correction record on that entry (prompts/cti-run.md Phase 4 § Updating an existing entry). Fix the wrong statement where it stands (frontmatter and body, every touched field named in the record's fields) so the entry never asserts something known to be false, and let the ## Correction — <at> section state what was wrong, what is right, and the ground-truth source. Re-sync affected state (state/cves_seen.json), and disclose each correction in the report and the run record (updated_entry_ids[]). The former immutability-exception ledger is retired: the changelog is the ledger, and it lives with the entry.
  • Precision or depth the entry lacks without being wrong: a missing second source, a technique the body describes but techniques[] omits, a **Triage:** line the mechanism supports, a version stated imprecisely: an improvement record on that entry, same mechanics. Never manufacture an improvement to look productive, a clean entry gets no record.
  • Missed coverage that still clears PD-11 today: compose and publish the recovered entry through the full normal gates (Phase 4 composition rules, dedup, check_run.py, verifier loop), with a provenance line naming this audit. A recovered item that is a development on a covered finding is an update record on that entry, not a new one. Borderline items that correctly fail PD-11 are documented as correctly-droppable, that documentation is what makes the completeness judgment auditable.
  • Systemic causes: fix prompts / tools / sources / agent definitions / memory under § META authority; the versioning rule applies in full (banner bump + CHANGELOG entry in the same commit). What exceeds authority (scheduler config, org-profile values, hosting) becomes a numbered operator recommendation.
  • Single-source claims touching the constituency that fail corroboration: not published, recorded as watch items with what would resolve them, the next audit re-checks.
  • Warnings on settled history (v3.28): a WARN whose subject cannot change without falsifying the record; a published run record's stall-length duration_seconds, an era-correct recorded confirmation waiver (run records stay immutable; entries do not, so a WARN on an entry is fixed through a changelog record, never acknowledged), is acknowledged in state/warning_acknowledgments.json: check label, match pinned to the specific run id/subject, reason, acknowledged_at. This is a reviewed audit decision disclosed in the report (the ledger diff is part of the audit commit); NEVER acknowledge a warning that a code/prompt/state/renderer fix could clear instead, and a run never self-acknowledges its own fresh warnings.

Phase 5; Audit report (skeleton-then-Edit)

Write docs/audits/<RUN_DATE>-quality-audit.md (earlier reports carry the -weekly-quality-audit suffix from the retired naming; read them with either name), mirroring the 2026-07-11 report's structure:

  1. Header + method: mandate, window, run-record link, sub-agents spawned, entries checked / primary URLs fetched.
  2. Verdict: lead with clean/total entries verified and the one-paragraph honest state of the pipeline.
  3. Findings, false or erroneous published intelligence (table: entry, defect, ground truth, fix, the fix names the correction / improvement record appended) + root causes. Minor imprecisions documented even when no record is warranted.
  4. Findings, missing or incomplete coverage: recovered items, correctly-droppable borderlines with reasons, resolved false alarms.
  5. Findings, systemic / operational (incl. fix-effectiveness results from Phase 3 item 7).
  6. Priority calibration (monthly fires only, Phase 3b output).
  7. Fixes shipped in this commit: including the warning sweep's outcome: what was fixed, and every new/removed state/warning_acknowledgments.json entry with its reason.
  8. Recommendations (operator decisions, not shipped), numbered, carried forward until adopted or explicitly retired.
  9. Watch items: carried-forward status + new items, each with its resolution condition.

Empty sections state "none found"; the absence is a result. Then write the run record: telemetry frontmatter (audit passes under sub_agents:), verification & coverage notes summarizing the report, entries_published = recovered (new) entries only, entries_updated + updated_entry_ids[] = the entries that received a correction / improvement / update record this fire.

Phases 5.5 → 7, State, gate, verify, publish

Read prompts/cti-run.md now (Phases 5, 5.5, 5.7, 6, 7) and execute verbatim with this run's RUN_ID: state lifecycle only where recovered entries touched it; python3 tools/check_run.py "$RUN_ID" --pre-verify exit 0 before the first verifier spawn. The Phase 5.7 loop (≥1 iteration ALWAYS, ≥2 for a CLEAN publish per the double-CLEAN gate, two consecutive CLEAN passes of the single verifier definition, cap 8, fail-open) is scoped to: this run's recovered entries + every entry it appended a changelog record to (read whole; new section and changed fields verified against the cited sources) + the run record + the audit report; the report makes checkable claims about published files, records and state, and a cold reader must confirm they hold on disk (the 07-11 audit's own iteration 1 caught exactly this class: a claimed repair log entry that wasn't there). Then stage specifics (incl. docs/audits/, work/<run-id>/, .claude/memory/ when touched), commit, sync origin/main, push with retry, Phase 7 publish polling, report publish: from the actual poll.


Quality gates (self-check)

  • [ ] Every window entry was covered by exactly one truth pass, or its batch's abandonment is logged as an audit coverage gap.
  • [ ] Both audit halves ran: truth passes (soundness) AND independent re-sweeps (completeness), or the cut is recorded with its cause (a stall, a blocked ladder, a dead transport, never elapsed time).
  • [ ] Every confirmed defect carries a root cause and either a shipped fix or a numbered operator recommendation, no orphan findings.
  • [ ] Every change to a published entry is a correction / improvement / update changelog record (with its section when reader-facing, internal: true and no section when metadata-only; updated_at bumped only by type: update) listed in updated_entry_ids[] and disclosed in the report; no silent edit, no second entry for a covered finding.
  • [ ] Previous audit's watch items re-checked and its fixes effectiveness-checked.
  • [ ] A legacy re-verification batch was checked and marked in state/legacy_review.json (v4.19), and every duplicate found was folded or linked.
  • [ ] Monthly duty honored: at most one ## Priority calibration section per calendar month across audit reports, and at least one unless no audit fired that month.
  • [ ] Recovered entries passed the full normal gates (dedup, check_run.py, verifier loop), audit provenance never lowers a bar.
  • [ ] Zero-warning sweep complete: python3 tools/check_run.py --all ends 0 warn · 0 fail (acknowledged excepted) and python3 site/build.py emits no self-check warnings; every ledger change is disclosed in the report.
  • [ ] No manufactured findings; clean components reported clean.
  • [ ] All shared gates from prompts/cti-run.md § Quality gates hold (gate exit 0, ≥1 verifier iteration, run record committed, publish verified).

Output

run: runs/YYYY-MM-DD/<run-id>.md
audit: docs/audits/YYYY-MM-DD-quality-audit.md
window: <start> → <end> · entries audited: N (clean: N) · runs reviewed: N
findings: erroneous: N · coverage gaps: N (recovered: N) · systemic: N · calibration: run | not-due
fixes shipped: N · entries corrected/improved (changelog records): N · watch items: N open
warnings: 0 open (N acknowledged) · build self-check: clean
commit: <short SHA> · push: ok (feature branch) | failed (<reason>) · publish: ok | main-only | pending (<reason>)

META, self-evolution authority

Same authority and process as prompts/cti-run.md § META. All hard invariants apply, plus:

  • A-INV-1: the audit changes a published entry only through a dated changelog record (correction / improvement / update; section when reader-facing, internal: true without one when metadata-only; updated_at moves only on type: update), never a silent edit, never a second entry for the same finding.
  • A-INV-2: findings are never manufactured; a clean week is the success outcome and is reported as such.
  • A-INV-3: the audit report and run record ship even when the audit finds nothing or a stall cuts it short; a silent audit is an operational failure.
  • A-INV-4: the audit ships fixes under the versioning rule and never weakens a hard invariant, concerns about an invariant become recommendations, not edits.
  • A-INV-5: the audit never blocks or races the scheduled intel runs; it works on published history plus its own artifacts, on its own feature branch (per-run paths make a concurrent intel fire safe for new entries; an entry both fires update in the same window is a file conflict the sync surfaces, keep both records in at order, never drop one).