CTIPilot

2026-09-13T1307Z-audit

One pipeline fire, in full · audit run of 2026-09-13 · sub-agent allocation and telemetry, per-iteration verification verdicts and findings, source-list edits, coverage gaps, bridge invocations, and the run's own verification & coverage notes: what was published, what was dropped at the borderline or judged not relevant (and why), single-source carve-outs, and contradictions. Rendered from runs/2026-09-13/2026-09-13T1307Z-audit.md.

Run telemetry

2026-09-13T1307Z-audit audit prompt v4.10 publish ok
1h 33m duration 0 published 8 updates
Claude Opus 5 (claude-opus-5) main agent
F1 Claude Sonnet 5 (claude-sonnet-5)
Items returned
2
Duration
8m 14s
Tool calls
0 WebFetch0 WebSearch12 bridge
Cited sources
none
G1 Claude Sonnet 5 (claude-sonnet-5)
Items returned
0
Duration
6m 50s
Tool calls
8 WebFetch17 WebSearch12 bridge
Cited sources
none
G2 Claude Sonnet 5 (claude-sonnet-5)
Items returned
1
Duration
9m 53s
Tool calls
4 WebFetch24 WebSearch16 bridge
Cited sources
none
G3 Claude Sonnet 5 (claude-sonnet-5)
Items returned
13
Duration
14m 56s
Tool calls
17 WebSearch46 bridge
Cited sources
none
truth-A Claude Sonnet 5 (claude-sonnet-5)
Items returned
17
Duration
9m 12s
Tool calls
1 WebSearch45 bridge
Cited sources
none
45 URLs checked
truth-B Claude Sonnet 5 (claude-sonnet-5)
Items returned
17
Duration
14m 57s
Tool calls
4 WebSearch2 bridge
Cited sources
none
45 URLs checked
truth-C stalled

Claude Sonnet 5

Past the 30-min wall-clock cap; abandoned.

truth-C1 stalled

Claude Sonnet 5

Past the 30-min wall-clock cap; abandoned.

truth-C1a stalled

Claude Sonnet 5

Past the 30-min wall-clock cap; abandoned.

truth-C1b Claude Sonnet 5 (claude-sonnet-5)
Items returned
5
Duration
6m 19s
Tool calls
not reported
Cited sources
none
truth-C2 Claude Sonnet 5 (claude-sonnet-5)
Items returned
8
Duration
10m 12s
Tool calls
0 WebFetch0 WebSearch30 bridge
Cited sources
none
33 URLs checked

Verification

✓ double-CLEAN · Sonnet 5 ×2 #1 NEEDS_FIXES · Sonnet 5 · t=1 e=1 a=0 #2 CLEAN · Sonnet 5 · t=0 e=0 a=1 #3 NEEDS_FIXES · Sonnet 5 · t=2 e=0 a=0 #4 CLEAN · Sonnet 5 · t=0 e=0 a=0 #5 CLEAN · Sonnet 5 · t=0 e=0 a=3

Deep dive

·

Entries this run published (0) and updated (8)

Sources changed (this run)

Edits this run made to sources/sources.json · promotions, demotions, new candidates, and fetch-method / category / reliability / url corrections (the run record's sources_changed[]). Paginated; 10 per page.

No source-list edits recorded for this run.

Coverage gaps (this run)

Sources this run's brief needed that returned no usable content via any documented recipe. Bridge-recovered or quiet-day sources do NOT appear here. (Distinct from the independent source-accessibility probe at the foot of this section, which probes all active sources regardless of what any run needed.)

No coverage gaps in this run · every source the brief needed returned usable content via its documented recipe.

Bridge invocations (this run)

5 bridge calls this run · these are successful bridge fetches (separate from "Coverage gaps" above).

5 other
  • ×5

Verification findings · all iterations

Per-iteration finding detail. Each table is one verifier pass · what was flagged, how the main agent remediated it, and the outcome. Walking the tables top-to-bottom shows the verifier's debugging trail across iterations.

Iteration #1 NEEDS_FIXES · 2 findings (truth=1, editorial=1, advisory=0) · Claude Sonnet 5 · ·

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The report's method paragraph misstated this run's own verifier telemetry three ways against the run record's sub_agents block: "five" retrospective truth passes where four returned, "eight" spawns atMethod paragraph rewritten: seven attempted, three blocked, four returned covering 47 entries (17/17/5/8); the two blocked-spawn entries and the one inventory m
F6
strengthen-primary-source
·
(low confidence) The gate's aggregator-only warning on this entry is disposed of in the run record's notes but was not mentioned in the audit report, so a reader of the report alone would not see thatAdded as its own numbered item in the report's "Fixes shipped" section, stating the PD-5 victim carve-out and that the entry is carried as-is.

Iteration #2 CLEAN · 2 findings (truth=0, editorial=0, advisory=1) · Claude Sonnet 5 · ·

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F3
claim-not-supported
·
The correction record and the audit report asserted a single EPSS percentile, 0.97366, for both 2026-07-24 and 2026-07-25; FIRST.org returns 0.973660000 for the first date and 0.973690000 for the secoBoth the record and the report give the percentile per date (0.97366 on 07-24, 0.97369 on 07-25). The probability 0.21621 is unchanged and identical on both dat
F11
editorial-advisory
·
(advisory, explicitly not raised as a defect) The CISA KEV source added for CVE-2026-48710's exploited/cisa-kev flags is the bulk catalog JSON rather than a per-CVE page. The verifier confirmed this iNone. The KEV catalog has no per-CVE URL, the bulk feed is the documented fetch path for cisa-kev, and changing one entry away from a store-wide convention woul

Iteration #3 NEEDS_FIXES · 2 findings (truth=2, editorial=0, advisory=0) · Claude Sonnet 5 · ·

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The load-bearing one. The main analysis still opened a paragraph with "Compensating controls, not patching, are the available lever", contradicting the entry's own now-patched frontmatter, actions[] aThe paragraph now opens by stating it describes the position at disclosure and directs the reader to patch first, pointing at the 2026-09-13 update. The dated C
F4
hallucinated-fact
·
(low confidence, low severity) The n=2 iteration block's findings array carried three records while its truth/editorial/advisory counters summed to two: the F6 record belonging to iteration 1 had beenF6 moved back into iteration 1, whose counters (truth 1, editorial 1) now match its two findings. Iteration 2 keeps F11 as its single counted advisory plus the

Verification & coverage notes

The run record's narrative body, verbatim. This is where the run accounts for its own judgement calls: every borderline drop and judged-not-relevant item with its reason, dedup decisions, single-source items and their carve-outs, contradictions, and per-source coverage gaps, so nothing the run considered disappears silently.

Verification & coverage notesrun record body

2026-09-13T1307Z-audit · audit · Opus 5 · window 168 h · 0 entries published

Verification & coverage notes

Audit report: docs/audits/2026-09-13-quality-audit.md. Window 2026-09-06T13:08Z → 2026-09-13T13:07Z (168.0 h), anchored on the previous audit record's started; seven intel fires; 50 entries in scope (26 published new, 24 older entries carrying an in-window changelog record).

Soundness: 42 of 50 entries verified clean, 2 factual errors, 6 imprecisions. The two factual errors are the September 2026 Patch Tuesday CVE count on 2026-09-09/windows-september-2026-two-exploited-lpe-zero-days-kev (stated "roughly 1,170"; BleepingComputer says 966 and ZDI "nearly 1,000", and 1,170 appears in neither) and the EPSS probability on 2026-07-24/laundry-bear-zimbra-zero-click-cve-2025-66376 (0.1201 against FIRST.org's 0.21621 for the entry's own dates, a magnitude error the 2026-09-06 audit's unit-only conversion carried forward). Both corrected through the entry's changelog.

The window's headline finding is that the verifier can be blocked reproducibly by content. The 2026-09-09T1726Z-intel fire published two entries with iterations: [] after four cti-verification spawns were terminated by the content-safety classifier, and its record asked the next audit for an independent pass. This audit attempted it and reproduced the trip three more times across three further framings, the full 15-entry batch (killed mid-flight), a 7-entry defensive-fact-checker reframe with manifest-file scope (killed on spawn), and the 2 entries alone with per-entry checkpointing (killed on spawn), while two sibling batches covering the other 13 entries of the same split completed normally. Seven blocked spawns across two fires and four framings establishes the trip as a property of the content, not the message. The main agent verified the three uncovered entries itself (batch D: the two 2026-09-09 entries plus 2026-09-12/jfrog-artifactory-…, which the first-pass inventory had missed) and found a factual error in one of them (the Patch Tuesday count above) while confirming everything else those entries claimed against MSRC per-CVE records, the KEV catalog, MITRE CNA records and the cited primaries.

Completeness inside the window: no gap. The mechanical KEV sweep found 14 in-window additions and 0 uncovered, every row resolving to a named entry, the second consecutive clean KEV window. G1 returned zero new items; G2 one borderline (Oomnium, a Zurich crowdfunding platform, correctly out of nexus); G3 thirteen candidates, all either already covered, correctly droppable under the v4.2 quality-over-quantity bar, or a development on a covered finding.

Two pre-window findings the re-sweeps surfaced. First and more serious: 2026-08-12/shieldbreak-defender-rogueplanet-patch-bypass-no-fix has told readers since August that no fix exists for CVE-2026-69414, in its headline, summary, action item and twice in its body. Microsoft's own record has named the fix since 2026-09-03, Malware Protection Engine 1.1.26080.3, with 1.26070.7 the last affected, and RL:O in the vector. That is the defect class where the whole remediation inverts. Fixed by an update record that moves the status, names the fixed engine build, rewrites the stale statements where they stand, and replaces the action item with an explicit engine-version check (the Defender engine updates on its own cadence and is not the OS patch level). The same record carries the ShieldCrash development with its attribution intact: a partial bypass claimed by the researcher's own repository, no new CVE, explicitly unfinished, and not confirmed by Microsoft.

Second: CVE-2026-27912 (ResetNightmare) entered state/cves_seen.json on 2026-08-09 and no entry ever covered it. Opened as a recovery candidate on the strength of its index title (Kerberos, low-privileged user to Domain Admin) and closed as a correct drop by the primaries: Microsoft patched it in April 2026, rates it CVSS 8.0 AV:A/…/E:U/RL:O, sets exploited: No and "Exploitation Less Likely", and it is not on CISA KEV, PD-11(b)'s excluded case. The in-window 0patch backport reaches only unsupported Windows Server estates under a third-party patch subscription, too narrow for an entry under the v4.2 bar. No entry recovered; the deep read is persisted under work/ so a future fire need not redo it. What it exposed is that nothing checks the index-into-entries direction (253 ids store-wide sit in that gap, mostly legitimately) raised as operator recommendation 2.

Entries updated (8), all through their changelogs, none silently: four correction records (Windows Patch Tuesday count; Zimbra EPSS, internal; Japan Digital Agency credibility 1→2 with its sourcing note; Revolut spliced quotation), three improvement records (Dell revision history; LiteLLM KEV citation, internal; NetScaler live-counter as_of, internal) and one update record (ShieldBreak / ShieldCrash). No entry was published new.

Fixes shipped. Prompts to v4.10 in lockstep with a CHANGELOG entry: a new exhausted-ladder rung in Phase 5.7 requiring the main agent to take the truth half of the gate on its own output when every spawn is blocked, and to record what it did and could not do; the "one iteration is mandatory" hard rule fenced to name that single exception. tools/check_run.py carries store severity for verification.iterations missing or empty under --all only (run scope still FAILs, tested against the 2026-09-09 record) because a published record is immutable and fail() never consults the acknowledgment ledger, so the FAIL was permanently unclearable. tools/kev_window_diff.py gains --run-id and writes work/<run-id>/kev-window.txt itself, because fourteen fires across two windows were asked to tee that file and none did.

Source-health work. The NCSC.ch carry-forward watch item is resolved and the recipe was genuinely broken: both BACS pages are Nuxt SPAs returning an empty shell to their recorded transports, which is why they stayed green at seven and six quiet periods while contributing nothing since 2026-06-18; ncsc-ch-incidents had its own note recording that the bridge returned "a JS-only shell" and was never switched. Both now pinned to the transport verified to hydrate them. The previous audit's recommendation 3 is discharged three of four: Volexity and Proofpoint were never broken (an RSS hunt on a feedless host, and a WebFetch summariser eating a listing), SocRadar works with its listing-date metadata recorded as unreliable, GreyNoise re-confirmed; Aqua Nautilus is not a record in sources/sources.json at all. ReliaQuest, IBM X-Force and Jamf Threat Labs are confirmed dark across extract, jina and bridge url and left active rather than demoted, so the blind spot stays visible.

Warning sweep. One new acknowledgment (the 2026-09-09 empty verifier block, with the reproduction evidence and an explicit statement of the fix deliberately not taken); existing 31 rows reviewed, none dead, none pruned; ledger now 32. check_run.py --all ends 0 warn · 0 fail (32 acknowledged) and site/build.py emits no self-check warnings.

  • Reduced-confidence note (aggregator-only, the check's own documented disposition): 2026-09-13/revolut-fake-government-request-kyc-breach cites two news hosts and no vendor or regulator primary. That is inherent to the story rather than a sourcing shortfall, every fact traces to Revolut's own customer notification and spokesperson statement, which is the PD-5 victim carve-out the entry already records as verification: single-source-victim, and no regulator has published. Carried as-is.
  • Monthly priority calibration: not due; the 2026-09-06 report carries the September section. Context only: high ran 63.6 % of operational entries this window (n=22) against a store-wide 51.7 %, with three criticals; flagged for the October pass, and no mis-prioritized entry was found among them.
  • Coverage gaps: inside-it-ch returns HTTP 429 site-wide on every transport (Insel Gruppe still blocked on it); urnerzeitung.ch 403; netzwoche.ch has no feed; reliaquest, ibm-xforce and jamf-threat-labs serve content-free shells.
  • ATT&CK pin: attack_data.py --check reports up to date, local v19.2 == upstream latest v19.2.
  • Watchdog: no fire in the window tripped the runaway-duration threshold, against five the previous window; the longest was 2.90 h.

← Operations dashboard · run-record contract: docs/pipeline.md