CTIPilot

2026-09-06T1308Z-audit

One pipeline fire, in full · audit run of 2026-09-06 · sub-agent allocation and telemetry, per-iteration verification verdicts and findings, source-list edits, coverage gaps, bridge invocations, and the run's own verification & coverage notes: what was published, what was dropped at the borderline or judged not relevant (and why), single-source carve-outs, and contradictions. Rendered from runs/2026-09-06/2026-09-06T1308Z-audit.md.

Run telemetry

2026-09-06T1308Z-audit audit prompt v4.9 publish ok
2h 51m duration 2 published 13 updates
Claude Opus 5 (claude-opus-5) main agent
G1 Claude Sonnet 5 (claude-sonnet-5)
Items returned
5
Duration
12m 00s
Tool calls
not reported
Cited sources
0 of 26 in slice
G2 Claude Sonnet 5 (claude-sonnet-5)
Items returned
2
Duration
10m 46s
Tool calls
not reported
Cited sources
0 of 26 in slice
G3 Claude Sonnet 5 (claude-sonnet-5)
Items returned
7
Duration
9m 09s
Tool calls
not reported
Cited sources
0 of 26 in slice
truth-A Claude Sonnet 5 (claude-sonnet-5)
Items returned
17
Duration
13m 48s
Tool calls
not reported
Cited sources
none
truth-B Claude Sonnet 5 (claude-sonnet-5)
Items returned
17
Duration
10m 36s
Tool calls
not reported
Cited sources
none
truth-C Claude Sonnet 5 (claude-sonnet-5)
Items returned
17
Duration
14m 37s
Tool calls
not reported
Cited sources
none

Verification

✓ double-CLEAN · Sonnet 5 ×2 #1 NEEDS_FIXES · Sonnet 5 · t=7 e=0 a=0 #2 NEEDS_FIXES · Sonnet 5 · t=5 e=1 a=0 #3 NEEDS_FIXES · Sonnet 5 · t=2 e=2 a=0 #4 NEEDS_FIXES · Sonnet 5 · t=5 e=0 a=1 #5 NEEDS_FIXES · Sonnet 5 · t=2 e=1 a=0 #6 NEEDS_FIXES · Sonnet 5 · t=4 e=0 a=1 #7 CLEAN · Sonnet 5 · t=0 e=0 a=2 #8 CLEAN · Sonnet 5 · t=0 e=0 a=0

Deep dive

·

Entries this run published (2) and updated (13)

Sources changed (this run)

Edits this run made to sources/sources.json · promotions, demotions, new candidates, and fetch-method / category / reliability / url corrections (the run record's sources_changed[]). Paginated; 10 per page.

No source-list edits recorded for this run.

Coverage gaps (this run)

Sources this run's brief needed that returned no usable content via any documented recipe. Bridge-recovered or quiet-day sources do NOT appear here. (Distinct from the independent source-accessibility probe at the foot of this section, which probes all active sources regardless of what any run needed.)

No coverage gaps in this run · every source the brief needed returned usable content via its documented recipe.

Verification findings · all iterations

Per-iteration finding detail. Each table is one verifier pass · what was flagged, how the main agent remediated it, and the outcome. Walking the tables top-to-bottom shows the verifier's debugging trail across iterations.

Iteration #1 NEEDS_FIXES · 7 findings (truth=7, editorial=0, advisory=0) · Claude Sonnet 5 · 10m 23s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
Removed the claim that the releases respond to Microsoft's Digital Crimes Unit pursuing legal action, it came from the entity registry's own summary, not from a fixed
F5
missing-citation
·
Same paragraph: two inline citations to The Hacker News added, one per quoted clause. fixed
F4
hallucinated-fact
·
CVE-2026-61409 corrected from cvss 7.5 / auth post-auth to Dell's own 7.3 and pre-auth; the id had been carried from the advisory's acknowledgements line withou fixed
F4
hallucinated-fact
·
CVE-2026-61409 affected/fixed narrowed to the Application component alone, matching its advisory row, which unlike the other three names no Appliance version; t fixed
F4
hallucinated-fact
·
T1685 removed. The entry describes abusing a security product's own remediation logic to escalate, which T1068 already carries; the only disabling described is fixed
F4
hallucinated-fact
·
T1078 removed. No source states an attacker obtained the SSH operator account; the described behavior is an existing account escalating locally, already covered fixed
F4
hallucinated-fact
·
The 2026-09-02 update section's fix-cadence sentence for CVE-2026-19318 still listed the incomplete set; corrected in place to include 2026.3.1, and the correct fixed

Iteration #2 NEEDS_FIXES · 6 findings (truth=5, editorial=1, advisory=0) · Claude Sonnet 5 · 15m 51s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
Quotation marks removed from 'continues to ghost them and refuses to engage in', which is the outlet's own third-person paraphrase rather than a quoted utteranc fixed
F4
hallucinated-fact
·
CVE-2026-48710's type, vector and auth aligned to Starlette's own advisory (PR:N/UI:N makes it pre-auth and zero-click, and the flaw is an auth bypass), resolvi fixed
F4
hallucinated-fact
·
The action item said the 2026.3 branch is affected by all four flaws, which overstates the per-hardware-line precision: WatchGuard places that band on the Defau fixed
F4
hallucinated-fact
·
The store-wide priority-calibration row was drafted before this fire published its own two entries and read 698/358; refreshed to the current 700/360, and the w fixed
F4
hallucinated-fact
·
The improvement record's fields omitted body although the record added a whole new changelog section; body added. fixed
F6
strengthen-primary-source
·
The new section cited an NVD per-CVE page as a live hyperlink, which is a blocked source pattern. The hyperlink is removed and the score attributed to NVD's rec fixed

Iteration #3 NEEDS_FIXES · 4 findings (truth=2, editorial=2, advisory=0) · Claude Sonnet 5 · 13m 52s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The em-dash figures were counted over a date-folder glob rather than over the window set the report itself defines, and were wrong: 35 of 39 in-window entries c fixed
F4
hallucinated-fact
·
The actions[] distribution was likewise measured on a different set and misstated: 18 of 39 (46.2 %) carry none, the mean is 0.82, and the longest list is two i fixed
F12
single-source-flag-missing
·
Credibility lowered from 1 to 2. Only the FalconFlank portion has two parties describing it consistently; PrettyPrague, HardBreacher and GreenSection rest on on fixed
F10
missed-angle
·
The registry summary carried an unsourced claim that Microsoft's Digital Crimes Unit threatened criminal action, and that text is where iteration 1's hallucinat fixed

Iteration #4 NEEDS_FIXES · 6 findings (truth=5, editorial=0, advisory=1) · Claude Sonnet 5 · 11m 06s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The ATT&CK density table's three historical rows were wrong. Recounted with site/content_model.py, the reference parser, which the finding's own figures match: fixed
F4
hallucinated-fact
·
cisa-kev was wrongly listed among the sources contributing no cited content: three in-window entries cite the KEV JSON feed directly, which is its documented fe fixed
F4
hallucinated-fact
·
The correction record's summary named an actions trim its own section never narrated. The section now explains the trim and names the three tasks that remain, s fixed
F4
hallucinated-fact
·
Same shape in reverse: the improvement record's summary claimed a sourcing-note wording change the section does not narrate. The clause is removed from the summ fixed
F4
hallucinated-fact
·
Iteration 3's note claimed it found no entry-level defect while its own findings list a rating-scope finding against a published entry. The note now says what t fixed
F11
editorial-advisory
·
Two aliases the fetched source and the new entry both name, INFINITE NIGHTMARE and MSNightmare, added to the record. Alias additions are append-only and the key fixed

Iteration #5 NEEDS_FIXES · 3 findings (truth=2, editorial=1, advisory=0) · Claude Sonnet 5 · 12m 58s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
This fire had rewritten the summary of a 2026-08-19 changelog record, which hard invariant #19 forbids: the EPSS correction was applied as a string replacement fixed
F4
hallucinated-fact
·
The correction record's summary still claimed a third change, a style fix in the analysis, that its section does not narrate. The clause is removed; the change fixed
F5
missing-citation
·
The PrettyPrague paragraph carried one citation across several sentences including a directly quoted vendor statement. Citations added per clause, matching the fixed

Iteration #6 NEEDS_FIXES · 5 findings (truth=4, editorial=0, advisory=1) · Claude Sonnet 5 · 10m 06s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F3
claim-not-supported
·
The correction section cited an undated FIRST.org EPSS URL to support a historical figure; because EPSS moves daily the link returns today's value, ten times th fixed
F3
claim-not-supported
·
Same defect, same fix: the dated query returns 0.5585 for 2026-08-18, confirming the figure the section states. fixed
F3
claim-not-supported
·
Same defect, same fix: the dated query returns 0.01368 for 2026-08-22, which independently corroborates the 0.0137 the unit conversion produces; the sentence no fixed
F4
hallucinated-fact
·
GreenSection was described as a crash in NVIDIA drivers where the source says only an NVIDIA memory-corruption bug affecting Vulkan and OpenGL applications. Nar fixed
F11
editorial-advisory
·
Not a new defect: two pipeline self-references inside 2026-08-13 and 2026-08-19 record summaries, already disclosed in the report. The verifier's structural poi declined

Iteration #7 CLEAN · 2 findings (truth=0, editorial=0, advisory=2) · Claude Sonnet 5 · 11m 29s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F11
editorial-advisory
·
The internal improvement record's summary said the CVE record disagreed with Starlette's advisory on three fields where the diff shows four; corrected to four, fixed
F11
editorial-advisory
·
Declined with a rebuttal put to the confirmation pass, which accepted it: The Hacker News carries role: primary although the source record notes it aggregates, declined

Verification & coverage notes

The run record's narrative body, verbatim. This is where the run accounts for its own judgement calls: every borderline drop and judged-not-relevant item with its reason, dedup decisions, single-source items and their carve-outs, contradictions, and per-source coverage gaps, so nothing the run considered disappears silently.

Verification & coverage notesrun record body

2026-09-06T1308Z-audit · audit · Opus 5 · window 167.9 h · 2 entries published

Verification & coverage notes

Quality audit over the 167.9 h since 2026-08-30T1312Z-audit, covering seven intel fires and 51 entries (37 published new, 14 carrying a changelog record inside the window). Full report: docs/audits/2026-09-06-quality-audit.md.

Soundness: 40 of 51 entries verified clean against freshly fetched primaries. Two factual errors, both in structured version data and both confirmed by the main agent against the authority before correction. WatchGuard: four of five CVE records omitted a second affected band (Fireware OS 2026.3 up to 2026.3.1) and its fix, so an appliance on a 2026.3.x build read as out of scope. HPE Aruba: CVE-2026-73749's 10.18 range was inverted, listing 10.18.0001 as the first affected build where HPE's own CVE record makes it the last. Nine further imprecisions, of which four took an improvement record and five are documented in the report without one.

A third defect class the truth passes did not look for and the main agent found: cves[].epss has never had defined units. docs/pipeline.md and prompts/entry-template.md carried the field as epss: null with no range and no source, and the store holds both conventions at once. ENISA's EUVD renders EPSS as a percentage (its API returns 0.71 where FIRST.org returns 0.00710, confirmed directly this fire) and FIRST.org returns the probability, so entries transcribed whichever their source showed. Two 100x-wrong values were published inside this window, and the second is the instructive one: on 2026-09-01 the verifier's iteration 1 set the correct FIRST.org probability and iteration 2 reverted it, reasoning that "the store's convention is a percentage number (confirmed by a pre-existing entry with epss: 1.37, impossible as a raw 0-1 probability)". A wrong legacy value taught a later verifier the wrong convention and it overrode a correct fix. Five entries corrected, the units defined normatively, and a cve-epss range check shipped.

Completeness: two recovered entries, both genuine blind spots. Chaotic Eclipse's unpatched local privilege escalations in CrowdStrike Falcon and Avast, with public working exploit code and no vendor fix, went unpublished for three days despite Truesec carrying it in a swept standard-tier feed. Dell's DSA-2026-382, a 105-CVE bundle whose top flaw replays one unauthenticated request indefinitely into ADMIN tokens with no workaround, went unpublished for six days. Everything else the three re-sweeps surfaced either matched published coverage or is documented as correctly droppable or backlogged.

KEV sweep, the v4.8 fix, took. All ten in-window CISA KEV additions were already covered by store entries; zero misses, against two recovered by the previous audit. The forensic half did not take: no fire wrote the work/<run-id>/kev-window.txt artefact the prompt specifies, though six of seven disclosed the sweep in prose.

The window's dominant systemic finding is verifier-loop convergence. No fire reached a confirmed CLEAN, against 2 of 30 over the preceding month; mean iterations rose to 7.6 against roughly 4.5 two windows ago; and the loop consumed 60 to 79 % of every fire's wall clock, which is the entire cause of the five runaway-duration warnings this fire acknowledged. The findings are mostly real rather than churn: 16 % of post-iteration-1 findings concern a previous iteration's own remediation, so 84 % are fresh. What that leaves unaddressed is a structural gap the report names: every one of the seven fires published under decision rule 5, whose final-iteration remediations no pass ever verified.

Coverage gaps this fire. G3 reached its cap with eight publisher listings unswept, and found four feed recipes returning nothing usable (Volexity, Proofpoint's Threat Insight blog, Aqua Nautilus, and SocRadar partially). G1 could not read ssd-disclosure (CAPTCHA on every transport). Eight essential-tier sources are green in state/source_health.json but contributed no cited content across all seven fires; two of them are the Swiss national authority's own pages, which the report flags for a recipe spot-check rather than treating as settled.

Priority calibration (monthly duty, owned by this fire). The 59.0 % high share the previous audit recorded did not persist: this window is 51.4 % of operational entries, September to date is 44.4 %, and the F16 signal is symmetric (two flags that a high was generous, two that one was under-calibrated). No calibration edit shipped.

Watchlist: not applicable; this deployment configures no product or supplier watchlist.

Coverage gaps: ssd-disclosure (CAPTCHA on every transport, no substitute); volexity, proofpoint, aqua-nautilus feed recipes returning no in-window items; eight G3 publisher listings unswept at the 45-min cap.

Essential-coverage: missed=none this fire (the audit does not run the full essential sweep; the per-fire misses over the window are reviewed in the report).

← Operations dashboard · run-record contract: docs/pipeline.md