CTIPilot

2026-09-20T1308Z-audit

One pipeline fire, in full · audit run of 2026-09-20 · sub-agent allocation and telemetry, per-iteration verification verdicts and findings, source-list edits, coverage gaps, bridge invocations, and the run's own verification & coverage notes: what was published, what was dropped at the borderline or judged not relevant (and why), single-source carve-outs, and contradictions. Rendered from runs/2026-09-20/2026-09-20T1308Z-audit.md.

Run telemetry

2026-09-20T1308Z-audit audit prompt v4.11 publish ok
2h 31m duration 1 published 10 updates
Claude Opus 5 (claude-opus-5) main agent
G1 Claude Sonnet 5 (claude-sonnet-5)
Items returned
6
Duration
9m 36s
Tool calls
0 WebFetch0 WebSearch34 bridge
Cited sources
14 of 14 in slice
G2 Claude Sonnet 5 (claude-sonnet-5)
Items returned
4
Duration
11m 28s
Tool calls
0 WebFetch20 WebSearch27 bridge
Cited sources
6 of 20 in slice
G3 Claude Sonnet 5 (claude-sonnet-5)
Items returned
6
Duration
17m 35s
Tool calls
6 WebSearch62 bridge
Cited sources
6 of 114 in slice
truth-A Claude Sonnet 5 (claude-sonnet-5)
Items returned
13
Duration
11m 41s
Tool calls
0 WebFetch0 WebSearch32 bridge
Cited sources
none
27 URLs checked
truth-B Claude Sonnet 5 (claude-sonnet-5)
Items returned
10
Duration
14m 14s
Tool calls
0 WebFetch2 WebSearch2 bridge
Cited sources
none
33 URLs checked
truth-C Claude Sonnet 5 (claude-sonnet-5)
Items returned
11
Duration
12m 21s
Tool calls
not reported
Cited sources
none
truth-D Claude Sonnet 5 (claude-sonnet-5)
Items returned
1
Duration
5m 56s
Tool calls
not reported
Cited sources
none

Verification

unconfirmed CLEAN · waived: Single CLEAN at the iteration cap (8). Iteration 8 returned CLEAN with zero find #? NEEDS_FIXES · Sonnet 5 · t=5 e=1 a=1 #? NEEDS_FIXES · Sonnet 5 · t=3 e=2 a=0 #? NEEDS_FIXES · Sonnet 5 · t=4 e=1 a=0 #? NEEDS_FIXES · Sonnet 5 · t=2 e=1 a=0 #? NEEDS_FIXES · Sonnet 5 · t=1 e=0 a=1 #? NEEDS_FIXES · Sonnet 5 · t=1 e=0 a=0 #? NEEDS_FIXES · Sonnet 5 · t=1 e=0 a=0 #? CLEAN · Sonnet 5 · t=0 e=0 a=0

Deep dive

·

Entries this run published (1) and updated (10)

Sources changed (this run)

Edits this run made to sources/sources.json · promotions, demotions, new candidates, and fetch-method / category / reliability / url corrections (the run record's sources_changed[]). Paginated; 10 per page.

No source-list edits recorded for this run.

Coverage gaps (this run)

Sources this run's brief needed that returned no usable content via any documented recipe. Bridge-recovered or quiet-day sources do NOT appear here. (Distinct from the independent source-accessibility probe at the foot of this section, which probes all active sources regardless of what any run needed.)

No coverage gaps in this run · every source the brief needed returned usable content via its documented recipe.

Verification findings · all iterations

Per-iteration finding detail. Each table is one verifier pass · what was flagged, how the main agent remediated it, and the outcome. Walking the tables top-to-bottom shows the verifier's debugging trail across iterations.

Iteration #? NEEDS_FIXES · 7 findings (truth=5, editorial=1, advisory=1) · Claude Sonnet 5 · 16m 49s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The correction record said the three NVD REST endpoints were replaced by Red Hat's pages, but only sources[] changed; all three were still cited inline in the analysis, attributed to a publisher the eAll three inline citations re-pointed at Red Hat's per-CVE advisory pages; the bridge-flaw sentence rewritten to what Red Hat states, since the function name an
F4
hallucinated-fact
·
The record was marked internal: true although both changed fields render on the entry page, the credibility code as a badge and the sourcing note as new reader-facing text.internal flag removed; the record now carries its `## Correction` section stating why the corroboration is editorial rather than independent.
F4
hallucinated-fact
·
The report's own fixes list said five correction and five improvement records; the disk carries six corrections and four improvements, and named the NTC record an improvement when it is a correction.Item 5 rewritten to the counts on disk, with the internal-record misjudgement disclosed.
F4
hallucinated-fact
·
The report claimed roughly 2,100 em dashes in entry source files; the store carries 8,415 across 906 of 917 files. The separate claim that none reach the rendered site was independently reproduced andFigure corrected to the measured 8,415 across 906 files.
F4
hallucinated-fact
·
(low confidence) The verifier's recount of the counter-mismatch finding gave 112 iterations across 60 records against the report's 140 across 68, with the denominators agreeing.Re-measured with the shipped check's own rule and corpus: 118 of 748 across 60 records. The 140 came from a looser pass over every file under runs/. Corrected i
F11
editorial-advisory
·
The run-record notes and the report use the words spawn and sub-agent, which the verifier read as forbidden workflow-internal language.Declined. The style rule scopes itself to reader-facing entry text (title, headline, summary, sourcing_note, body, changelog sections); the run record is the op
F12
single-source-flag-missing
·
(low confidence, advisory) `verification: single-source-national-cert` on a data-protection authority rather than a national CERT.Declined. prompts/verification.md defines the value for a national CERT OR government cybersecurity authority as primary disclosing party for its own jurisdicti

Iteration #? NEEDS_FIXES · 5 findings (truth=3, editorial=2, advisory=0) · Claude Sonnet 5 · 12m 59s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F10
missed-angle
·
The same release carries a sixth unauthenticated CVSS 10.0 flaw, CVE-2026-87230 in Hyperion Financial Management's Security component, which the entry never mentioned while framing itself around five.Verified against Oracle's own risk matrix (HTTP, remote-no-auth Yes, 10.0, PR None, UI None) and added: title, headline, summary, cves[], affected_products[], t
F3
claim-not-supported
·
The entry cited NCSC-NL's advisory page for a rating of high on both likelihood and damage; that page shows a single priority field, and the dual marker appears only in the advisory feed the entry doeClaim narrowed to the priority rating the cited page carries, and the 'only one of eleven' framing dropped.
F4
hallucinated-fact
·
This fire's own improvement record left a duplicated clause in the sourcing note while its summary claimed a clean plain-language rewrite.Duplicated clause removed; the note now reads as the record describes.
F4
hallucinated-fact
·
(low confidence) The corrected em-dash figure was still one occurrence out.Recounted over entries/*/*.md after this run's own additions: 8,416 occurrences across 906 of 917 files. Corrected in the report.
F8
needs-more-research
·
(low confidence) The CVE-database-API backlog count was right for sources[] records but missed an entry that cites the blocked pattern only inline in its body, which is exactly what the new body-scan Re-measured including body links: 12 entries, 18 sources[] records, 15 body links, one of them body-only. The report and recommendation 4 now carry both figures

Iteration #? NEEDS_FIXES · 5 findings (truth=4, editorial=1, advisory=0) · Claude Sonnet 5 · 9m 41s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The second action item still said all five flaws after iteration 2 moved the rest of the entry to six, contradicting the entry's own frontmatter.Action item corrected to six.
F4
hallucinated-fact
·
The coverage notes still described five flaws and repeated the NCSC-NL likelihood-and-damage framing that iteration 2 established the cited page does not carry.Notes moved to six flaws with Hyperion named, and to the priority rating the cited page shows.
F4
hallucinated-fact
·
The report's published-recovery paragraph and its verdict paragraph carried the same two stale claims, contradicting the report's own account of the iteration-2 fix.Both paragraphs swept: six CVEs listed, the NCSC-NL claim narrowed, and the immutable slug's disagreement with the corrected count called out.
F8
needs-more-research
·
(low confidence) The new Hyperion cves[] record omitted its affected-version string, which Oracle's matrix gives and which all five sibling records carry; the first action item punted on it too.Version 11.2.26.0.000 added to the record and to the action item. Re-reading the row also showed Hyperion's availability impact is rated none where the other fi
F4
hallucinated-fact
·
(low confidence) The 18 sources[] records figure for the CVE-database-API backlog did not reproduce; the verifier's count with the shipped matching logic gave 16.Declined after re-enumeration: printing every matching record one by one gives 18 across 12 entries, and the list is in this run's transcript. The report now no

Iteration #? NEEDS_FIXES · 3 findings (truth=2, editorial=1, advisory=0) · Claude Sonnet 5 · 7m 08s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The sourcing note still asserted NCSC-NL's own high likelihood and high damage rating, the claim iteration 2 found the cited page does not carry. The body was fixed then; this sibling frontmatter fielSourcing note moved to the high priority the advisory actually assigns. Third surface of the same propagation failure, and the last one.
F4
hallucinated-fact
·
(low confidence) The 48 Oracle CVE-index records figure did not reproduce under any counting method; the verifier's closest was 29.Accepted and root-caused: the substring used to count them matched Adobe ColdFusion on 'fusion'. A word-boundary count gave 28, which iteration 6 then showed st
F8
needs-more-research
·
(low confidence) The body grouped all six flaws under the single-sign-on, directory and application-server tier; Hyperion Financial Management is a financial-consolidation application from a separate Framing split: five components named as the identity and application-server tier, Hyperion named separately as the finance estate's exposure.

Iteration #? NEEDS_FIXES · 2 findings (truth=1, editorial=0, advisory=1) · Claude Sonnet 5 · 6m 23s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The oracle-cpu source note this run appended still said the September release carried five unauthenticated CVSS 10.0 flaws, the count corrected everywhere else. A fourth surface of the same propagatioNote corrected to six. Every surface carrying the count has now been swept: entry frontmatter, entry body, action items, sourcing note, run record, audit report
F11
editorial-advisory
·
(low confidence, advisory) Recommendation 5's count of entries carrying a production-process self-reference in sourcing_note reproduced as 81, not 80.Accepted after re-measuring with the gate's own _SELF_REF_RE, which is broader than the hand-written pattern the first count used: 81. Corrected in the report w

Iteration #? NEEDS_FIXES · 1 finding (truth=1, editorial=0, advisory=0) · Claude Sonnet 5 · 7m 19s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The corrected Oracle CVE-index figure of 28 still over-counted: two records match the word Oracle as cryptography rather than as a vendor, a padding oracle in Apache Tomcat and a quick-check oracle inCorrected to 26 Oracle-product records, with both misses named in the report so the figure's history is visible rather than tidied away.

Iteration #? NEEDS_FIXES · 1 finding (truth=1, editorial=0, advisory=0) · Claude Sonnet 5 · 11m 38s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The report summarised its eight imprecisions in three buckets that summed to seven and conflated an entry counted as a declined finding with one counted as an improvement.The summary is replaced by an itemization: the two fixed through a changelog record are named, the six recorded without one are named with the defect in each, a

Verification & coverage notes

The run record's narrative body, verbatim. This is where the run accounts for its own judgement calls: every borderline drop and judged-not-relevant item with its reason, dedup decisions, single-source items and their carve-outs, contradictions, and per-source coverage gaps, so nothing the run considered disappears silently.

Verification & coverage notesrun record body

2026-09-20T1308Z-audit · audit · Opus 5 · window 168 h · 1 entry published

Verification and coverage notes

Weekly quality audit over 2026-09-13T13:07Z to 2026-09-20T13:08Z. Full report: docs/audits/2026-09-20-quality-audit.md.

Soundness. 35 entries in scope, 22 verified clean, 5 factual errors, 8 imprecisions. All four truth passes returned, the first window in three with no blocked verifier spawn; the 2026-09-09 Windows entry that killed seven spawns across the two previous fires was isolated in a batch of one and verified without incident. The headline number is not comparable to the previous window's 42 of 50: this audit asked every pass to re-check status and fixed as of today rather than as of publication, and to treat citation adjacency as its own defect class, and three of the five errors are those two shapes.

The costliest error is a superseded remediation. 2026-08-04/cve-2026-20079-cisco-secure-fmc-auth-bypass-root-hotfix named per-train hot fixes for an exploited CVSS 10.0 authentication bypass; Cisco replaced them with the September hardening releases in revision 2.6 on 2026-09-16, in the same advisory that this entry's own 2026-09-18 update was reading for two sibling CVEs. Two of the five errors share that cause, a changelog record written against its own delta without re-reading the rest of the page it was already on.

Three findings were checked and declined with the rebuttal recorded in the report: a T1213 mapping on the Swiss Bitcoin Pay incident, a T1553 mapping on DDRop, and the national-authority carve-out on the AEPD entry.

Completeness. The three re-sweeps returned eight items the six fires never surfaced. One is published here: Oracle's September 2026 Critical Security Patch Update, six unauthenticated CVSS 10.0 flaws across WebLogic Server, Access Manager, Forms, Internet Directory, Platform Security for Java and Hyperion Financial Management, released 2026-09-15 and relayed by NCSC-NL on 2026-09-16 at priority Hoog. No run record in the window mentions Oracle at all, as an entry or as a drop.

Seven further verified items are on state/coverage_backlog.md with sources, discovery traces and per-row gating notes, and their research is committed under work/2026-09-20T1308Z-audit/. The cut is the wall-clock watchdog, applied in the prompt's stated priority order: the truth passes and the re-sweeps both completed, the systemic review completed, and publication was limited to the one recovered item whose absence would be a blind spot on the critical or high signal. Publishing eight entries would have put the fire past its budget with the report and record unwritten.

The KEV channel is clean for the third consecutive window: seven in-window additions, every one covered, and all six fires wrote the kev-window.txt artefact that fourteen fires across the two previous windows never produced.

Systemic. The window's largest finding is that the source rotation had stopped rotating. last_successful_fetch moves only when a source is fetched and used, so a source swept every fire that yields nothing keeps its stale date and stays pinned to the head of a stable oldest-first ranking. Six consecutive fires drew an almost identical S3 slice of roughly fourteen sources out of 111, and talos, sentinellabs and kaspersky-securelist were allocated to no fire all week. Five of the six research publications the re-sweep recovered came from publishers no sub-agent was ever given.

Three enforcement surfaces were found narrower than the rules they enforce: the reader-text check never looked at the body, its PD-number half was never implemented, and the blocked-source list covered the CVE databases' web pages but not their APIs (12 entries carried 21 such records, five of them primary). Verifier per-iteration counters do not reconcile with their own findings lists on 118 of 748 iterations store-wide, and two published records report zero residuals on a NEEDS_FIXES final iteration that carries a finding. All four are fixed in this commit; all four fixes are run-scope or version-gated so the store-wide scan stays at zero.

One fire of six reached a confirmed CLEAN, breaking a two-window streak of zero, so the carried watch item's redesign trigger did not fire. No fire tripped the runaway-duration threshold for the second consecutive window. All seven records carry publish_status: ok.

Warning sweep. One new acknowledgment (2026-09-14/gtg-27005-ai-drone-swarm-weapons-engineering, research-kind empty techniques[], the fourth row of a reviewed class); all 32 existing rows re-checked and none pruned; ledger now 33. check_run.py --all ends 0 fail with 33 acknowledged, and site/build.py emits no self-check warnings. One warning stays open by design: this fire's own verification-confirmation fail-open, a telemetry fact about this run that cannot be cleared without falsifying the record, and one a run is never permitted to acknowledge for itself.

Coverage gaps: none in the audit's own sweep. G1 reached roughly 14 of its 58 slice sources within its budget after the Cisco bundle and Oracle investigations, which is recorded as an audit coverage limitation rather than a pipeline finding.

← Operations dashboard · run-record contract: docs/pipeline.md