CTIPilot

2026-09-27T1308Z-audit

One pipeline fire, in full · audit run of 2026-09-27 · sub-agent allocation and telemetry, per-iteration verification verdicts and findings, source-list edits, coverage gaps, bridge invocations, and the run's own verification & coverage notes: what was published, what was dropped at the borderline or judged not relevant (and why), single-source carve-outs, and contradictions. Rendered from runs/2026-09-27/2026-09-27T1308Z-audit.md.

Run telemetry

2026-09-27T1308Z-audit audit prompt v4.12 publish ok
2h 37m duration 1 published 8 updates
Claude Opus 5 (claude-opus-5) main agent
G1 Claude Sonnet 5 (claude-sonnet-5)
Items returned
5
Duration
12m 10s
Tool calls
10 WebSearch28 bridge
Cited sources
8 of 22 in slice
G2 Claude Sonnet 5 (claude-sonnet-5)
Items returned
4
Duration
10m 25s
Tool calls
1 WebFetch27 WebSearch28 bridge
Cited sources
3 of 14 in slice
G3 Claude Sonnet 5 (claude-sonnet-5)
Items returned
18
Duration
17m 53s
Tool calls
0 WebFetch7 WebSearch55 bridge
Cited sources
16 of 32 in slice
truth-A Claude Sonnet 5 (claude-sonnet-5)
Items returned
14
Duration
11m 21s
Tool calls
0 WebFetch0 WebSearch8 bridge
Cited sources
none
40 URLs checked
truth-B Claude Sonnet 5 (claude-sonnet-5)
Items returned
16
Duration
12m 54s
Tool calls
0 WebFetch0 WebSearch
Cited sources
none
truth-C Claude Sonnet 5 (claude-sonnet-5)
Items returned
14
Duration
10m 24s
Tool calls
0 WebFetch0 WebSearch3 bridge
Cited sources
none
30 URLs checked
truth-D Claude Sonnet 5 (claude-sonnet-5)
Items returned
10
Duration
7m 04s
Tool calls
0 WebFetch0 WebSearch17 bridge
Cited sources
none
17 URLs checked

Verification

#? CLEAN · Sonnet 5 · t=0 e=0 a=3 #? NEEDS_FIXES · Sonnet 5 · t=3 e=0 a=1 #? NEEDS_FIXES · Sonnet 5 · t=0 e=3 a=1 #? NEEDS_FIXES · Sonnet 5 · t=1 e=3 a=1 #? NEEDS_FIXES · Sonnet 5 · t=1 e=1 a=0 #? NEEDS_FIXES · Sonnet 5 · t=1 e=1 a=0 #? NEEDS_FIXES · Sonnet 5 · t=3 e=0 a=0 #? NEEDS_FIXES · Sonnet 5 · t=1 e=0 a=0

Deep dive

·

Entries this run published (1) and updated (8)

Sources changed (this run)

Edits this run made to sources/sources.json · promotions, demotions, new candidates, and fetch-method / category / reliability / url corrections (the run record's sources_changed[]). Paginated; 10 per page.

No source-list edits recorded for this run.

Coverage gaps (this run)

Sources this run's brief needed that returned no usable content via any documented recipe. Bridge-recovered or quiet-day sources do NOT appear here. (Distinct from the independent source-accessibility probe at the foot of this section, which probes all active sources regardless of what any run needed.)

No coverage gaps in this run · every source the brief needed returned usable content via its documented recipe.

Verification findings · all iterations

Per-iteration finding detail. Each table is one verifier pass · what was flagged, how the main agent remediated it, and the outcome. Walking the tables top-to-bottom shows the verifier's debugging trail across iterations.

Iteration #? CLEAN · 3 findings (truth=0, editorial=0, advisory=3) · Claude Sonnet 5 · 12m 01s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F11
editorial-advisory
·
(low confidence) affected_products lists only Microsoft Entra ID, although the source frames the design flaw as equally a Windows-side one: the reachable surface is the shared Windows Runtime activatiAccepted after re-reading the source. affected_products is now ["Microsoft Windows", "Microsoft Entra ID"]; product:microsoft-windows already exists in the regi
F11
editorial-advisory
·
(low confidence) the v4.12 entry's rotation figures read as self-contradictory standalone: "64 never did" alongside "22 of the 115 were handed to no sub-agent at all", without the qualifier that sourcAccepted. The CHANGELOG sentence now carries the qualifier explicitly, matching the framing the audit report uses.
F11
editorial-advisory
·
(low confidence) a pre-existing changelog record dated 2026-09-11, not written by this fire, uses the word "frontmatter" in a reader-rendered summary field.Declined, correctly: earlier updates[] records are append-only and are the audit trail, so this fire has no authority to rewrite one. The defect class is alread

Iteration #? NEEDS_FIXES · 4 findings (truth=3, editorial=0, advisory=1) · Claude Sonnet 5 · 14m 24s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The report described all 54 audited entries as what "the week's fires produced", 36 new and 18 updated. The seven intel fires produced 43 of them (35 new, 8 updated); the other 11 are the previous audThe 54 stands and is the correct audited population; the framing did not. Method and verdict now state the 43/11 split explicitly and name batch D as the fix-ef
F4
hallucinated-fact
·
Three rows this fire added (Dyfed-Powys Police, DIVD, Maileva) duplicated rows the 2026-09-25 and 09-26 intel fires had already opened for the same items, against the file's own stated convention thatThe three duplicate rows were removed and the re-check findings appended as dated notes on the original rows instead. Open rows are 22 over 22 distinct topics;
F4
hallucinated-fact
·
(low confidence) The update section called one typosquat target "a press-freedom media organization". Volexity's post names China Digital Times and the Center for American Progress but never characterAccepted. The section now names both organisations as the source does, which is fully supported and also more useful to a reader, and drops the characterisation
F11
editorial-advisory
·
The new entry omitted five frontmatter keys that docs/pipeline.md documents as standard and every comparable entry carries: deep_dive, deep_dive_category, org_triage, watchlist_hit, migrated_from. TheAll five backfilled at their correct null or false values. Worth noting for the operator: this is a schema-completeness gap the mechanical gate does not enforce

Iteration #? NEEDS_FIXES · 4 findings (truth=0, editorial=3, advisory=1) · Claude Sonnet 5 · 7m 28s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F9
surface-contradiction
·
The sub_agents block carried no G3 record although the run record's own prose names G3 twice and gap-G3.yaml, G3.started_at and G3.ended_at are all on disk. The skeleton was written before G3 returnedG3 block added with its timestamps, scope, 32 publishers attempted, 16 with in-window content, 18 items returned and its telemetry, all read off gap-G3.yaml and
F9
surface-contradiction
·
Iteration 2's counters recorded its one F11 finding as editorial rather than advisory, inconsistent with iteration 1's counting of F11 in the same record. The four-finding total was unaffected.Iteration 2 recounted as truth 3 / editorial 0 / advisory 1.
F10
missed-angle
·
One of G1's five returned items, Rapid7's teaser naming more than 50 Zimbra Collaboration Suite vulnerabilities for business email compromise, appeared in no artefact: not the report, not the run recoDispositioned in the report's coverage section as not publishable (no CVE, no advisory, no patch, no exploitation, so nothing a reader can act on) and carried a
F11
editorial-advisory
·
(low confidence) The entry self-discloses an inherited accuracy gap in an append-only changelog record. Noted for completeness; no action requested.Declined, correctly: earlier updates[] records are the audit trail and cannot be rewritten. This fire's correction record already documents the gap where a read

Iteration #? NEEDS_FIXES · 5 findings (truth=1, editorial=3, advisory=1) · Claude Sonnet 5 · 13m 34s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The report's "correctly droppable, verified" line said four items had each been reached by a sweep, gated and correctly left out. They came from the incident re-sweep's return message, not from gap-G2Reworded to say what is true: the sweep reported reaching and gating them, the facts given put them out of scope, and this audit did not re-derive that independ
F9
surface-contradiction
·
G3's telemetry (websearch_calls 7, bridge_fetches 55) has no supporting artefact: unlike G1 and G2, gap-G3.yaml carries no self_telemetry block, so those two numbers rest on the return message alone aBoth the provenance and the approximation are now recorded, on the telemetry line and in work/<run-id>/subagent-returns.md. G3's sources_attempted, sources_used
F9
surface-contradiction
·
The research-items row said seventeen publications remained open, but one of the five Huntress items it counted was published this fire as a standalone entry, which the row's own text says. The true oRow corrected to sixteen and four Huntress, with the arithmetic stated in the row: eighteen returned, one published as an entry, one shipped as a changelog reco
F11
editorial-advisory
·
(low confidence) entities[] omitted product:microsoft-windows although affected_products names Microsoft Windows and that registry key exists.Accepted. entities[] now carries both product keys, matching the pattern comparable entries use.
F9
surface-contradiction
·
(low confidence) Recommendation 6's figure of 11 entries and 29 occurrences could not be reproduced, landing at 10/24, 12/28 or 22/37 depending on which pattern set was applied.Re-measured and the figure holds at exactly 11 entries and 29 occurrences, 15 in sources[] and 14 as body links. The irreproducibility was the pattern set: this

Iteration #? NEEDS_FIXES · 2 findings (truth=1, editorial=1, advisory=0) · Claude Sonnet 5 · 6m 05s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F9
surface-contradiction
·
Iteration 4's counter block read truth 1 / editorial 2 / advisory 2 while its own findings list is one F4, three F9 and one F11, which maps to 1 / 3 / 1. The total was right and the split was not, theIteration 4 recounted as 1 / 3 / 1, and the convention is now stated once in the notes: counters are normalised by F-code across every iteration rather than tra
F4
hallucinated-fact
·
(low confidence) techniques[] carried T1078.004, Valid Accounts: Cloud Accounts, although the source describes only theft and reuse of an OAuth token. It never describes use of account credentials disAccepted and removed. The mapping is now T1218, T1550.001 and T1528. This is the same evidence floor the audit applied when it ADDED ids to three other entries

Iteration #? NEEDS_FIXES · 2 findings (truth=1, editorial=1, advisory=0) · Claude Sonnet 5 · 9m 35s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
The report and the run record both named ncsc-ch-focus among the essential sources that contributed nothing all window and called it unreadable for two consecutive windows. The run's own artefacts conBoth documents corrected. ncsc-ch-incidents is genuinely dark and is named as such; ncsc-ch-focus is not, and the case is now reported as the more interesting o
F9
surface-contradiction
·
(low confidence) The stated counter convention put F13 to F15 in the editorial class; the verifier definition classes them as truth. No iteration produced an F13 to F15 finding, so no count was actualThe note now quotes the definition's own mapping: truth = F1 to F4 plus F13 to F15, editorial = F5 to F10 plus F12 and F16 to F18, advisory = F11. Every recorde

Iteration #? NEEDS_FIXES · 3 findings (truth=3, editorial=0, advisory=0) · Claude Sonnet 5 · 10m 54s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
·
Claimed the discipline-drift line's actions figure is wrong: 18 actions and mean 0.50 rather than the recorded 19 and 0.53.DECLINED with evidence. Re-counted two independent ways over the 36 window entries, by YAML parse and by listing every entry carrying a non-empty actions[]: the
F4
hallucinated-fact
·
(low confidence) Claimed only 12 vulnerability-kind entries exist this window, not the thirteen the report says, and that the 1.7 density reproduces against 21 ids over 12 entries.DECLINED on the count, ACCEPTED on what it exposed. There are 13: the thirteenth is 2026-09-20/oracle-september-2026-cspu-five-unauthenticated-cvss-10, which th
F4
hallucinated-fact
·
(low confidence) Both documents said the sources.json edit changed fetch_method, the failure counter and the notes, but the diff shows no counter value differing pre or post.ACCEPTED. Both records already carried consecutive_fetch_failures: 0, so setting it to 0 changed nothing; the diff is fetch_method and notes only. Both document

Iteration #? NEEDS_FIXES cap-breach · 1 finding (truth=1, editorial=0, advisory=0) · Claude Sonnet 5 · 6m 59s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F14
?
·
(low confidence) The update section said the typosquat pages, plural, load most of their content from the sites they copy and add the exploit components in a hidden iframe. Volexity demonstrates that Accepted and narrowed. The section now attributes the iframe delivery to the page that was still live and says plainly that the same mechanism is inferred rathe

Verification & coverage notes

The run record's narrative body, verbatim. This is where the run accounts for its own judgement calls: every borderline drop and judged-not-relevant item with its reason, dedup decisions, single-source items and their carve-outs, contradictions, and per-source coverage gaps, so nothing the run considered disappears silently.

Verification & coverage notesrun record body

2026-09-27T1308Z-audit · audit · Opus 5 · window 168 h · 1 entry published

Verification & coverage notes

Quality audit over 2026-09-20T13:08:12Z to 2026-09-27T13:08:37Z (168 h), anchored on the previous audit record. Full report: docs/audits/2026-09-27-quality-audit.md.

Soundness. Four retrospective truth passes covered all 54 entries in the window (36 new, 18 carrying a changelog record from it), each re-verified against freshly fetched primaries. 46 of 54 clean. Two factual errors, six imprecisions, no hallucinated source, no broken link, no IOC, no unrated entry. Batch D additionally re-checked all ten entries the previous audit corrected or improved: 10 of 10 fixes held, and both defects it surfaced were pre-existing errors that fire never touched.

Completeness. Three independent re-sweeps returned 27 items the week's fires had not published: 18 research-blog publications, 5 vulnerability items and 4 incident items. One was published as a new entry, one shipped as a changelog record on an existing entry, and eight coverage-backlog rows carry the rest with per-row gating notes and publish conditions. The KEV sweep was clean: 10 additions in the window, 10 of 10 already covered, kev-window.txt written by all seven fires.

Systemic. The v4.11 rotation fix worked and stopped one step short of the cause. Consecutive-fire slice overlap fell to 8 % on S3 and the four previously-starved research publishers were all swept, but the underlying ranking never advanced, so slices cycled with period 3: lag-3 overlap measured 82 to 92 % across all four domains, and 64 of 115 research sources were allocated to no fire at all. Fixed in v4.12 by ranking on a rotation cursor that an attempt advances, derived from run records rather than a new source field; simulated at 112 of 115 reached in seven fires with zero consecutive overlap. Separately, the store-wide YAML-portability check was extended to run records, which surfaced one record of 196 unreadable by any standards-compliant parser since May.

Operational. ncsc-ch-focus and ncsc-ch-incidents were both pinned to a reader pool with zero live keys and both fail their health probe on that pin; fetch_source.py extract reads both in full through trafilatura-direct, so both are re-pinned to bridge. ncsc-ch-incidents is genuinely dark, last contributing 2026-09-14. ncsc-ch-focus is not: the 2026-09-23 fire read it and published 2026-09-23/ncsc-ch-google-recovery-oauth-app-password-persistence from it, so the bad pin was masked by the fetch ladder working around it. Only fetch_method and the notes changed on either record; last_successful_fetch and the quiet counters stay as the fires left them, since this fire probed the transport rather than using either page for an entry. Reader-pool state noted as context only, per the standing rule that a dead pool is a normal condition rather than an incident.

Watch items: six resolved (verifier-loop convergence, rotation-fix effectiveness, the NCSC.ch pages, Boston Scientific, the KEV artefact, content-driven verifier blocking), seven still open, four new.

Warning sweep: four new acknowledgment rows, all settled history on immutable run records, including the previous audit's own unconfirmed-CLEAN fail-open, which that audit explicitly deferred to this one. All 33 existing rows still silence a live warning; none pruned. Ledger now 37 rows.

Calibration: not due (September's pass is in the 2026-09-06 report); continuity figures recorded in the report.

Essential-coverage: not applicable to an audit fire; the re-sweeps' source slices and per-publisher reachability are recorded in work/2026-09-27T1308Z-audit/gap-G1.yaml, gap-G2.yaml and gap-G3.yaml.

Coverage gaps: G2 did not separately drill cert-at, cert-fr avis-recent, or thirteen rotation sources, having prioritised the mandatory watch-item re-checks; inside-it-ch 429'd again; the Dyfed-Powys force statement page and one franceinfo article 403'd on every transport and were substituted with outlets quoting them directly. G3 found trellix, kela-cyber and esentire listings unreachable or empty by every transport tried, and paradigm-shift-research genuinely empty in-window; recipe work is operator recommendation 2. G1 did not attempt github-advisory (the egress proxy blocks github.com at session level) and had no package-specific lead to justify an OSV query.

Verification outcome. Eight iterations, the hard cap, ending on NEEDS_FIXES with one residual. The chain ran CLEAN, then NEEDS_FIXES at 4, 4, 5, 2, 2, 3 and 1 findings; iteration 1's CLEAN went unconfirmed because iteration 2 refused it and was right to, finding two wrong numbers in the audit report and three duplicated coverage-backlog rows. The entries themselves have been clean since iteration 2: every finding from iteration 3 onward was in the report's or this record's own bookkeeping, and the recurring failure mode was a self-referential number or narrative claim that the run's own artefacts contradicted. Iteration 8 re-derived essentially every such number from scratch and found them exact, settled both of iteration 7's disputed declines in favour of the declines, and stated that it would publish this output as it stands. The single residual is its own low-confidence F14, which was remediated before commit: the wording is narrowed, so the residual is recorded rather than outstanding. What iteration 8 said it would still check at a ninth pass, and which therefore did not get an independent read: the full 112-distinct-technique-id claim across all 54 window entries (it spot-checked the 13 vulnerability entries), the truth-pass and coverage-sweep YAML artefacts against this report's defect and gap tables (it re-verified the corrections against their primary sources directly instead), the Oracle, IBM MQ, Langflow and Adobe citations on the five new backlog rows, and the per-fire iteration-count and priority-calibration continuity figures. The next audit should treat those four as unverified claims of this fire.

Verification counters. Each iteration's truth / editorial / advisory counts are normalised by F-code using the verifier definition's own mapping, truth = F1 to F4 plus F13 to F15, editorial = F5 to F10 plus F12 and F16 to F18, advisory = F11, rather than transcribed from each pass's own severity call. Independent passes label the same code differently, and iterations 3 and 5 each flagged an un-normalised block; normalising once makes the chain readable across iterations. Every block still sums to the length of its own findings list, which is what the gate enforces.

ATT&CK pin: tools/attack_data.py --check reports local v19.2 equal to upstream latest v19.2. No drift.

← Operations dashboard · day page 2026-09-27 · run-record contract: docs/pipeline.md