ctipilot.ch

2026-08-02T1309Z-audit

One pipeline fire, in full · audit run of 2026-08-02 · sub-agent allocation and telemetry, per-iteration verification verdicts and findings, source-list edits, coverage gaps, bridge invocations — and the run's own verification & coverage notes: what was published, what was dropped at the borderline or judged not relevant (and why), single-source carve-outs, and contradictions. Rendered from runs/2026-08-02/2026-08-02T1309Z-audit.md.

Run telemetry

2026-08-02T1309Z-audit audit prompt v3.30 publish ok
40m 02s duration 5 published 2 updates
Claude Opus 5 (claude-opus-5) main agent
G1 Claude Sonnet 5 (claude-sonnet-5)
Items returned
7
Duration
18m 19s
Tool calls
9 WebFetch6 WebSearch28 bridge
Cited sources
none
G2 Claude Sonnet 5 (claude-sonnet-5)
Items returned
1
Duration
11m 37s
Tool calls
14 WebFetch17 WebSearch11 bridge
Cited sources
none
G3 Claude Sonnet 5 (claude-sonnet-5)
Items returned
2
Duration
12m 57s
Tool calls
12 WebFetch6 WebSearch24 bridge
Cited sources
none
truth-B1 Claude Opus 5 (claude-opus-5)
Items returned
20
Duration
14m 50s
Tool calls
not reported
Cited sources
none
truth-B2 Claude Sonnet 5 (claude-sonnet-5)
Items returned
11
Duration
12m 51s
Tool calls
not reported
Cited sources
none
truth-B3 Claude Opus 5 (claude-opus-5)
Items returned
20
Duration
14m 59s
Tool calls
not reported
Cited sources
none
truth-B4 Claude Sonnet 5 (claude-sonnet-5)
Items returned
20
Duration
15m 13s
Tool calls
not reported
Cited sources
none

Verification

unconfirmed CLEAN · waived: Single CLEAN at the 8-iteration cap. Iteration 8 (Sonnet) returned CLEAN with ze #1 NEEDS_FIXES · Opus 5 · t=3 e=2 a=3 #2 NEEDS_FIXES · Sonnet 5 · t=2 e=0 a=1 #3 NEEDS_FIXES · Opus 5 · t=4 e=0 a=3 #4 NEEDS_FIXES · Sonnet 5 · t=1 e=0 a=1 #5 NEEDS_FIXES · Opus 5 · t=3 e=1 a=2 #6 CLEAN · Sonnet 5 · t=0 e=0 a=0 #7 NEEDS_FIXES · Opus 5 · t=2 e=0 a=1 #8 CLEAN · Sonnet 5 · t=0 e=0 a=0

Deep dive

Sources changed (this run)

Edits this run made to sources/sources.json · promotions, demotions, new candidates, and fetch-method / category / reliability / url corrections (the run record's sources_changed[]). Paginated; 10 per page.

No source-list edits recorded for this run.

Coverage gaps (this run)

Sources this run's brief needed that returned no usable content via any documented recipe. Bridge-recovered or quiet-day sources do NOT appear here. (Distinct from the independent source-accessibility probe at the foot of this section, which probes all active sources regardless of what any run needed.)

Source (uncovered)URL triedMethod chainStatus / classWhat the agent did instead
github-advisoryhttps://github.com/advisories/GHSA-vmg2-rwj3-8rq2url (direct — github.com is egress-proxy-blocked)osv vuln CVE-2026-54363 (upstream HTTP 404 — not a package-ecosystem advisory)bsi-csaf WID-SEC-2026-2607 (succeeded, but carries ids and version ranges only, no per-CVE score or mechanism)jina200 content-not-renderednone available — the coordinating CERT document was used to confirm the identifiers and the fixed release, but it carries neither the CVSS scores nor the hardco

Verification findings · all iterations

Per-iteration finding detail. Each table is one verifier pass · what was flagged, how the main agent remediated it, and the outcome. Walking the tables top-to-bottom shows the verifier's debugging trail across iterations.

Iteration #1 NEEDS_FIXES · 8 findings (truth=3, editorial=2, advisory=3) · Claude Opus 5 · 12m 05s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
cves[] bound three of four identifiers to the wrong vulnerabilities — the entry mapped ascending CVE ids onto the source's TL;DR ordering, but the page prints an explicit mapping table giving 65879 = cves[] rebuilt from the source's mapping table with auth levels and CWEs corrected; title, summary and body updated; all four cves_seen.json record titles corre fixed
F3
claim-not-supported
cves[].cvss carried the discloser's own CVSS 4.0 self-scores where the same page states the CNA scores are authoritative — 'Where the two differ, the CNA's number is the one that travels with the CVE'cves[].cvss now carries the CNA figures (9.2 / 9.8 / 8.2 / 8.3); the body quotes the source's own statement on which score governs; sourcing_note updated. fixed
F5
missing-citation
A closing sentence asserted that this research stream's earlier findings 'have repeatedly reached CISA's exploited-vulnerabilities catalogue within weeks'. The entry's only source contains no mention Sentence removed and replaced with two claims the cited page does carry — the June 2026 icon-upload zero-day exploited in the wild, and the product's own 'most fixed
F8
needs-more-research
The cited page names a fifth identifier against the same version range — CVE-2026-65876, an unauthenticated SQL injection through the loadMoreArticles catid parameter, 9.2 Critical — which the entry oCVE-2026-65876 added to cves[], to the body with the discloser's caveat that it neither reported nor tested it, and to cves_seen.json. fixed
F4
hallucinated-fact
The report and run record stated '61 operational + 11 strategic' against a 71-entry window; the window is 60 + 11. The priority-calibration table's store-wide and monthly rows had also been computed wCorrected to 60 + 11 throughout; the whole calibration table recomputed from disk via site/content_model.load_entry with zero parse failures across all 1,110 en fixed
F11
editorial-advisory
affected_products doubled the vendor into the product name ('Marimo Marimo', 'WordPress WordPress'), which matches neither source's naming and fragments the product surface against values the store alCorrected to 'Marimo Notebook' and 'WordPress'. fixed
F11
editorial-advisory
The correction carried entities: [] while the entry it updates links actor:knaithe-knyuan and tool:hermes-ai-agent, so it would not appear on either entity's timeline.Both keys added. fixed
F11
editorial-advisory
cves[].type: auth-bypass on a CVSS 10.0 pre-auth flaw whose impact Adobe records as arbitrary code execution and whose own title says the same — defensible as Adobe's weakness category, but it understtype changed to rce. fixed

Iteration #2 NEEDS_FIXES · 3 findings (truth=2, editorial=0, advisory=1) · Claude Sonnet 5 · 10m 52s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
Whitespace-level quote-fidelity defect in the evidence[] quote and its body echo: 'read the entire database , password hashes included' carries a space before the comma that the live page does not havQuote corrected in evidence[] and in the body; a whitespace-faithful extraction (tag stripping that does not insert spaces) written to work/<run-id>/txt.spb.fai fixed
F4
hallucinated-fact
The calibration table's prior-window row (58 operational / 13 high / 22.4 %) could not be reproduced under any plausible window definition.Root cause was a date-only filter that spilled across both audit boundaries. Recomputed on the previous audit's actual started-timestamp bounds: 43 operational fixed
F11
editorial-advisory
The actor:knaithe-knyuan record's summary carried the same understated confirmed-impact scope this run was correcting in the entry.Summary corrected to the source's full four-CVE confirmed-impact list, with the correction attributed. fixed

Iteration #3 NEEDS_FIXES · 6 findings (truth=4, editorial=0, advisory=3) · Claude Opus 5 · 12m 38s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
The report still carried the 'database , password' whitespace defect that iteration 2 fixed only in the entry, and rendered the CNA-score quote with a straight apostrophe where the page has a curly onBoth corrected in the report. fixed
F4
hallucinated-fact
The watch-items row still carried the pre-correction action-density figures (0.87 against 0.48) while the report's own finding 3 had already been corrected to 0.53 -> 0.88; disk gives 53/60 and 23/43.Watch-items row corrected to 0.88 against 0.53. fixed
F4
hallucinated-fact
The entry dated its Unit 42 primary 2026-07-31, which is the page's article:modified_time; the dateline, article:published_time and JSON-LD datePublished all give 2026-07-30, as do the entry it updatesources[].date, event_date and the inline citation date corrected to 2026-07-30. fixed
F3
claim-not-supported
Factual-error row 2 attributed a phrase to the 2026-07-21 entry that is verbatim only in the W30 weekly.Row rewritten to quote each entry's own wording and to say which carries which. fixed
F11
editorial-advisory
The v3.30 entry said 'Five changes' after a sixth was added.Corrected to six. fixed
F11
editorial-advisory
Carried the `rce` tag although none of the five flaws is remote code execution.Tag removed. fixed

Iteration #4 NEEDS_FIXES · 2 findings (truth=1, editorial=0, advisory=1) · Claude Sonnet 5 · 9m 58s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F3
claim-not-supported
sourcing_note still said 'The CVSS figures are the discloser's own CVSS 4.0 assessments as printed in the advisory', contradicting the cves[] block that iteration 1's own F3 remediation had switched tsourcing_note rewritten to state that cves[] carries the CNA figures on the discloser's own instruction, and to name the discloser's lower self-scores as the bo fixed
F11
editorial-advisory
Factual-error row 2 said the W30 weekly 'repeats the phrasing verbatim' where the weekly's wording is a close paraphrase rather than an exact match.Replaced with the weekly's own verbatim phrase. fixed

Iteration #5 NEEDS_FIXES · 4 findings (truth=3, editorial=1, advisory=2) · Claude Opus 5 · 13m 47s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F3
claim-not-supported
CVE-2026-65879's 9.8 is attributed to the Joomla CNA, but that record's CNA metrics block is empty and the 9.8 is a CISA-ADP CVSS 3.1 score sitting among four CVSS 4.0 figures — and the body then rankThe cross-scale ranking was removed, but the accompanying sourcing_note wrongly asserted that no authority reachable this run could settle the scales — an asser superseded
F4
hallucinated-fact
cves[].cvss carried 9.8 for CVE-2026-39987 from Unit 42's table; the owning marimo GHSA/CNA record gives CVSS 4.0 9.3 (AV:N/AC:L/AT:N/PR:N/UI:N/VC:H/VI:H/VA:H).Corrected to 9.3 with the vector and CWE-306 read from the owning record; cves_seen.json title and primary_source_url repointed to the GHSA. fixed
F4
hallucinated-fact
The entry said the fixed version was not stated and its action told readers to treat exposure as a compromise-assessment target 'rather than a patching task'. CVE-2026-39987 was fixed in marimo 0.23.0fixed: set to 0.23.0 with the KEV listing and the store's own prior coverage named; status[] gains cisa-kev and patch-available; the action now reads patch-then fixed
F3
claim-not-supported
The correction quoted the 07-21 entry as 'autonomously rediscovering and weaponising the already-patched'; that entry's wording is 'to autonomously rediscover and weaponise the already-patched'.Quote corrected to the entry's own wording. fixed

Iteration #7 NEEDS_FIXES · 3 findings (truth=2, editorial=0, advisory=1) · Claude Opus 5 · 16m 37s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
sourcing_note claimed 'no authority reachable from this run could confirm which scale each belongs to'. False: tools/fetch_source.py url https://cveawg.mitre.org/api/cve/<id> returns every owning recoQueried all five records directly and confirmed: Joomla CNA cvssV4_0 9.2 / 8.2 / 8.3 / 9.2 on -65766 / -65877 / -65878 / -65876, and cna.metrics null on -65879 fixed
F4
hallucinated-fact
The 9.8 on CVE-2026-65879 was attributed to the Joomla CNA in the summary, sourcing_note and cves[], and propagated to state/cves_seen.json and the audit report. The CNA assigned that identifier no meCNA attribution removed everywhere it appeared; cves[] gains a cvss_note recording that the figure is CISA-ADP CVSS 3.1 and must not be ranked against its CVSS fixed
F11
editorial-advisory
Iterations 3 and 5 declare more findings than their findings[] arrays list; the declared counts match the verifiers' own on-disk reports, so the arrays were compressed in transcription rather than theLeft as-is and disclosed here: the per-iteration verification.iterN.findings.yaml files in work/ are the complete record, and the counts are the ones the verifi skipped

Verification & coverage notes

The run record's narrative body, verbatim. This is where the run accounts for its own judgement calls — every borderline drop and judged-not-relevant item with its reason, dedup decisions, single-source items and their carve-outs, contradictions, and per-source coverage gaps — so nothing the run considered disappears silently.

Verification & coverage notesrun record body

2026-08-02T1309Z-audit · audit · Opus 5 · window 168 h · 5 entries published

Audit run — 2026-08-02

Weekly quality audit over the window 2026-07-26T13:08:25Z → 2026-08-02T13:09:58Z (~168 h), anchored on the previous audit record. Scope: 71 published entries (60 operational + 11 W30 weekly strategic, of which 9 were the previous audit's own recoveries) across 10 run records. Full report: docs/audits/2026-08-02-weekly-quality-audit.md.

Verification & coverage notes

Soundness — 58 of 71 entries verified clean (81.7 %), against 56 % last window, on a genuine 2:2 Opus/Sonnet split and with the per-citation adjacency sweep applied exhaustively rather than as a sample. 162 primary and authority URLs fetched across the four passes. Three factual errors and ten imprecisions; six defects reach a machine surface.

The previous window's dominant defect class is materially reduced. Attribution defects — a true fact cited to a source that does not carry it — ran at 11 in 57 entries last window. This window they run at 4 in 71, and the two Opus batches covering 40 operational entries found none at all. Both surviving instances sit in the weekly strategic set.

Three factual errors, all corrected by new update entries rather than in-place edits. An inverted account of the Searchlight Cyber GPT5.6 experiment (an original pre-auth discovery reported as a rediscovery of an already-patched chain, a framing the store contradicted three days earlier in its own WP2Shell entry); a fabricated quotation attributed to Unit 42 that also dropped 11 confirmed compromises from the campaign's impact count; and a Joomla wave entry that records a constrained anonymous file-write as remote code execution and credits a disclosure to a research team whose own page says it did not make it. The first two are published as corrections here; the third is documented, because the field it corrupts is not inside the enumerated repair class and the audit declined to widen that class on its own authority — see the report's recommendation 3.

Two in-place repairs, both inside the logged immutability-exception class (wrong ATT&CK id, the 2026-07-11 precedent): 2026-07-26/ifage-geneva-dragonforce-data-published-student-records and 2026-07-31/exfilsquad-uk-department-for-education-pnld-breach each carried a technique id no cited source supports, and each retains a supported mapping after removal. Logged in .claude/memory/entry-immutability-exceptions.md.

Completeness — three genuine gaps recovered, one confirmed but unpublishable, several correctly-droppable. The re-sweeps re-researched the window independently: G1 across KEV, national CERTs and vendor PSIRTs, G2 across incidents and regulators with the watch-item re-checks, G3 across 36 research publishers. The recoveries are SP Page Builder for Joomla, Adobe Campaign Classic, and the Phoenix Contact CHARX charging controllers. All three in-window CISA KEV additions were independently re-enumerated and each was already covered — the KEV discipline shipped last week holds with no counter-example.

One confirmed gap could not be published and is recorded as such. Gladinet CentreStack carries six new CVEs including a hardcoded cryptographic key that forges domain-admin tokens, on a product whose identical bug class was mass-exploited in 2025 — a strong candidate on mechanics alone. Its per-CVE authority is a set of GitHub advisories, and github.com is egress-blocked for direct fetches while the reader returned only page chrome. The coordinating national-CERT document confirms the identifiers and the fixed release but carries neither the scores nor the mechanism, and the mechanism is the whole reason the item clears the bar. Publishing it would have meant transcribing a score from a research agent's summary rather than from the record that owns it, so it was not published. Carried as a watch item.

The SP Page Builder recovery is deliberately a new entry rather than an update, and the gate flags the shared Joomla-extension-wave tag for exactly the right reason, so it is worth answering in full. It shares that tag with the 2026-07-26 Balbooa Gridbox entry and the 2026-08-01 Aimy Captcha entry, but it is a different product from a different vendor with its own four identifiers, its own affected estate and its own fixed release, and neither existing entry tells a defender running SP Page Builder anything actionable. Treating each disclosure in this stream as a delta on the last one is precisely the failure this audit root-caused and fixed in v3.30: the consolidation rule is about a campaign recurring, not about independent per-product disclosures arriving from one publisher.

Correctly-droppable borderlines, documented so the completeness judgement is auditable: Progress MOVEit Transfer 2026.0.3 (four flaws, top score 7.5, the auth-enforcement one adjacent-network and high-complexity, no exploitation — two national CERTs escalating is a CERT's own risk framing, not exploitation evidence, and the 2026-07-31 fire's own drop reasoning stands); Apache Traffic Server's 38-CVE batch (unauthenticated scores are high but there is no exploitation, no public exploit, and no named mechanic forcing a timeline ahead of the patch cycle); the Xen Project's ten advisories (guest-to-host preconditions, routine XSA cadence, scoring not yet published at release); and CrowdStrike's Astaroth WhatsApp Web spambot (Brazil and LATAM victimology, carried only for a transferable trusted-platform-automation technique).

Watch items answered. The unconfirmed French leak-site wave remains uncorroborated after a full re-research — no Interior Ministry, ANSSI/CERT-FR or named-SDIS statement and no high-reliability journalism — and is closed on continued silence. Zscaler's TELESHIM Part 2 has not published and carries forward. The IFAGE and ANCPI stories have no in-window delta. No duplicate registry key was created from a Google unified cryptonym; the one that appeared was added as an alias on the existing actor record, which is the correct handling.

Machinery. No runaway runs (longest 2.69 h against a ~3 h threshold). Publish follow-through 10 of 10. Gap-derived windows self-healed across the one off-cadence gap with a 24 h floor throughout and no record-less day. The verifier rotation held perfectly in every fire — no same-model consecutive pair anywhere in the window. Four of nine fires reached a genuine two-model confirmed clean verdict, one published on the fail-open at the iteration cap, and four published with low residuals.

Two systemic findings drove prompt changes, and one drove none. The backup weekly fire of 2026-07-27 stood down correctly as a duplicate week — but only discovered the primary at the pre-push sync, after composing nine entries and running eight verification iterations, because the duplicate-week check runs at preflight only. A stream of Joomla-extension disclosures from one research publisher has now produced a recovered miss in three consecutive audits, because the consolidation rule written for long-running campaign activity was being applied to independent per-product vulnerability disclosures. Both are fixed in v3.30. The third finding — that the editorial half of the verification gate produced one priority-calibration and one action-item finding across roughly forty fires while both underlying metrics moved — is reported with instrumentation rather than a rule change, because this audit's own calibration review found no mis-prioritised entry to point at.

Reader pool. Seven live keys, 70 M tokens, refilled by the operator at 09:37–09:51Z today. The pool was exhausted for the 04:09Z fire, which is the third refill cycle in fifteen days; the standing recommendation is unchanged and resized in the report.

← Operations dashboard · run-record contract: docs/pipeline.md