ctipilot.ch

2026-08-03T0110Z-weekly

One pipeline fire, in full · weekly run of 2026-08-03 · sub-agent allocation and telemetry, per-iteration verification verdicts and findings, source-list edits, coverage gaps, bridge invocations — and the run's own verification & coverage notes: what was published, what was dropped at the borderline or judged not relevant (and why), single-source carve-outs, and contradictions. Rendered from runs/2026-08-03/2026-08-03T0110Z-weekly.md.

Run telemetry

2026-08-03T0110Z-weekly weekly prompt v3.30 publish ok
1h 09m duration 0 published 0 updates
Claude Opus 5 (claude-opus-5) main agent
W1 Claude Opus 5 (claude-opus-5)
Items returned
7
Duration
20m 19s
Tool calls
17 WebFetch11 WebSearch35 bridge
Cited sources
10 of 29 in slice
W2 Claude Opus 5 (claude-opus-5)
Items returned
6
Duration
21m 49s
Tool calls
12 WebFetch13 WebSearch24 bridge
Cited sources
12 of 33 in slice

Verification

unconfirmed CLEAN · waived: single CLEAN at iteration cap — iteration 8 is the hard cap and iteration 7 retu #? NEEDS_FIXES · Opus 5 · t=5 e=1 a=1 #? NEEDS_FIXES · Sonnet 5 · t=1 e=0 a=0 #? NEEDS_FIXES · Opus 5 · t=1 e=0 a=2 #? CLEAN · Sonnet 5 · t=0 e=0 a=0 #? NEEDS_FIXES · Opus 5 · t=1 e=0 a=1 #? NEEDS_FIXES · Sonnet 5 · t=1 e=0 a=0 #? NEEDS_FIXES · Opus 5 · t=1 e=0 a=2 #? CLEAN · Sonnet 5 · t=0 e=0 a=0

Deep dive

Entries published (this run)

Empty run · no new verified signal; only the run record was published (a healthy outcome).

Sources changed (this run)

Edits this run made to sources/sources.json · promotions, demotions, new candidates, and fetch-method / category / reliability / url corrections (the run record's sources_changed[]). Paginated; 10 per page.

No source-list edits recorded for this run.

Coverage gaps (this run)

Sources this run's brief needed that returned no usable content via any documented recipe. Bridge-recovered or quiet-day sources do NOT appear here. (Distinct from the independent source-accessibility probe at the foot of this section, which probes all active sources regardless of what any run needed.)

Source (uncovered)URL triedMethod chainStatus / classWhat the agent did instead
gtig-mandianthttps://cloud.google.com/blog/topics/threat-intelligence/rssrssjinawebfetchNone
RSS unreachable on both feed paths. The HTML listing returned titles without publication dates, so no in-window filtering was possible. The sub-agent attributed
A targeted WebSearch recovered the in-window GTIG naming-system post. Broader GTIG/Mandiant actor reporting for the week remains unassessed.
ccn-cert-eshttps://www.ccn-cert.cni.es/es/seguridad-al-dia/comunicados-ccn-cert.htmlbridge:urlbridge:jina404
404 on the comunicados listing path. Re-tested after the run: the reader reaches the host and returns a genuine 404 for this path, so this is a recipe gap rathe
none — this source has been an unresolved gap across several recent fires (recorded in the 2026-07-27, 2026-07-28 and 2026-07-29 intel runs, and attempted again
volexityhttps://www.volexity.com/blog/feed/rssjinaNone
feed returned non-JSON via the bridge feed parser. Re-tested after the run: the reader is reachable for this feed and rotates to a live credential, but returns
none — coverage gap.
proofpointhttps://www.proofpoint.com/us/rss/threat-insightrssNone
feed did not parse; no alternate feed path found inside the time box.
The in-window Proofpoint TA488 research was already covered operationally on 2026-07-31; broader TA-series reporting for the window is unassessed.
inside-it-chhttps://www.inside-it.ch/de/rssbridge:feedbridge:feed(/feed)bridge:jina403 transport-403
upstream 403 on both feed paths; the reader proxy relayed an upstream block.
substituted netzwoche.ch (direct, HTTP 200) — Swiss trade press on summer break since 2026-07-17, no policy items either way.

Bridge invocations (this run)

5 bridge calls this run · these are successful bridge fetches (separate from "Coverage gaps" above).

5 other
  • fetch_source.py (cisa page / feed / cisa-kev) ×1
  • fetch_source.py (ncsc-csh recent 40/60) ×1
  • fetch_source.py (bridge) ×1
  • fetch_source.py url ×1
  • fetch_source.py url (jina fallback) ×1

Verification findings · all iterations

Per-iteration finding detail. Each table is one verifier pass · what was flagged, how the main agent remediated it, and the outcome. Walking the tables top-to-bottom shows the verifier's debugging trail across iterations.

Iteration #? NEEDS_FIXES · 7 findings (truth=5, editorial=1, advisory=1) · Claude Opus 5 · 8m 43s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
completed/duration_seconds asserted a time that had not occurred and contradicted the run's own main.ended_at checkpoint.Re-captured the end checkpoint with a real date -u (02:20:15Z); both fields re-derived.
F4
hallucinated-fact
apple-security recorded as an unreached coverage gap; W1's findings file states it was reached and cleared, with an in-window 2026-07-27 Apple release wave.Rewritten to match W1; the release-wave fact preserved and only the per-release pages plus the ZDI review recorded as undrilled.
F4
hallucinated-fact
group-ib claimed as status -> active; the semantic diff shows it was already active on HEAD.Status-change claim dropped from sources_changed, the notes body and bridge_uses; the verified bridge recipe retained.
F14
quantifier-without-source
'second consecutive weekly' affected by the Sonnet classifier trip is refuted by two intervening weeklies, including this week's own primary.Reframed as intermittent, counterexample named, operator recommendation narrowed away from a pin rebind.
F14
quantifier-without-source
'ccn-cert-es unworked for a second consecutive run' understates a longer unresolved span.Quantifier replaced with the correct span in both places it appeared.
F10
missed-angle
Residual-coverage list omitted the in-window SBOM minimum-elements item, which the primary also does not carry.Added as the ninth residual bullet; heading count corrected.
F11
editorial-advisory
jina blast-radius arithmetic named three of six fetch_failures plus a non-member item.Fourth record named explicitly; the OT/ICS lab surface stated separately as a W1 coverage gap.

Iteration #? NEEDS_FIXES · 1 finding (truth=1, editorial=0, advisory=0) · Claude Sonnet 5 · 6m 32s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
The iteration-1 fix named inside-it-ch as a fourth jina key-exhaustion casualty, but its record shows the reader reached the host and relayed a genuine upstream 403 — a different failure mode. CorrectCount corrected to three of six; inside-it-ch given its own causal sentence.

Iteration #? NEEDS_FIXES · 3 findings (truth=1, editorial=0, advisory=2) · Claude Opus 5 · 9m 11s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
AI Act residual bullet claimed prohibitions, AI-literacy, transparency and penalties all apply from 2 August 2026; prohibitions and AI literacy have applied since 2 February 2025 and penalties since 2Bullet rewritten with correct dates, plus an explicit warning that findings.W2.yaml item 3 repeats the error.
F11
editorial-advisory
W2 sources_used omitted media.defense.gov and cyberresilienceact.eu, both cited by its own items.Both hosts added.
F11
editorial-advisory
Coverage-gaps bullet never named the seven OT/ICS research-lab feeds that went unread.All seven named with their failure mode.

Iteration #? NEEDS_FIXES · 2 findings (truth=1, editorial=0, advisory=1) · Claude Opus 5 · 10m 58s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
AI Act bullet bound Chapter VII governance and Chapter III Section 4 notified bodies to 2 August 2026; Article 113(b) applies both from 2 August 2025, the same tranche as the penalties the sentence alBullet rewritten to set out all four Article 113 tranches explicitly; heading changed from 'enforcement layer live' to 'general date of application reached'.
F11
editorial-advisory
NCSC-UK bullet called the blog a procurement requirement; it advocates to buyers rather than mandating.Reworded to a buying criterion NCSC is telling procurers to demand.

Iteration #? NEEDS_FIXES · 1 finding (truth=1, editorial=0, advisory=0) · Claude Sonnet 5 · 4m 23s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
Bullet claimed the omnibus amends Article 113 point (c) only; the run's own capture shows it also replaces point (a) — carving Article 5(1) sub-points out to 2 December 2026 — and adds a new point (d)Paraphrased timetable removed; the omnibus's own amending text now quoted from the capture and literal-substring checked (3 quotes, 0 failures), with an instruc

Iteration #? NEEDS_FIXES · 3 findings (truth=1, editorial=0, advisory=2) · Claude Opus 5 · 12m 01s

F-codeSectionItem · URL/quoteVerifier summaryRemediation · outcome
F4
hallucinated-fact
Claimed the jina key pool was exhausted and the rotation returned nothing; the run's own capture shows the rotation retrieving 170 KB of Official Journal text, and jina-usage reports 6 of 7 keys live jina-reader record removed from fetch_failures (recovered, so not an unrecovered failure); the three affected gap causes restated from evidence; the operator-ac
F11
editorial-advisory
source_health snapshot accounting said 'both' where the file gained three snapshots.All three named with timestamps.
F11
editorial-advisory
'the week's six agent incidents' — two of the six entries are corrections and one a detection mapping.Reworded to 'six agent-thread entries'.

Verification & coverage notes

The run record's narrative body, verbatim. This is where the run accounts for its own judgement calls — every borderline drop and judged-not-relevant item with its reason, dedup decisions, single-source items and their carve-outs, contradictions, and per-source coverage gaps — so nothing the run considered disappears silently.

Verification & coverage notesrun record body

2026-08-03T0110Z-weekly · weekly · Claude Opus 5 · 0 entries published

Verification & coverage notes

Disposition: duplicate-week. This backup fire stood down; the primary weekly for 2026-W31 published. No strategic entries ship from this run.

ISO week 2026-W31 (2026-07-27 00:00 UTC → 2026-08-02 24:00 UTC). Gap to the previous weekly run record = 7 days; the last weekly that published content before this week was 2026-07-26T2309Z-weekly, so window_days = 8.

Why the run executed in full before standing down. The Phase 0 duplicate-week guard ran at 01:10Z against a freshly fetched origin/main (then at c1f0ab5) and found no -weekly record carrying week: 2026-W31, so the backup proceeded correctly on the evidence available. The primary run 2026-08-02T2311Z-weekly had in fact completed at 00:06:31Z, about an hour before this fire started, but its record had not yet been promoted to main. It appeared at 6a04d5d when the pre-verifier guard re-ran at 02:10Z — exactly the race the v3.30 pre-verifier re-check was added to catch. The guard did its job: the stand-down was reached before the verifier loop rather than at the Phase 6 pre-push sync, saving the eight-iteration cost the 2026-07-27 stand-down paid.

What was withdrawn. Fifteen strategic entries had been composed and had passed the mechanical gate (check_run.py --pre-verify: 35 pass · 9 warn · 0 fail). They were deleted before commit. Two entity registrations made for them (policy:eu-ai-act-digital-omnibus-2026, report:intrinsec-enterprise-llm-threat-atlas-2026) were reverted with them, so entities_added is empty and the registry is unchanged from main. The research returns, triage, verified-quote ledger and fetched primaries are retained under work/2026-08-03T0110Z-weekly/ as the forensic surface and as the input to the residual list below.

Residual coverage — the nine in-window items this run verified that the primary's fifteen entries do not carry. Established by grepping the primary's published entry set on main. These are not duplicate coverage and are the substantive output of this fire; they are candidates for the next weekly or, where they have operational character, for an intel run:

  • CERT Intrinsec's two-part DFIR artefact map for autonomous coding agents (OpenCode 2026-07-27, OpenAI Codex 2026-07-31) — no match in the primary's set. Closes the responder-side gap left by the week's six agent-thread entries: where prompt history, per-session interaction logs and agent credentials sit on disk. Also a new collection-target exposure, since the same directory holds API keys and access tokens.
  • Intrinsec's Enterprise LLM Threat Atlas (2026-07-30, dual MITRE ATLAS + ATT&CK mapping) — no match. Its own risk ranking, tool/agent abuse and supply chain first, was independently corroborated by this week's incident record.
  • Group-IB's PAM-as-anti-forensics intrusion (2026-07-30) — no match. Trusted-third-party initial access, pam_rootok weaponised to impersonate low-privileged users as a deliberate forensic smokescreen, logging suppression and a self-unlinking miner.
  • NCSC UK on forensic observability for network devices (2026-07-29) — no match. Makes edge-appliance forensic capability a buying criterion NCSC is telling procurers to demand — it advocates rather than mandates — with an international reference architecture in progress. Pairs directly with the week's edge-exploitation record.
  • AI Act: the general date of application was reached on 2 August 2026, and Regulation (EU) 2026/1744 — the Digital Omnibus on AI — entered into force on 27 July — no match in the primary's set. The omnibus amends AI Act Article 113's third paragraph in three places, and the amending text (captured this run at work/2026-08-03T0110Z-weekly/pages/eurlex-omnibus.txt) reads: point (a) is replaced by “Chapters I and II shall apply from 2 February 2025, with the exception of Article 5(1), first subparagraph, points (ba) and (bb), and Article 5(1a) and (1b) which shall apply from 2 December 2026”; point (c) is replaced by “Chapter III, Sections 1, 2, and 3, with the exception of Article 6(5), shall apply from” 2 December 2027 for systems classified as high-risk under Article 6(2) and Annex III, and 2 August 2028 for those classified under Article 6(1) and Annex I; and a new point (d) is added, “Articles 102 to 110 shall apply from 27 July 2026”. A future entry must take every date from Article 113 as amended and quote it rather than paraphrase. This record's own attempts to summarise the timetable were wrong three separate times across six verification iterations — first on which duties attach to 2 August 2026, then on the 2 August 2025 tranche, then on the omnibus's own scope — and findings.W2.yaml item 3 carries the original error too. The relevant unamended tranches are in Article 113 itself, which is short.
  • CI Fortify joint OT-isolation guidance (CISA / ASD ACSC / NCSC UK / Canadian Centre, 2026-07-28) — no match. The obligation-side counterpart to the primary's own water-PLC entry.
  • Germany's NIS2 registration forbearance lapsing 31 July, with roughly 11,500 of ~29,500 entities registered per in-window trade reporting — the primary mentions NIS2 only in its looking-ahead entry, so the enforcement-phase transition is at most partially carried.
  • The updated SBOM minimum elements (CISA and international partners, published in-window; joint TLP:CLEAR guidance dated 2026-07-28 with the news announcement 2026-07-29) — no match. A grep of the primary's full entry set for SBOM and bill of materials returns nothing, including inside its own CRA guidance entry. It resets the baseline that CRA Annex I software-bill-of-materials duties and European procurement specifications converge on, extending scope to open-source, AI software and SaaS and adding component hash, licence, tool name and generation context as required elements.
  • The GTIG actor-naming change is partially covered — it appears inside the primary's open-source-supply-chain status entry rather than as its own item. Its registry-hygiene consequence is already handled: the 2026-07-27 stand-down added the SANDWORM RELIC alias to actor:sandworm.

Verification outcome — eight iterations, ten defects fixed, published on an unconfirmed CLEAN at the cap. The loop ran the full eight iterations against a run record and nothing else, and it was worth every one: it removed three fabricated or inverted telemetry claims (a completed timestamp that had not yet occurred, an apple-security coverage gap that inverted what the research agent actually reported, and a group-ib status change that never happened), two unsourced quantifiers, one missed residual item, and — twice over — substantive factual errors that would have propagated. The AI Act residual bullet was wrong three separate times on three different points before the paraphrase was abandoned in favour of quoting Article 113 directly. And iteration 7 refuted this run's own headline operator finding, showing the jina key pool was not exhausted at all. Iteration 4 returned CLEAN and iteration 5 refused to confirm it, finding a further defect — which is precisely the failure mode the double-CLEAN gate exists to catch, working as designed. The final CLEAN at iteration 8 is unconfirmed because iteration 7 was NEEDS_FIXES and the cap left no room for a confirmation pass; verification.confirmation_waived records it and check_run.py carries the corresponding WARN, which is a truthful fact about this run rather than a defect to suppress. Two process notes for the audit: every verifier iteration reported that its environment could not write report files, so all eight reports were transcribed by the main agent into work/2026-08-03T0110Z-weekly/verification.iter1.findings.yaml; and iteration 2 inadvertently re-ran source_health.py, which is not read-only.

Operator recommendation — the preflight guard cannot see a completed-but-unpromoted primary. This is the second consecutive weekly cycle disrupted by the same mechanism, and the residual risk after v3.30 is now precisely characterised: the Phase 0 guard greps origin/main only, so a primary that has finished its pipeline but whose auto-merge has not yet landed is invisible to it. Concrete fix for the next prompt revision — extend the Phase 0 guard to also grep the remote claude/** feature branches for the week label (git ls-remote --heads origin 'claude/*', then git grep each), and stand down on a match there too. That would have ended this fire at 01:11Z instead of 02:10Z. Recorded as a recommendation rather than executed here: a prompt edit requires a banner bump plus a CHANGELOG entry across all three banner-versioned prompts in one commit, which is not work a stood-down fire should land alongside zero entries.

Research sub-agent model pin overridden — the pinned Sonnet definition was blocked four times by the real-time cyber safeguard. Both W1 and W2 failed to spawn on the cti-research definition's Sonnet pin, twice each: the initial parallel spawn, and a retry after both spawn envelopes were rewritten as minimal on-disk briefs specifically to reduce the classifier surface. All four attempts returned the same safeguard error before the agent produced any output, which establishes the trip is on the pinned model rather than on the spawn message. Both were then re-spawned with an explicit model: opus override and completed normally, returning 7 and 6 items. Recorded as a deviation from the definition's pin under the documented classifier-block exception precedent. The trips are intermittent rather than consecutive, and the record should not be read as a failing pin. The 2026-07-26 W30 weekly did lose its W1 domain entirely to the same safeguard, but the two weekly fires in between both ran all four Sonnet-pinned research spawns successfully — 2026-07-27T0110Z-weekly (W1 5 items, W2 2 items, both returned) and, decisively, the 2026-08-02T2311Z-weekly primary for this very week roughly two hours before this fire (W1 7 items, W2 2 items, both returned on Sonnet 5). So the pin works most of the time and failed four times in a row on this one fire. The override recovered both domains and no coverage was lost. The operator decision this supports is therefore narrower than a rebind: treat an all-attempts-blocked spawn as a known intermittent condition with the Opus override as the documented fallback, and consider the Cyber Verification Program named in the error if the frequency rises. Rebinding cti-research to Opus on the strength of this fire alone would be an overcorrection that the same week's primary refutes.

Operator action — one jina reader credential is exhausted; the pool is not. This run's notes initially claimed the whole key pool was exhausted and that this was the root cause of most of its coverage gaps. The verification loop refuted that, and the correction matters because the wrong version would have sent an operator to top up a pool that does not need it while leaving the real causes unexamined. What is true: the primary credential …hxnTiF returns HTTP 402 — balance exhausted, so every reader call wastes a round-trip on it before rotating. What is false: that the rotation fails. tools/fetch_source.py jina-usage reports 7 keys, 6 live, and roughly 50.8 million tokens of remaining balance, and this run's own capture proves the rotation working — pages/eurlex-omnibus.err records rotating to the next credential followed by # fetched via jina reader fallback, and pages/eurlex-omnibus.txt holds the 170 KB Official Journal text it retrieved. The GTIG/Mandiant, Volexity, CCN-CERT and OT/ICS-lab gaps are real, but they are not attributable to an unusable reader: re-testing after the run showed the Volexity feed rotating to a live credential and returning zero items, and the CCN-CERT listing path returning a genuine 404 through the reader. Their true causes were not established this run and are the thing to investigate. The operator action is narrow: remove or top up …hxnTiF so the wasted round-trip stops.

State changes retained. The source-lifecycle work is independent of the withdrawn entries and is kept: the certvde promotion (a state-digest-driven duty that would otherwise go unactioned for another week), the group-ib bridge recipe confirmed and recorded (its status was already active and is unchanged), the NCSC-NL news-feed recipe note, and one new candidate. source_health.py was run and its snapshot retained — 173/173 probed, and the single needs-demote flag it raised was against this run's own new candidate record, whose URL was corrected from /news (404) to /blog (200) and re-probed clean in the same run, per the standing repair order. The sweep now ends with zero unsolved sources. (Iteration 2 of the verification loop re-ran source_health.py while cross-checking this claim, so state/source_health.json carries one additional probe snapshot beyond this run's own; all three 2026-08-03 snapshots (01:59:42Z, 02:02:11Z and 02:26:56Z) are genuine probes from this session, and the aggregate figures of the last two are identical.)

ATT&CK pin freshness (weekly maintenance duty): tools/attack_data.py --check → up to date, local v19.1 == upstream latest v19.1. No update required. Two revoked ids were caught by the gate in the withdrawn entries and are worth recording for future composition: T1562.001 is revoked in favour of T1685, and T1070.002 in favour of T1685.006.

Closed-source intake: intel/ carries only README (no in-window drops) — no W3 spawned.

  • Watchlist: products checked=0, hits=0; suppliers checked=0, hits=0 (none configured — the sweep is a no-op).
  • Coverage gaps: jina reader (the primary credential is exhausted and wastes a round-trip per call, but the pool rotates to six live keys — NOT the cause of the gaps below, contrary to this run's initial reading); gtig-mandiant (feed unreachable, listing undated; broader actor reporting unassessed); volexity, proofpoint, socket-dev (feed parse failures); the whole OT/ICS research-lab surface — dragos, claroty-team82, nozomi, harfanglab, withsecure-labs, trellix, orange-cyberdefense — all returned non-JSON from the bridge feed parser and were recorded by the sub-agent as unable to escalate to the reader on an exhausted key pool — a diagnosis the verification loop later refuted, so the surface went entirely unread this run for a cause that remains unestablished and energy, water and transport are on this deployment's sector list; ccn-cert-es (404 on the listing path; an unresolved gap across several recent fires, recipe gap); apple-security (NOT a gap — the rotation-priority source was reached and cleared: the Apple security-releases index was fetched via the bridge and shows an in-window release wave dated 2026-07-27 covering iOS/iPadOS 26.6, macOS Tahoe 26.6, macOS Sequoia 15.7.8, macOS Sonoma 14.8.8, tvOS 26.6, watchOS 26.6, visionOS 26.6 and Safari 26.6. What was left undrilled was the per-release security-content pages and ZDI's 2026-07-30 July Apple update review, both vulnerability-triage work rather than the research horizon's remit); inside-it-ch (403 both paths, substituted netzwoche.ch); coe-cybercrime, finma, bakom-ofcom, europol, cnil-fr, edpb, ico-uk, govcert-at, cert-at (reached, no in-window policy content — the Swiss policy surface was genuinely quiet, with BACS on its summer awareness series, the Bundesrat in recess and FINMA's last cyber item dated 9 July).
  • Essential-coverage: all essential sources in both slices were attempted; no essential miss beyond the gaps recorded above.
  • Date trap recorded for future runs: the Sekoia blog listing renders migration dates rather than publication dates after the blog.sekoia.io → sekoia.com move — one article displayed 2026-07-29 on the listing while its JSON-LD datePublished, byline and independent coverage all say 1–2 July. Corroborate Sekoia dates externally before any in-window decision.
  • One self-caught process error in W1, no impact on output: an Apple advisory URL was constructed by inference and resolved to an unrelated 2025 page. It was never cited.

← Operations dashboard · run-record contract: docs/pipeline.md