The intelligence pipeline, data model (v4, normative)
This document is the single normative specification of the content model:
per-finding entries (one living record per finding, with a dated
changelog), the entity registry, and per-run run records. Every
producer (the run prompts, the migration tools) and every consumer
(site/build.py, tools/check_run.py, the verifier agents) implements
exactly this contract. If code and this document disagree, this document
wins and the code is the bug.
v4.0 (2026-08-27) in one paragraph. Two routines remain; the intel run
(prompts/cti-run.md, any cadence) and the quality audit
(prompts/quality-audit.md); the weekly strategic routine is retired, its
/weekly/ pages are gone, and its entries were deleted outright on
2026-08-29 (with them went the horizon axis, the weekly_section field
and the synthesis/outlook kinds). A finding has exactly one entry for its
whole life: developments, corrections and improvements are appended to that
entry as timestamped updates[] changelog records with matching body
sections, updated_at floats the entry back to the top of the live brief,
and the old update_of second-entry mechanism is retired (the historical
update entries were folded into their roots by tools/migrate_updates.py).
See § Entry lifecycle.
Why this model exists
v2 produced one monolithic Markdown brief per day. That capped intelligence
latency at the routine cadence: something disclosed at 09:00 waited for the
next morning's fire. v3 turns the product into a pipeline: the run
prompt (prompts/cti-run.md) can fire any number of times per day, each
fire publishes only the new verified signal since the previous fire as
individual entry files, and the "brief" is a rendering over a reader-
chosen time window (default: last 24 h). Because every finding is a
standalone file with complete structured metadata, downstream automation
(notification hooks on priority: critical, sector feeds, entity timelines,
trend analytics) consumes the pipeline directly, no Markdown scraping.
Two properties are non-negotiable and carried over from v2 unchanged:
- More runs must not mean more content. Entry volume is governed by a strict relevance/actionability gate (see § Relevance discipline), not by a numeric target or ceiling: the rolling-24-hour window carries exactly the entries that clear that gate, however few or many that is. Firing more often changes latency, never volume, dedup guarantees a re-scan of the same window republishes only the new delta. A run that finds nothing new publishes nothing but its run record, that is a healthy outcome.
- Everything published passed the same gates: two-source verification, fake-news guard, URL truth, taxonomy validation, the mechanical self-check, and the adversarial verifier loop.
Repository layout
entries/YYYY-MM-DD/<slug>.md # one finding per file; folder = UTC date of discovered_at
entries/README.md # short contract pointer (this file is normative)
entities/registry.yaml # global entity registry: actors, campaigns, malware, tools, incidents, reports
entities/README.md # registry contract pointer
attack/enterprise-attack.json # pinned MITRE ATT&CK release (see § The ATT&CK layer)
attack/README.md # ATT&CK dataset contract + update procedure
runs/YYYY-MM-DD/<run-id>.md # one run record per fire: frontmatter = telemetry, body = verification notes
runs/README.md # run-record contract pointer
state/cves_seen.json # flat fast-lookup CVE index (kept from v2)
state/source_health.json # source accessibility snapshots (kept from v2)
state/warning_acknowledgments.json # audit-reviewed ledger of settled-history check_run.py WARNs (v3.28)
sources/sources.json # curated source list (kept from v2)
work/<run-id>/ # per-run forensic artefacts (kept from v2)
site/content_model.py # THE shared parser/loader/validator for entries, registry, runs
Retired from v2 (no backwards compatibility): briefs/ (migrated into
entries/ by tools/migrate_briefs.py, then deleted),
state/covered_items.json (coverage is now derived by scanning entries/),
state/deep_dive_history.json (derived from entries with deep_dive: true),
state/run_log.json (replaced by runs/).
Run identity; multiple runs per day
run_id = <YYYY-MM-DD>T<HHMM>Z-<fire> fire ∈ { intel, audit } (weekly: legacy records only)
e.g. 2026-07-03T0412Z-intel runs/2026-07-03/2026-07-03T0412Z-intel.md
- UTC, minute precision. Lexically sortable. Deterministic: a same-minute
retry computes the same
run_idand updates the same record in place (idempotent retry, same rationale as v2's sha8 scheme). - The suffix names the fire type and the frontmatter
kindcarries the same value. A v4+ fire isinteloraudit(content_model.ACTIVE_RUN_KINDS);weeklyremains in the validated vocabulary only so the historical weekly run records keep validating; the weekly routine was retired in v4.0 andcheck_run.pyFAILs a new record carrying it. (v3.24 made audit fireskind: audit; before that they carriedkind: intelwith the-auditsuffix as the only discriminator.) Consumers distinguish run types bykind; the run-id suffix stays as the human-readable mirror. work/<run-id>/uses the identical string.- Migrated v2 runs keep their historical ids (
2026-07-03-04ba8283,2026-W26-b78503e7) as filenames underruns/<date>/; only new runs use the timestamped form. Consumers treatrun_idas an opaque sortable string and read timing from the frontmatter, never by parsing the id.
Entry files, the atomic intelligence unit
Path: entries/<YYYY-MM-DD>/<slug>.md where the folder date is the UTC date
of discovered_at and <slug> is kebab-case, [a-z0-9-], ≤ 60 chars,
unique within the day. The entry id is path-derived:
<YYYY-MM-DD>/<slug> (e.g. 2026-07-03/coolify-cve-2026-34038-rce).
There is no id frontmatter field; the path is the identity.
One entry per finding, for the finding's whole life. The entry is the
single living record of the finding: a later run (an intel fire or the
quality audit) that learns something new about it, finds an error in it, or
can make it more precise edits this same file, appending a dated
updates[] changelog record and a matching ## <Type> — <at> body section,
bringing the frontmatter to the current truth, and (for a material new
development, a type: update record) moving updated_at.
There is never a second entry for the same finding, and there is never a
silent edit: every change is a record the reader and the gate can see.
Three things never change once published, the entry id/path,
discovered_at (first publication) and run_id (the originating fire).
Full contract: § Entry lifecycle.
Frontmatter, strict YAML subset
The frontmatter block is parsed by site/content_model.py (stdlib-only,
no PyYAML). It accepts a strict subset of YAML: 2-space indentation, no
tabs, no flow style except [] / inline [a, b] lists of plain scalars,
- list items (scalar or single-level mapping), one level of nested
mapping for block fields, >/| block scalars, null/true/false
literals, full-line comments only. Producers MUST stay inside this subset;
tools/check_run.py fails the commit on anything the parser rejects.
---
schema: 1
kind: vulnerability # see § Kinds
title: "CVE-2026-34038 — Coolify: authenticated command injection to RCE (CVSS 9.9)"
headline: "Coolify ships an emergency fix for a CVSS 9.9 authenticated command-injection RCE"
summary: >
Self-contained 1–3 sentence summary naming products, regions and CVEs.
This is the TL;DR bullet body, the RSS description, and the notification
text — a reader who sees ONLY this must know what is affected and why it
matters.
discovered_at: "2026-07-03T04:21:09Z" # UTC moment the finding was FIRST published — never changes
updated_at: null # == `at` of the last NON-INTERNAL `type: update` record; null while there is
# none (corrections, improvements and internal records never move it).
# max(discovered_at, updated_at) is the entry's activity moment, the live
# brief's sort key
event_date: "2026-07-02" # date of the underlying event / primary publication
run_id: 2026-07-03T0412Z-intel # the ORIGINATING fire — never changes; updating runs appear in updates[]
priority: high # critical | high | notable | routine — see § Priority
immediate_action: null # or the block below — presence ⇔ priority: critical
# immediate_action:
# title: "Patch Coolify to ≥ v4.0.0-beta.469 now"
# action: >
# One-to-three sentences: the specific time-critical defender action
# (emergency patch, isolation, credential rotation, emergency rule).
tags: [vulnerabilities, rce, patch-available] # taxonomy themes ∪ nexus
regions: [global] # taxonomy regions
sectors: [technology] # taxonomy sectors (may be empty)
entities: [] # registry keys, e.g. [actor:shinyhunters, campaign:fortibleed]
techniques: [] # MITRE ATT&CK ids the sources support (T####[.###]) — the
# CANONICAL mapping surface (active ids per the pinned
# attack/enterprise-attack.json); every id must name a behavior
# a cited source supports and the body does not contradict
# (inline T-ids only where essential); [] when the entry maps none
affected_products: [] # official product names ("Vendor Product" strings — what an
# alert or asset inventory would name); [] when not
# product-specific
cves: # [] when the entry carries no CVE
- id: CVE-2026-34038
cvss: "9.9" # string; "n/a" when unassigned
epss: null # FIRST.org EPSS PROBABILITY as a quoted decimal in [0, 1]
# ("0.0047"), never a percentage ("0.47"), never the
# percentile ("55.85"), never a suffix ("0.27 (EUVD)") —
# take the API's `epss` field, not its `percentile` field,
# and put provenance in sources[]/sourcing_note. null when
# not looked up. `check_run.py` `cve-epss` enforces the range.
type: rce # taxonomy cve_types
vector: zero-click # taxonomy cve_vectors
auth: post-auth # taxonomy cve_auth
status: [patch-available] # taxonomy cve_status
affected: "≤ 4.0.0-beta.462"
fixed: "4.0.0-beta.469"
sources:
- url: "https://github.com/coollabsio/coolify/security/advisories/GHSA-qqrq-r9h4-x6wp"
publisher: "coollabsio GHSA"
date: "2026-07-02"
role: primary # primary | corroborating — first source is the most primary
closed_sources: [] # [{title, provider, date, ref}] — intel/ drop citations, never URLs (no TLP gate)
evidence: # quotes binding claims to fetched sources — quote is ALWAYS English (v4.2)
- quote: "An authenticated remote command injection vulnerability (CWE-78) in Coolify…"
publisher: "coollabsio GHSA"
source_url: "https://github.com/coollabsio/coolify/security/advisories/GHSA-qqrq-r9h4-x6wp"
# the page the quote is from — expected on every web-sourced
# quote (the gate's quote-literal check searches it first)
- quote: "first reported a data leak on 7 August (translated from German)" # non-English source:
original: "erstmals am 7. August einen Datenabfluss gemeldet" # quote = marked English
publisher: "Der Tagesspiegel" # translation; original =
# verbatim source text (the
# verifier greps THIS)
verification: multi-source # multi-source | single-source | single-source-national-cert |
# single-source-victim | contradicted
sourcing_note: null # human clause, e.g. "victim-own SEC 8-K disclosure carve-out"
confidence: high # high | medium | low
references: [] # entry ids this entry builds on (a distinct finding that shares a CVE
# with an older entry MUST list it here — the explicit, gate-checked
# statement that it is not a duplicate; see § Dedup)
deep_dive: false # true ⇒ this entry IS the deep-dive treatment
deep_dive_category: null # taxonomy-free rotation slug when deep_dive: true (see prompt)
org_triage: null # or {category: P1, rationale: "…"} on triage-kind entries when a scheme is defined
classification: # NATO Admiralty code. Vulnerability is a triage kind: it carries
reliability: A # org_triage + classification: null ONLY while the profile configures a
credibility: 1 # triage scheme; with none configured (the shipped profile) it carries
# this block like every other kind, e.g. {reliability: B, credibility: 2}
# reliability A–F (of the sourcing) + credibility 1–6 (of the item),
# assessed independently — config/org-profile.yaml `classification:`.
watchlist_hit: false # true only when inclusion was driven by an org-profile watchlist match
actions: [] # imperative, entry-specific defender actions (strings) — feed § Action Items
updates: [] # the changelog — append-only, oldest first (§ Entry lifecycle):
# updates:
# - at: "2026-07-05T04:40:12Z" # UTC; > discovered_at and > the previous record's at
# run_id: 2026-07-05T0410Z-intel # the fire that made the change
# type: update # update | correction | improvement
# summary: > # 1–3 self-contained sentences: what changed and why —
# CISA added CVE-2026-34038 to KEV on 2026-07-04; status moved to exploited. # the timeline row
# fields: [cves, priority, summary] # optional: frontmatter fields changed in place ("body" for
# # an edit to the main analysis)
# merged_from: null # migration provenance only (a v3 update_of entry folded here)
migrated_from: null # v2 provenance (briefs/YYYY-MM-DD.md) — migration tool only
---
Body: the full analysis in Markdown, followed — when the entry has been
updated — by one `## <Type> — <at>` section per `updates[]` record, in the
same order (§ Entry lifecycle). Inline source links at the point of
claim (`([Publisher, YYYY-MM-DD](URL))`), defender takeaway, detection and
hardening concepts — ATT&CK mappings live in `techniques[]`, and an inline
T-id appears in prose only where essential (deep-dive kill chains, a
mapping that is itself the finding) — the same technical register and
depth as a v2 brief item, described as
observable behavior (telemetry classes in vendor-neutral terms; platform
artifacts as examples) so a human analyst or an automated triage agent
can match an alert against it. Threat/incident/research bodies close with
`**Defender takeaway:**` and, where the cited mechanism supports a
benign-lookalike discriminator, a `**Triage:**` line. Deep-dive entries
carry the complete deep-dive narrative (Background paragraph, kill chain,
hunt concepts, mitigation). No IOCs, no rule code, no vanity metrics,
English only.
Field semantics and hard rules
headline: bold-lead TL;DR headline, ≤ 120 chars (the gate WARNs above 120 on a new entry; 160 is the parser's hard limit), no trailing period.summary: the load-bearing standalone digest. Never empty.discovered_at: the moment this pipeline first published the finding, set once, never backdated, never changed by an update. The folder date MUST equal its UTC date.updated_at/updates[]: the entry's changelog (§ Entry lifecycle).updates[]is append-only, oldest first;updated_atMUST equal theatof the last non-internaltype: updaterecord (null while there is none: corrections, improvements and internal records never move it). The entry's activity moment ismax(discovered_at, updated_at)(content_model.entry_activity_ts): it orders the live brief, the briefbook and the feeds, so an update floats the entry back to the top.event_date: recency anchor of the underlying event (primary-source publication date). Drives staleness checks;discovered_atdrives windows.entities: every value MUST resolve to a key inentities/registry.yaml. New entities are added to the registry in the same commit. Never invent a second key for a known entity, check aliases.cves[]: one record per CVE, always withtype/vector/auth/statusfrom the taxonomy. Multi-CVE items carry one record per CVE (the v2 "per-CVE breakdown" is now structural). Axis semantics:vectorencodes the victim-interaction requirement (zero-click= attacker- initiated, no victim interaction; independent of auth state),authencodes the authentication precondition; an authenticated, no-interaction bug is correctlyvector: zero-click+auth: post-auth.techniques[]: the entry's MITRE ATT&CK technique ids, validated againstT####/T####.###(format, FAIL) and against the pinned ATT&CK datasetattack/enterprise-attack.json(existence + lifecycle: FAIL on the run's own entries, WARN store-wide; see § The ATT&CK layer). This is the canonical mapping surface: the machine retrieval layer for alert-triage consumers (given an alert mapped to a technique, the matching entries are a field lookup), and the sole input to the derived entity/CVE TTP profiles, the/attack/matrix and the Navigator-layer exports, a technique missing here is invisible to all of them. Use active ids only (revoked ids resolve forward viarevoked_by, but new entries reference survivors). Every id must name a behavior a cited source supports and the body does not contradict: the mapping follows the sources, not the length of the prose, so a short entry maps a source-stated chain as completely as a long one. Inline T-ids in the body appear only where essential, and a bare ID list in prose is a defect. An id no cited source supports is a hallucination. Entries that predate this field (the migrated/early-v3 tail) carry their mappings as in-prose T-ids only; consumers derive their effective set viacontent_model.entry_technique_ids(frontmatter ∪ dataset-known prose ids); the tail is not bulk-rewritten, so that derivation path stays.affected_products[]: official vendor product names as plain strings ("Citrix NetScaler ADC","Adobe ColdFusion"), the names an alert, asset inventory, or CMDB would carry. Generalizes the CVE-onlyaffected/fixedversion fields to campaign/threat entries; empty when the entry is not product-specific.sources[]: ≥ 1 unlessclosed_sourcesis non-empty. First entry is the most primary (vendor PSIRT > vendor research blog > research-lab post > regulator filing > victim disclosure > national CERT/CSIRT > MITRE/NVD > ENISA EUVD > news). Homepage / listing / category / per-CVE-database URLs are FAIL-blocked (same pattern list as v2, intools/check_run.py).evidence[]: required when any CVEstatusincludesexploitedand on everyimmediate_actionentry. Each quote must be a verbatim substring of a page fetched this run, attributed to a listed source's publisher. English-only (v4.2, operator directive 2026-08-28): a non-English source is quoted as a marked English translation inquote("… (translated from German)"), with the verbatim source-language text in the optionaloriginalfield, the verification surface. The renderer shows onlyquote; reader-facing prose quotes the same way and never carries untranslated non-English text.verification/sourcing_note:single-source*values replace the v2[SINGLE-SOURCE]heading flag; renderers surface them as badges.update_of: RETIRED in v4.0. The gate FAILs any non-null value: developments and corrections areupdates[]records on the existing entry, never a second entry. (A long-running campaign's routine drip is consolidated, typically about one update record a week; every material development ships when it lands.)references[]: entry ids this entry builds on. It is also the explicit dedup statement: a genuinely distinct finding whosecves[]intersect an existing entry's MUST list that entry here, or the gate FAILs the new entry as a duplicate (§ Dedup across runs).actions[]: only actions derived from this entry's own content, held to the do-now bar (prompts/cti-run.mdPhase 4 §actions[], v3.19): concrete, self-contained, start-now tasks, never generic advice, never a restatement of the body's detection/hardening guidance. Empty is the normal case for many entries. The rendered brief's § Action Items is the union over the window, so every marginal action dilutes it for the reader.migrated_from: non-null marks a v2-brief import. Migrated entries may carry placeholderevidence[], emptyentities/actions/techniques, and news-register bodies; machine consumers (triage agents, exporters) should treatmigrated_from != nullas a lower-fidelity tier and prefer native entries when both cover a topic. An audit may lift a migrated entry through animprovementrecord like any other entry; the provenance flag itself never changes.org_triage/classification: every entry carries exactly one classification scheme, selected by kind. Triage kinds (classification.triage_kindsinconfig/org-profile.yaml, defaultvulnerability) carryorg_triage: {category, rationale}andclassification: nullonly while the profile configures a triage scheme; with none configured (the shipped profile) they carry the Admiralty block like everything else. Every other kind carries the NATO Admiraltyclassification: {reliability, credibility}(letter A–F for the sourcing, number 1–6 for the item, assessed independently) andorg_triage: null. Both schemes and the kind split are config-driven; the gate FAILs an out-of-vocabulary code and the verifier flags mis-placement (F16 / F17). There is no TLP gate anywhere, everything underintel/is processable.priority+immediate_action: see next section.
Priority, the notification surface
| value | meaning | rendering |
|---|---|---|
critical |
"stop reading and act now", the v2 Immediate-Action bar, unchanged and still intentionally extremely high | callout above TL;DR; immediate_action block REQUIRED; notification hooks fire |
high |
leads the window, a reader who reads only the TL;DR must see it | TL;DR bullet (headline + summary) |
notable |
standard item | section body |
routine |
marginal but worth the record (e.g. hygiene CVE kept for awareness) | section body, after notable |
priority: critical ⇔ immediate_action present (both directions,
enforced by tools/check_run.py). The bar for critical is ALL of: newly
disclosed or newly weaponised; actively exploited right now or mass
exploitation imminent / campaign underway with confirmed impact; defender
action time-critical to the hour or day. Criticals are rare *by
construction*, that bar is extreme, not because a count caps them. Two
critical entries in a rolling 24 h is legitimate only when each
independently clears every element of the bar.
Kinds, what renders where
kind |
brief section (content_model.KIND_DAILY_SECTION) |
writable by v4+ runs |
|---|---|---|
threat |
§ Active Threats, Trending Actors, Notable Incidents & Disclosures | yes |
incident |
§ Active Threats (incident / disclosure flavour) | yes |
vulnerability |
§ Trending Vulnerabilities | yes |
research |
§ Research, Reports & Policy | yes |
annual-report |
§ Research, Reports & Policy (one-time treatment per PD-9) | yes |
policy |
§ Research, Reports & Policy, a regulatory action or deadline with a transferable obligation for the constituency (PD-11 c) | yes |
Every kind is writable; content_model.ACTIVE_KINDS is KINDS. The
weekly routine's own kinds (synthesis, outlook) went with its entries
on 2026-08-29, along with the horizon axis and weekly_section;
check_run.py FAILs an entry that re-grows any of them. Orthogonal flags
relocate an entry at render time: deep_dive: true ⇒ § Deep Dive (and not
its kind section); an entry with a changelog record dated inside the
rendered day/window additionally appears in § Updates to Prior Coverage
(rendered from that record, see § Rendering).
Entry lifecycle, one living entry per finding
A finding is published once and then maintained in place. The same
file carries the original analysis, every later development, every
correction and every improvement, each as a dated, attributed changelog
record. This replaces the v3 "immutable entry + update_of second entry"
model (retired 2026-08-27, operator decision): a reader, human or triage
agent; opens one URL and sees the current state of the finding and how
it got there, and an update surfaces on the live brief exactly like a new
finding would.
The changelog record
updates:
- at: "2026-07-05T04:40:12Z" # UTC; strictly later than discovered_at and than the previous record
run_id: 2026-07-05T0410Z-intel # the fire that made the change — resolves to a run record
type: update # update | correction | improvement (vocabulary below)
summary: > # 1–3 self-contained sentences: what changed and why. This is the
… # text the live timeline row, the day page's § Updates and the feed
# item show — a reader who sees ONLY this knows what moved
fields: [cves, priority] # optional: the frontmatter fields this record changed in place;
# "body" when the main analysis itself was edited (corrections)
internal: false # optional (v4.2): true = a pipeline-internal fix (metadata /
# frontmatter hygiene, structured-field corrections with no
# reader-facing delta). Internal records have NO body section,
# are never rendered anywhere on the site, and never move
# updated_at — the changelog documents them for the operator only
merged_from: null # optional, migration provenance only: the v3 update_of entry id
# that was folded into this record (the build redirects its old URL)
updated_at: "2026-07-05T04:40:12Z" # == at of the last `type: update` non-internal record; null when none
Every non-internal record pairs 1:1, in order, by at with a body
section headed exactly ## <Type> — <at> (Update, Correction or
Improvement, an em dash, the record's at verbatim,
content_model.update_section_heading); an internal: true record has no
section. The section carries the delta only, inline-cited like any other
prose, never a recap of the entry, and never pipeline internals: field
names, run mechanics and record-keeping narration ("this entry's cves[]
record carried …") do not belong in reader-facing text; a change with
nothing to tell the reader is an internal record. The main analysis is
everything above the first such heading and must remain a complete,
readable entry on its own.
Only type: update moves updated_at (v4.2, operator directive
2026-08-28). A material new development re-floats the entry to the top of
the live brief; a correction or improvement does not, it changes the
entry in place, its section (when reader-facing) renders on the entry page
and in the day page's § Updates by its record at, but the finding's
position in the live timeline stays where the story last moved.
updated_at therefore mirrors the last non-internal type: update record
and is null when the entry has none.
type |
when | what changes |
|---|---|---|
update |
a material new development on the finding, new actor, victim, CVE in the chain, patch shipped, exploitation-status change (incl. a KEV listing of a not-yet-exploited CVE), confirmed law-enforcement action | the section states the development; the frontmatter moves to the new current state: cves[].status/fixed, affected_products, entities, techniques, tags, actions[] (replace, never accumulate; the list is the current do-now set), and priority/immediate_action when the bar changes in either direction; headline/summary are revised only when a reader who sees only the summary must now know something different (e.g. now exploited) |
correction |
the entry stated something wrong; a claim its source does not support, an inverted mechanism, a wrong version/date/score/id, a mis-attributed quote | the wrong statement is fixed where it stands (frontmatter and/or body, fields names them, "body" included) so the entry never asserts something known to be false, AND the section records what was wrong, what is right, and the ground-truth source. The reader can see both the corrected text and the correction note; git carries the exact diff |
improvement |
precision or depth added without reversing a claim, a second independent source, a technique mapped that the body already described, a **Triage:** line the mechanism supports, a fixed version stated to vendor precision, a Background paragraph |
the section states what was added and on what basis; touched frontmatter listed in fields |
Rules, all enforced by tools/check_run.py (entry-updates, silent-edit) unless noted:
- Provenance never moves. The entry id/path,
discovered_at,run_idandmigrated_fromare never edited. Folder date ==discovered_atdate forever, whateverupdated_atsays. - No silent edits. Every change to a published entry's file, any field,
any prose; ships with a changelog record whose
run_idis the editing fire. At the gate, an entry modified in the working tree relative toHEADthat carries neitherrun_id == <this run>(new) nor anupdates[]record withrun_id == <this run>FAILs (silent-edit). - One record per fire per entry. A fire that changes an entry in
several ways writes one record covering all of them; two fires write two
records. Records are append-only and strictly increasing in
at; a record is never edited or removed by a later fire; the changelog is the audit trail. The entry's content carries no such immutability (operator directive 2026-08-28): a later fire may revise the frontmatter, the main analysis, and the text of earlier## <Type> — <at>sections alike (a wrong earlier update is fixed where it stands) provided the fire's own record cleanly declares the change (a furthercorrectionrecord whosefieldsname what moved,bodyincluded). - Sources travel with the change. A section's inline citations are
sources[]records like any other claim's; new sources are appended tosources[](the first record stays the original primary), new verbatim quotes toevidence[]. The verifier reads the whole entry and checks the new section and every changed field against them. - The main analysis stays current. An
updatethat supersedes a statement in the main analysis (a "no patch" claim after the patch shipped, "PoC only" after exploitation) edits that statement too, a minimal,fields: [body]-declared edit, so the entry never contradicts itself; the section explains the change. Developments are never absorbed into the analysis silently: the section is where the reader learns what happened when. - Who may update. Any intel run (its own dedup decision, PD-8) and the quality audit (soundness corrections, completeness improvements). The audit's former "immutability-exception ledger" is retired, the changelog is the ledger, and it lives with the entry.
- Cadence discipline. A long-running campaign's routine drip is
consolidated, typically about one
updaterecord a week, and every material development ships when it lands; bookkeeping that changes nothing a reader would act on (acisa-kevflag on a CVE the entry already calls exploited) is not a record.
How the pipeline reacts to an update
- Live brief (the landing page
/,data/briefbook.json, brief.js): the entry is in the window iff its activity moment is; it renders in the run group of the fire that made the latest record, flaggedUPDwith the record's type andsummaryshown under the headline; the feed head's "updated" count includes it. An entry appears once, at its latest activity. - Day pages (
/daily/<date>/): § Updates to Prior Coverage lists every entry with a record dated that UTC day, rendered from the record (type, time, summary, the section body, link to the entry); the entry's kind section still shows it only on itsdiscovered_atday. - Entry permalink: "First published <discovered_at> · Updated
<updated_at>" in the meta line; each
## <Type> — <at>section renders as a timestamped, type-badged block; a Revision history panel lists the records (type · time · run link · summary · changed fields). - Feeds:
feed-items.xmland the sector feeds emit one item per entry (pubDate=discovered_at) and one item per changelog record (guid=<entry url>#update-<at>,pubDate=at, title prefixed by the type, description = the record'ssummary, content = the section). data/alerts.json: an entry enters the 7-day window by activity moment and carriesupdated_at+ a compactupdates[](at,type,summary), so a hook can alert on a critical/high entry's update.- Redirects: for every record with
merged_from, the build emits a meta-refresh stub at the folded entry's old permalink pointing at the living entry (noindex, excluded from the sitemap). - Run records:
entries_updatedcounts the entries a fire appended a record to andupdated_entry_ids[]names them (§ Run records); the run's detail page and the ops dashboard list them beside its new entries.
Relevance discipline; volume follows relevance, not cadence or a count
Entry volume is not fixed; there is no per-run, per-day, or
rolling-24-hour target or ceiling. The rolling 24-hour window across all
runs carries exactly the entries that clear the intel run's strict
relevance/actionability gate (prompts/cti-run.md PD-11), however few or
many the window's genuine signal turns out to be. A quiet day is a handful
of entries or none; a day with several unrelated actively-exploited edge
RCEs plus a home-region incident is legitimately larger. The reader is
protected from overflooding by the gate, not by a quota: every entry
must earn its place, and a marginal item is dropped no matter how much room
a numeric budget would have allowed.
The gate is applied for two properties whose weight differs by severity (v4.2, operator directive 2026-08-28):
- Sound: everything published is relevant, accurate, and actionable; very low false positives; no marginal, off-scope, or unverified item. Applies with full force to every entry.
- Complete: everything genuinely relevant to the reader's job is published; very low false negatives; a reader relying on ctipilot.ch alone has no blind spot on anything that matters to their work. Applies with full force to the critical and high-severity signal (an exploited exposure, an active campaign or confirmed incident touching the constituency); below that bar, completeness yields to quality and a marginal awareness item is better dropped or held to two sentences.
A missed critical or high item is the worst failure the brief can have, and a silent one, since the reader never sees what they were not told, so completeness is verified deliberately (the intel run's Phase 2 completeness sweep; the verifier's coverage + missed-angle checks), not assumed.
- Each
vulnerabilityentry must demand action beyond the regular patch cycle, actively exploited, imminent mass exploitation, pre-auth RCE on an exposed edge with public PoC, or another out-of-band response. A CVE the normal patch cadence already handles, with no exploitation or exposure-driven urgency, is out of scope even at high CVSS. - Deep-dive treatment is reserved for an item that earns the long form
(see the intel prompt's Phase 3 criteria); it is rare by construction, not
by quota. Category rotation is derived from the last 30 days of
deep_dive: trueentries. A day may carry none or several, each on its own merit. priority: criticalis governed by its own extreme bar (§ Priority), not by a count; criticals stay rare because the bar is high.- Every run reads the window's already-published entries first (including earlier runs the same day) and publishes only the delta, so more runs mean lower latency, never more content. An empty run publishes only its run record.
tools/check_run.pyreports the rolling-24-hour composition (operational count, deep dives today, criticals) for the operator's awareness; it does not flag a count as an exceedance.
Entity registry, entities/registry.yaml
The global controlled list of named things the pipeline tracks, so every entry links the same real-world entity to the same key and duplicates cannot creep in. Research and verification agents read it; the main agent extends it (same commit as the entries that need the new key).
schema: 1
entities:
- key: actor:shinyhunters
type: actor # actor | campaign | malware | tool | incident
# | report | trend | policy | product
name: "ShinyHunters"
aliases: ["UNC6240"] # every public alias; dedup checks match against these too
# ambiguous_labels: [] # optional: own name/aliases that never
# phrase-match prose (see below)
nexus: null # taxonomy nexus value when publicly attributed, else null
summary: >
One-to-three sentence definition: who/what this is, first public
reporting, why the pipeline tracks it.
first_seen: "2026-05-12" # first pipeline coverage (entry date)
relations: # optional: typed, directed, evidence-bound
# graph edges (§ Relationships below)
- to: "tool:shinysp1d3r-ransomware"
type: uses # controlled vocabulary — direction matters
source: "2026-06-14/some-entry-slug" # entry that establishes the edge
note: "one-clause basis (optional)"
# merged_into: <key> # optional: tombstone — this record was merged
# into the named canonical entity (see below)
Entity types: actor | campaign | malware | tool | incident | report |
trend | policy | product (trend tracks named vulnerability/technique waves,
policy tracks named regulatory items, both inherited from v2 coverage
tracking; product records are derived from affected_products[], § Products).
Rules: key is <type>:<kebab-slug>, globally unique, never renamed once
published (entries reference it). Aliases must not collide with another
entity's key, name, or aliases (check_run.py FAILs). CVEs are NOT
registry entities; state/cves_seen.json and per-entry cves[] carry the
CVE model. Regions, sectors and theme tags stay in site/taxonomy.yaml.
Definitions follow sourcing rules: the summary states only what cited
public reporting supports (attribution stays claim-attributed).
Naming convention (uniform across the registry): name is the concise
canonical entity name only, the name of the actor/campaign/tool itself,
never the reporting vendor, never a headline sentence, never a list of
alternates. Every other public name goes in aliases (which feeds both
dedup matching and the site's phrase-based entry↔entity attachment).
summary is the 1–3-sentence English definition carrying the
who/what/so-what plus the attributing source and date.
Ambiguous labels: explicit-key attachment only. The site attaches an
entry to an entity when the entry keys it in entities[] or when the
entity's name or an alias appears in the entry's title, headline or body
(word-boundary match, single-token labels case-sensitive). That second path
breaks when a label is also ordinary vocabulary or another thing's name:
the actor that calls itself "fingerprint" collected 24 unrelated entries
about TLS and device fingerprinting before this rule existed, UNC6671's
alias "Falcon" matched every CrowdStrike Falcon mention, "Troy" matched
Troy Hunt and "Everest" matched Everest Forms. Casing cannot separate
them ("Payload delivery:" opens a sentence). Such labels go in the record's
ambiguous_labels list: they stay the entity's display and dedup labels,
but they never phrase-match, so only an explicit entities[] key attaches
an entry through them (content_model.prose_match_labels, and the same rule
governs check_run.py's relation-evidence check). Every value must be the
record's own name or one of its aliases (compared case-insensitively, and
validate_registry FAILs anything else). Set it when registering a name
that is an English word, a person's first name, or another vendor's product
name, and make sure every entry about the entity keys it explicitly.
Merging duplicates, merged_into tombstones. Because keys are
permanent and hundreds of published entries reference them by key, a
duplicate entity is never deleted while any entry references it. Instead the losing record becomes a
tombstone: it keeps its key and gains merged_into: <canonical-key>.
Semantics enforced by content_model.validate_registry (surfaced as FAILs
by check_run.py): the target must exist and must not itself be a
tombstone (no chains); tombstones are exempt from the name/alias collision
check (their labels legitimately move to the canonical record). Consumers
resolve through tombstones via content_model.resolve_entity_key: the site
attaches a tombstone's entries to the canonical entity's page (the
tombstone keeps a stub permalink pointing forward), and cross-run dedup
treats old and canonical keys as the same entity. New entries MUST
reference the canonical key, never a tombstone. When tombstoning, move the
losing record's relations[] onto the canonical record (dropping edges the
canonical record already carries, and retargeting registry-wide edges that
pointed at the loser); a tombstone carries no relations, and no relation
targets one. An entity referenced by zero entries (orphan) that turns out
to be a duplicate may simply be deleted, fold its names into the
canonical record's aliases and migrate its edges first.
Products, affected software as entities
affected_products[] names the software an entry concerns, at the precision
a responder needs: "Microsoft SharePoint Server 2019", "Microsoft SharePoint
Server Subscription Edition". That is exactly the wrong granularity for a
pivot (nobody wants one page per release) so every string ALSO resolves to
a product entity, and a product then sits beside actors, malware and
campaigns: its own permalink under /entities/product:<slug>/, a coverage
timeline of every vulnerability, incident and campaign that touched it, an
aggregated ATT&CK profile, and a node in /graph/.
The entry is never rewritten. Resolution happens at render time, in two
stages (content_model.product_key):
- the registry's own
product:records: theirnameandaliasesare the curated merge surface. Six spellings of SharePoint fold ontoproduct:microsoft-sharepointbecause that record lists them as aliases; - a mechanical fallback: drop a trailing release year, dotted version or edition word, then slugify. Bare integers are never stripped, so "Microsoft 365" and "Dynamics 365" survive intact while "…Server 2019" and "ColdFusion 2025" fold.
tools/sync_products.py keeps the registry's product block in step with the
store: it reads every affected_products[] string, upserts one record per
product, and preserves every curated field it finds (name, summary,
aliases, relations, merged_into). --check reports drift and exits 1;
check_run.py --all warns when a spelling resolves to no record.
Two consequences worth stating plainly:
- Merging two products is an edit, not a migration. Add the loser's
spelling to the winner's
aliases(or tombstone it withmerged_into) and re-run the tool. Every entry that named the old spelling follows, because none of them ever stored the key. - Products never phrase-match prose. A product name is ordinary technical
vocabulary; an entry that says "a PHP deserialization bug" is not coverage
of PHP, so unlike every curated entity type, a product attaches only
where the entry itself declared it in
affected_products[]. That keeps the actor/campaign graph from drowning in generic software nodes.
A product record carries no summary unless an operator writes one: it is a
derived index node, not an analytical claim. Vendor-only strings
("Microsoft", "Linux") never become entities, a node attached to a third of
the store is not a pivot. Products are absent from the STIX export: the
faithful STIX shape is the software SCO, which carries none of the SDO
properties the export writes.
Relationships, the threat graph
Entity relationships are typed, directed, evidence-bound edges carried
in each registry record's optional relations[] list. They replaced the
untyped related: [] key list (removed without backward compatibility);
validate_registry FAILs a record that still carries related. The graph
has exactly two edge classes, and every edge's derivation is explicit:
- Curated edges (
relations[]in the registry), a connection a cited source states: "this actor operates this campaign", "this campaign deploys this malware". Each edge names its relationship type from the controlled vocabulary below and cites the entry whose sourced reporting establishes it. - Derived edges (computed by
site/build.py, never stored), a connection the entry store implies: two entities referenced by the same entry (co-occurrence, weight = shared-entry count), an entity and a CVE carried by the same entry, an entity and an ATT&CK technique via the derived TTP profiles (§ The ATT&CK layer). Derived edges are recomputed on every build and always carry their supporting entry ids, they can never drift from the store. Evidence-quality gate (build.derived_edge_qualified): only focused reporting creates a derived edge;annual-reporttreatments are excluded, because they mention many unrelated entities by construction: two names sharing a quarterly ransomware ranking is summarization, not a connection. Curated edges are unaffected (each carries its own establishing entry).
Curated edges assert what happened; derived edges surface *what the store connects*. Renderers keep the two visually distinct (curated edges carry their type label; derived edges are labelled by their derivation), and an analyst reading any edge can always answer "why does this edge exist?"; either "entry X's cited source states it" or "these N entries reference both".
Curated edge shape
relations:
- to: "actor:shinyhunters" # target registry key — MUST exist, MUST be
# canonical (never a tombstone)
type: attributed-to # controlled vocabulary below
source: "2026-06-14/<slug>" # entry id whose cited reporting establishes
# the connection — MUST resolve; the entry's
# date doubles as the edge's first-seen date
note: "GTIG attributes the wave to ShinyHunters" # optional one-clause basis
Relationship vocabulary (controlled, content_model.RELATION_TYPES)
Directed types read subject → object: the edge lives on the subject's
record and to names the object. Renderers show every edge from both ends
(the object's page shows the inverse reading). Symmetric types are stored
once, on either endpoint, declaring the mirror edge too is a FAIL
(duplicate), and renderers/exports surface it on both endpoints anyway.
type |
subject types → object types | reading (inverse reading) |
|---|---|---|
attributed-to |
campaign, incident, malware, tool → actor | subject is attributed to actor (actor's attributed activity) |
uses |
actor, campaign, incident → malware, tool | subject deploys/operates the malware or tool (used by) |
exploits |
actor, campaign, incident → trend, product | subject exploits the named vulnerability/technique wave, or the product itself (exploited by). CVE-level exploitation is a derived edge; the entry that carries both the entity and the cves[] record is the evidence; CVEs are not registry entities. |
affects |
campaign, incident, malware, tool, trend → product | subject reaches the named software (affected by). A product is the thing attacked, never the attacker: nothing points out of one except related-to / documented-in. |
part-of |
incident, campaign → campaign, trend | subject belongs to the larger campaign/wave (includes) |
variant-of |
malware, tool → malware, tool | subject is a variant/fork/derivative of the object (has variant) |
successor-of |
actor→actor, campaign→campaign, malware→malware, tool→tool, policy→policy | subject continues/rebrands/replaces the object (succeeded by) |
collaborates-with |
actor ↔ actor (symmetric) | the two actors cooperate (shared operations, hand-offs) |
overlaps-with |
actor, campaign, malware, tool ↔ same set (symmetric) | cited reporting states technical/infrastructure/TTP overlap short of attribution or identity |
documented-in |
any non-report type → report | the report profiles the subject (documents) |
related-to |
any ↔ any (symmetric) | fallback, a source-stated connection none of the typed relations fits; prefer a typed relation whenever one applies |
Semantics guardrails: attributed-to is for responsibility claims (keep
the claim attributed in the note/entry, per the sourcing rules);
overlaps-with is the honest middle ground when researchers report shared
infrastructure or tooling without asserting identity, never upgrade an
overlap claim to attributed-to or successor-of beyond what the cited
source states. A suspected same entity is not a relation at all, that is
an alias or a merged_into tombstone.
Hard rules (enforced by content_model.validate_registry, surfaced as FAILs by check_run.py)
typemust be in the vocabulary; subject/object entity types must satisfy the type's endpoint constraints.tomust exist, must be canonical (not a tombstone), and must not be the record itself. Tombstones must not carryrelations[], move edges to the canonical record when merging.sourceis REQUIRED and must be a valid entry id (YYYY-MM-DD/<slug>) that resolves to an existing entry; this is what makes every curated edge evidence-bound and dates it.check_run.pyadditionally WARNs when the source entry references neither endpoint in itsentities[](the edge is still legal; the establishing entry may predate one endpoint's registration, but the mismatch is worth an operator's glance).- No duplicate edges: the same
(subject, type, object), for symmetric types the same unordered pair, appears once in the whole registry. New corroboration of an existing edge does not add a second edge; material evolution of the relationship (e.g. overlap upgraded to attribution by new reporting) replaces the edge'stype/source/notein place, relations are registry state, not immutable entries. - Relations are otherwise append-only in spirit: edges are added when a cited source establishes a connection, in the same commit as the entry that carries the evidence.
The graph rendering, /graph/ + data/graph.json
The full graph ships as data/graph.json (all canonical entities, covered
CVEs, mapped ATT&CK techniques, curated + derived edges) and renders at
/graph/ as an interactive, self-contained (strict-CSP, no external
libraries) canvas exploration surface. Exploration is seeded: the
analyst names one or more starting nodes (search, an entity-page deep link
?focus=<id>, or the most-connected directory), and the view renders
exactly the subgraph the analyst has grown from those seeds, the direct
neighbourhood by default (?hops= widens to 2 hops or the full connected
component), extended node by node via expand, and nothing else: nodes
outside the grown view are not drawn at all, not even dimmed; with no
seed, nothing is drawn. Within the view:
type-filtering (entities / CVEs / techniques as a toggleable layer),
curated/derived edge toggles (both also bound reachability), hover
neighborhoods, a node detail panel (summary, typed relations, supporting
entries, including connections outside the current view), re-seeding from
any node, and shortest-path tracing between two nodes; "how is this actor
connected to this CVE?" answered visually, every hop backed by an edge
whose provenance is one click away. Entity pages render the same edges in
prose form: typed curated relations grouped by relationship reading, each
with its source entry link, followed by the derived co-occurrence list.
The ATT&CK layer, pinned dataset + derived TTP mappings
CVEs, actors, campaigns and every other entity get their MITRE ATT&CK technique profile by derivation, never by assertion: an entity maps a technique exactly when a published entry ties them together. The layer has three parts:
- The pinned dataset,
attack/enterprise-attack.json(contract: attack/README.md; writer:tools/attack_data.py). A compact, committed extraction of one specific ATT&CK Enterprise release: technique id → name, tactics, first-paragraph definition, sub-technique parentage, platforms, and lifecycle flags. Pinning matters because releases drift, v19 renamed Defense Evasion into Stealth + Defense Impairment (new TA0112) and every release revokes ids. Revoked/deprecated techniques are kept, flagged, withrevoked_byforwarding, the ATT&CK analogue of the registry'smerged_intotombstones, and for the same reason: the store is not rewritten when the pin moves, so an id cited before MITRE revoked it must keep resolving (content_model.resolve_technique_id). Updating the pin is an explicit, diff-reviewed act:tools/attack_data.py --check(drift detection; quality-audit duty) /--update(rewrite + change summary for the commit body) /--selftest(offline invariants; also enforced bycheck_run.py). - Per-entry effective techniques,
content_model.entry_technique_ids. The union of the entry'stechniques[]frontmatter (canonical, v3.17+) and dataset-known T-ids in its body prose (the only mapping surface of the pre-v3.17 tail, which is not bulk-rewritten), revoked ids resolved forward. Exposed per entry indata/briefbook.jsonanddata/alerts.jsonastechniques[]. - Derived aggregations (
site/build.py). Per entity AND per CVE:{technique id: [supporting entry ids]}, evidence-bound, rendered as the entity page's ATT&CK section (collapsed by default, grouped by tactic in official matrix order, definitions from the pin, entry links) and exported as a per-entity ATT&CK Navigator layer (entities/<key>/attack-layer.json, layer format 4.5, score = supporting-entry count). The/attack/page renders the full matrix heat-shaded by store-wide coverage, carries the per-technique definitions-and-evidence directory, and offers the client-side multi-entity overlap view (union / overlap≥2 / common-to-all) overdata/attack.json, Navigator-layer semantics without leaving the site, plus layer export of any comparison. Only entries ABOUT the entity contribute (anentities[]key, anaffected_products[]string for a product, acves[]id for a CVE). An entry that names the entity in passing ("sells footholds to Qilin, Akira and Rhysida") still appears on the entity page's story timeline, tagged mention, but it lends the entity none of its techniques, and none of its products, CVEs, sectors or action items either: it documents someone else's behavior.
Run records, runs/YYYY-MM-DD/<run-id>.md
One file per fire, written in the run's final phase. Frontmatter is the
complete machine-readable telemetry record (the v2 run_log.json entry,
relocated); the body is the human-readable verification & coverage
notes, the v2 brief § 7, relocated to a dedicated, per-run home.
Run records are immutable once their fire completes (unlike entries, which
are living records; a run record is telemetry about one fire and has no
"current state" to maintain), with exactly two same-fire in-place updates
permitted: the same-minute retry (idempotent run_id) and the Phase 7
publish-status amendment, after the publish poll, the fire updates
publish_status/publish_checked_at/publish_note in place, commits
run: <run-id> publish-status, and re-pushes the feature branch
(fire-and-forget; auto-merge promotes it). No other field is ever edited
after commit, and no later fire edits an earlier fire's record.
An optional stood_down: <reason> field (non-empty string) marks a fire that
legitimately aborted before Phase 1 spawned any research/verification
workers, e.g. the quality audit's duplicate-audit guard (gap since the last
audit < 72 h). Such a fire still writes a run record (run-record-per-fire is
never waived) but carries an empty sub_agents block, since no sub-agents
ran; the mechanical gate exempts the sub_agents requirement only when
stood_down is set. The mandatory verification iteration still runs (scoped
to the run record). Normal fires omit stood_down.
---
schema: 1
run_id: 2026-07-03T0412Z-intel
kind: intel # intel | audit (matches the run-id suffix; `weekly` = legacy records)
date: "2026-07-03"
started: "2026-07-03T04:12:03Z"
completed: "2026-07-03T04:31:40Z"
duration_seconds: 1177
model: "…" # main-agent friendly name (self-ID: harness prompt line; env vars as marked fallback)
model_id: "…"
prompt_version: "v4.17"
window_hours: 24 # gap-derived recency window this run covered (24 h floor)
gap_hours: 7 # hours since the previous run record
entries_published: 3 # NEW entry files this run (run_id == this run)
entries_updated: 1 # existing entries this run appended an updates[] record to
updated_entry_ids: # v4.0: their ids — len == entries_updated; [] when none
- 2026-07-01/some-earlier-entry
deep_dive: null # entry id of a deep-dive entry published this run, or null
sub_agents: # S1–S4 (+S5); audit fires: truth-pass and re-sweep workers
S1:
model: "…"
model_id: "…"
started_at: "…"
ended_at: "…"
duration_seconds: 279
sources_attempted: [cisa-kev, bsi-de]
sources_used: [cisa-kev]
items_returned: 2
returned: true
telemetry: {webfetch_calls: 8, websearch_calls: 0, bridge_fetches: 14}
fetch_failures: [] # rich v2 shape: {id, url_tried, fetch_method, status_code,
# error_class, error_message, attempted_methods, mitigation_applied, covered_anyway}
bridge_uses: [] # {id, method, outcome}
sources_changed: [] # {id, change, from, to, reason}
entities_added: [] # registry keys added this run
entries_dropped_by_verification: 0
publish_status: pending # pending | ok | main-only — machine-auditable publish outcome.
# Written `pending` at the Phase 6 commit; the SAME fire amends
# it in place after Phase 7's poll (ok = run record on main AND
# site rebuilt, or site polling disabled; main-only = record on
# main but the site rebuild never confirmed) and pushes the
# amendment. A record still `pending` on main means the fire died
# before Phase 7 or the amendment push failed — an operator signal
# either way. Absent on records that predate v3.14.
publish_checked_at: null # UTC timestamp of the Phase 7 poll that set publish_status
publish_note: null # free-text reason detail (e.g. "site polling disabled",
# "auto-merge pending at deadline")
verification_iterations: 2
verification_residual_count: 0 # never 0 when the final iteration was NEEDS_FIXES
verification:
confirmation_waived: null # optional (v3.23+): non-null string ONLY when the run published
# on a CLEAN that no second pass confirmed (single CLEAN at the
# iteration cap, confirmation spawn blocked) — the reason, verbatim. Normal confirmed-CLEAN
# publishes omit it. `check_run.py` FAILs an unconfirmed final
# CLEAN on v3.23+ records unless this (or the cap) explains it.
iterations: # v3.23+: a CLEAN publish requires the final TWO iterations both
# CLEAN — so a CLEAN publish has ≥2 iterations. v4.1+: both run
# the single `cti-verification` definition (generic `sonnet` pin); the
# two-different-models requirement of v3.23–v4.0 is retired
- n: 1
model: "…"
model_id: "…"
subagent_type: cti-verification # the verifier definition spawned (single since v4.1)
started_at: "…"
ended_at: "…"
duration_seconds: 240
verdict: CLEAN # CLEAN | NEEDS_FIXES
truth: 0 # F1–F4 + F13–F15
editorial: 0 # F5–F10 + F12 + F16–F18
advisory: 0 # F11
claims_in_scope: 214 # v4.17: the claim-ledger claims this pass had to answer for
claims_checked: 214 # v4.17: rows in verification.iter<N>.claims.yaml (== in-scope
# for a complete pass; an incomplete pass's CLEAN never counts)
findings: [] # rich per-finding records, v2 shape
- n: 2 # the confirmation pass — an independent cold read, also CLEAN
model: "…"
model_id: "…"
subagent_type: cti-verification
started_at: "…"
ended_at: "…"
duration_seconds: 210
verdict: CLEAN
truth: 0
editorial: 0
advisory: 0
findings: []
---
## Verification & coverage notes
The v2 § 7 content, per run: borderline drops with reasons, single-source
items and their carve-outs, reduced-confidence inclusions, contradictions,
out-of-window drops, stalled sub-agents, and the parseable lines —
`Coverage gaps: …`, `Watchlist: …`, `Closed-source intake: …`,
`Essential-coverage: …`.
The rendered window brief concatenates the run-record bodies of every run
in the window as its § Verification Notes, newest first. The Ops dashboard
is built entirely from runs/** frontmatter.
Dedup across runs; how overlap is prevented
- Preflight scan. Every run builds
work/<run-id>/prior_coverage.jsonby scanningentries/for the last 14 days plus everything already published today (multiple-runs-a-day is just more records in the same scan). Records carry: entry id, title, headline,summary, kind, CVE ids, entity keys, primary URL,discovered_at. Thesummarymakes the file a load of every in-window brief, not just a key list. - Compose-time dedup (in-context, 14 days). The main agent
Reads the fullprior_coverage.json, every in-window brief loaded into context (each record also carriesupdated_at, its record count and the last record's summary), and never writes a new entry for a candidate whose CVE ids or entity keys match an in-window entry from any run in those 14 days: the candidate is either anupdates[]record on that entry (material delta) or nothing. - Metadata check (store-wide, older than 14 days). Coverage older than
the 14-day in-context window is caught by the store-wide CVE index
(
state/cves_seen.json, surfaced ascves.idsin the state summary), not an in-context read; an old CVE re-surfacing is still recognised and handled the same way (a record on the existing entry). - Fetch-time dedup. Research sub-agents read
prior_coverage.jsonbefore fetching and skip already-covered items unless they hold a material delta (returned asnovelty: update-of:<entry-id>). - Mechanical gate.
tools/check_run.pyFAILs a new entry whose CVE set intersects any existing entry (store-wide, any age) unless the new entry lists that entry inreferences[], the explicit, reviewable statement that this is a distinct finding building on the older one, not a duplicate, and WARNs on entity-key overlap inside 14 days. Entries updated by the run are exempt (the record IS the dedup decision).
Rendering; the brief is a query
/(the landing page): the live rolling brief, rendered as a run-grouped, reverse-chronological timeline keyed on each entry's activity moment (max(discovered_at, updated_at)): a new finding appears under the run that published it, an updated finding under the run that updated it (flagUPD, the record's type + summary shown), each entry once at its latest activity. Reader picks last N hours (6 / 12 / 24 / 48 / 72) via the window selector or loads older findings; default 24 h. Every run in the window appears, including quiet (0-finding) ones. The default window ships server-rendered (full content, no-JS readable); JS re-renders the timeline client-side fromdata/briefbook.json(last ~35 days of entries by activity + run records). Each timeline row carries priority / CVE / exploited / updated badges, a linked headline, provenance, and a clickable source list./daily/YYYY-MM-DD/: one settled page per completed UTC day (the still-rolling day lives only in the landing-page brief), in the classic editorial section order: TL;DR → Active Threats → Trending Vulnerabilities → Research, Reports & Policy → Updates to Prior Coverage (every entry with a changelog record dated that day, rendered from the record) → Deep Dive → Action Items, with a collapsible Verification block./daily/is the newest-first archive; daily RSS keys on these./entries/YYYY-MM-DD/<slug>/: per-entry permalink: current state first, the changelog sections as timestamped blocks, a Revision history panel; "First published … · Updated …" meta. Folded v3 update entries' old URLs redirect here (merged_from).- Feeds:
feed-items.xml(one item per entry,<pubDate>=discovered_at(true discovery latency, not commit time) plus one item per changelog record,<pubDate>= itsat) + the sector slices fromconfig/branding.yamlfeeds.sector_slices(this deployment:feed-public-sector.xml) + the daily digest feed. There is no weekly feed. data/alerts.json: last 7 days (by activity moment) ofcritical/highentries with headline, summary, immediate_action, entities, CVEs, techniques,updated_atand compactupdates[]: the notification-hook surface./attack/+data/attack.json+entities/<key>/attack-layer.json
the ATT&CK coverage matrix, its client-side overlap dataset, and the per-entity Navigator layer exports (§ The ATT&CK layer).
/graph/+data/graph.json: the interactive threat graph over all canonical entities, covered CVEs and (toggleable) ATT&CK techniques: curated typed edges + derived co-occurrence/CVE/technique edges, each with its provenance (§ Relationships)./stix/, STIX 2.1 bundle endpoints (site/stix_model.py), the store compiled to STIX 2.1 on every build, pulled like the RSS feeds:bundle.json(full corpus),recent.json(briefbook-window activity, reference-closed, the poll target for TIP platforms such as OpenCTI),entities.json(core entity graph),sector-<slug>.json(the RSS sector slices),extension-schema.json. Onereportper entry; the registry as intrusion-set/campaign/malware/tool/incident (+ trend → grouping, policy/report → report); one sharedvulnerabilityper CVE;techniques[]as MITRE's canonical attack-patterns;relations[]as SROs. Every id is uuid5 over the permanent store key (entry id, registry key, CVE id) under the branding-URL namespace, so one finding keeps one object forever, the shared-object graph carries the cross-brief connections, and re-ingestion is idempotent. Deliberately no TAXII server: a static host cannot satisfy the TAXII 2.1 media-type / header / filtering MUSTs.- Entity pages, trends, ops, search: all derived from entries +
registry + runs, same URLs as v2. An entity page is laid out for the
analyst who pivots onto it to decide what to do: at-a-glance tiles
(coverage, latest activity, peak priority, targeted sectors and regions,
sources), then § Action items (the
actions[]andimmediate_actionof every entry about the entity, newest first) and § Defender insights (each entry's**Defender takeaway:**, with its**Triage:**and detection guidance one click away), then the typed relationships, the story timeline with passing mentions tagged, the hunting pivots (CVEs exploited first, affected products, tags), the collapsed ATT&CK profile, and the embedded entries. Covered techniques are searchable.
The mechanical gate, tools/check_run.py
Replaces tools/check_brief.py. Read-only, stdlib-only, exit 0 required
before the verifier spawns and before every commit. Validates: frontmatter
parses and every field is schema- and taxonomy-valid; folder-date/
discovered_at/slug consistency; source-URL block-list + liveness (honouring
work/<run-id>/url-liveness.tsv); evidence shape/presence; evidence quotes
literal-searched on their cited page (quote-literal, v4.13, WARN on a quote
that is not a contiguous passage of the page) and every CVE id in a cited
clause found on that clause's cited page (citation-cve, v4.13, WARN), both
over the direct transports only with bodies cached under
work/<run-id>/quote-bodies/ (--page-checks-since DATE runs the pair over
a date range, the audit's pre-pass); priority ⇔
immediate_action consistency; entity refs resolve; registry integrity
(incl. typed-relation vocabulary, endpoint constraints, canonical targets
and source-entry resolution, § Relationships);
the entry lifecycle (updates[] shape, updated_at mirror, strictly
increasing at, 1:1 body-section pairing, record run_id resolution, no
silent edit; an entry modified without a record for the modifying run
FAILs; update_of retired; any non-null value FAILs; entries carry no
deleted legacy field); references[] resolution;
cross-run dedup (store-wide CVE overlap FAILs unless declared in
references[]); run counters vs disk (entries_published,
entries_updated + updated_entry_ids, deep_dive); rolling-24 h
composition (reported, not gated on a count); CVE sync with
cves_seen.json; IOC scan; run-record completeness (incl. verification
counters and prompt-version cross-check against prompts/CHANGELOG.md);
sources/sources.json shape (incl. Admiralty A–F reliability_codes);
closed-source citation traceability to intel/ (no TLP gate); org-triage and
Admiralty-classification vocabulary/placement; the ATT&CK layer (pinned
dataset present + invariant-clean, FAIL; techniques[] ids unknown /
revoked / deprecated in the pin, FAIL on the run's own entries (v3.21),
WARN store-wide; prose-mapped ids missing from the frontmatter, WARN); and
the site smoke tests (site/test_build.py).
v4.17 additions (run scope, gated on the run's prompt version):
changelog-fields (a record's fields must name every frontmatter field
its fire changed, plus body for an analysis edit; an analysis edit of more
than 12 words carried only by internal: true records FAILs, since an
internal record may only re-point a citation or re-word a few words); dedup-extended (two new entries
of one fire sharing a CVE, or a new entry sharing another entry's primary
source URL plus an entity or most of its title, FAIL unless declared in
references[]); run-integrity (iteration counter vs iterations listed,
duration_seconds vs the stamps, a NEEDS_FIXES final pass with no truth or
editorial finding); entry-shape (a new body with no inline citation FAILs;
headline over 120 characters, main analysis over 550 words or 1,100 for a
deep dive, an uncited sources[] URL, an inline citation date differing from
its sources[] record, no-patch beside a fixed version, a critical
immediate_action naming no fixed version, WARN); exploitation-consistency
(WARN on a sentence denying exploitation of a CVE the entry marks exploited);
reader-text-internals extended to em dashes, KEV remediation deadlines and
pipeline vocabulary across every reader-facing field; defanged indicators in
the IOC scan; cves_seen.json duplicate records and duplicate registry keys
FAIL; an exception inside any check is a check-crashed FAIL rather than a
crashed script.