Cisco Talos: malware now embeds natural-language instructions aimed at LLM triage pipelines, from a copied comment to template-sprayed prompts across four families, and a Russia-aligned group uses it too
Talos: malware tells the analysing LLM to skip it; the text is plain and therefore detectable
Analysis
Talos, which tracks AI-integrated malware in its CAIRN toolkit, calls the class "AI-analysis evasion": malware that embeds natural-language instructions to influence automated analysis, aimed at the layer above packers and anti-debug checks, the pipeline that extracts text from a sample and submits it to a language model for triage, classification or reverse-engineering help (Cisco Talos, 2026-10-08). It traces four families and 84 samples collected from January 2025 to July 2026: a PowerShell reverse shell (FRUITSHELL) carrying a two-line comment that says there is no need to analyze the file, a comment later reused verbatim in nine further scripts from at least four actors, among them a script set that adds an AMSI bypass and runtime compilation (Cisco Talos, 2026-10-08). PLOTSAFE generates the sentence from a template and keeps it alive in Go builds with a dummy function that references the strings; HOLLOWCLAD repeats one refusal instruction across seven LLM chat-template formats so one copy may match the scanner's own template, and adds fake honeypot and licence warnings; MANTLEMAZE claims invented government contracts and certifications (Cisco Talos, 2026-10-08).
ESET found the technique in the hands of a named group: the Russia-aligned UAC-0099 put a decoy request for guidance on building a nuclear weapon into the comment of a VBScript, aiming to trip the safety guardrails of an LLM-powered code scanner so it stops before the malicious code that downloads the MATCHBOIL loader (ESET, 2026-09-10). Talos tested the strings against five local models, 135 matched pairs per string: its headline figure is that the best techniques steered the outcome in the attacker's favour in about 35% of runs, a net rate (pairs shifted toward benign minus pairs shifted toward malicious, over all pairs), while the cheapest technique, a direct instruction to ignore the sample, worked almost universally and the more complex ones had little effect or backfired by making models more suspicious; Talos calls the overall impact a mixed bag (Cisco Talos, 2026-10-08). It says core conventional detection is unaffected (Cisco Talos, 2026-10-08).
Triage: Talos' discriminator is that legitimate software has no reason to embed instructions telling an analyzer to refuse analysis, invoke copyright law or claim government contracts (Cisco Talos, 2026-10-08).
Cited evidence
malware that embeds natural-language instructions to influence automated analysis
Legitimate software has no reason to embed instructions telling an analyzer to refuse analysis, invoke copyright law, or claim government contracts.
no single LLM engine should have the sole authority to decide that a piece of code is safe
Sources2
AI-generated · no human review · this permalink is the shareable record for the finding · verify operationally critical claims against the linked primary source.