In late 2025, a threat actor group dubbed NULLRIFT deployed a loader that evaded seven separate AV engines across a financial sector target. The secret: the payload rewrote its own function signatures and variable names every 90 seconds using a fine-tuned LLM running on a compromised cloud instance. This is not a thought experiment anymore. AI-generated polymorphic malware is operational, and your signature-based detection is already behind.
How AI Mutation Engines Actually Work
Traditional polymorphic malware used pre-written mutation routines — swap opcodes, insert junk instructions, re-encrypt the payload. The mutation pool was finite. Defenders eventually catalogued every variant.
AI changes the ceiling. A local LLM (think a quantized Llama or Mistral model) can be prompted to semantically rewrite a function while preserving its behavior. The logic stays identical. The code looks brand new. No mutation table to exhaust.
Here is what a simplified mutation request looks like when an attacker feeds a keylogger function to a local model API:
# Attacker's mutation call — run on compromised host 192.0.2.44
# Model: llama3-8b-q4 served via llama.cpp HTTP server
curl -s http://127.0.0.1:8080/completion \
-H "Content-Type: application/json" \
-d '{
"prompt": "Rewrite this Python function with different variable names, restructured loops, and added benign-looking dead code. Preserve exact behavior.\n\ndef capture_keys(hook_id, event_type, kb_data):\n if event_type == 256:\n log_buffer.append(chr(kb_data[0]))\n return ctypes.windll.user32.CallNextHookEx(hook_id, event_type, kb_data)",
"max_tokens": 300,
"temperature": 0.9
}'
The model returns a functionally identical function — maybe now called process_input_event() with a dummy timestamp_check() block inserted for noise. Temperature set to 0.9 means high variability: run this fifty times, get fifty distinct outputs.
A defender watching this host would see llama.cpp consuming GPU cycles and a process making local HTTP calls. That behavioral fingerprint — not the malware’s code — is your detection hook.
What the Mutation Looks Like on Disk and in Memory
Each generated variant gets compiled or interpreted on the fly, then executed. On a Windows target belonging to user jmartin at hostname WKSTN-047, an EDR pulling process telemetry might surface this:
[EDR Telemetry Export — WKSTN-047 \ jmartin — 2026-10-08 03:17:42 UTC]
Process Tree:
python3.exe (PID 4812)
└─ spawned by: svchost_updater.exe (PID 3201) [UNSIGNED]
File Writes (last 60s):
C:\Users\jmartin\AppData\Local\Temp\~df3a1.py [CREATED]
C:\Users\jmartin\AppData\Local\Temp\~df3a1.py [DELETED 4s later]
Network (PID 4812):
127.0.0.1:8080 ESTABLISHED [local LLM API]
192.0.2.17:443 ESTABLISHED [C2 — exfil]
Hash of ~df3a1.py: a9f3c2b1... [NO MATCH in VirusTotal]
Hash of prior variant (03:16:11): 7e820fd4... [NO MATCH]
Two things stand out immediately. First, an unsigned process called svchost_updater.exe is spawning Python — nothing legitimate does that. Second, temporary Python files are being written and deleted on a tight loop, with a new hash every cycle. No signature will catch these because no two files are identical.
What you do next: pivot on the parent process. Pull svchost_updater.exe from disk, check its PE header, run it through CAPE sandbox. The mutation engine itself — the LLM binary and its model weights — will be somewhere on that host. That payload is static and hashable. The AI model does not mutate itself. That is your anchor.
Detection Strategy: Behavior Over Bytes
Chasing hashes here is a trap. The payload hash changes faster than your feed updates. Shift your detection logic to three behavioral indicators:
- Ephemeral script files: Any scripting-language file created and deleted within 10 seconds under a user temp directory warrants an alert. Write a SIGMA rule for it today.
- Localhost model API traffic: Processes making high-frequency HTTP connections to 127.0.0.1 on non-standard ports while also holding an external connection should be flagged automatically.
- Unsigned parent processes: Legitimate Windows services do not spawn interpreters. Enforce allowlisting on what can launch
python.exe,node.exe, orpowershell.exe.
On the network side, the C2 connection is often the weakest link. The LLM handles mutation locally precisely to avoid sending code over the wire — but exfiltration still has to happen. JA4 fingerprinting on that TLS session to 192.0.2.17 may reveal a client signature inconsistent with any known browser or application on that host.
Memory forensics is the other lever. The generated payload executes in-process. A Volatility dump of PID 4812 will contain the plaintext mutated source before Python’s compiler discards it. If you have EDR memory scanning enabled, tune it to flag executable regions in Python’s heap — unusual but detectable.
The attacker’s advantage is code novelty. Your advantage is that behavior leaves marks that no LLM prompt can erase.
What To Do Now
Pull your SIEM’s process-creation logs for the last 30 days and run this one query: find any instance where python.exe, node.exe, or wscript.exe was spawned by a parent process that is not on your approved software list. Sort by rarest parent binary first. If you find unsigned parents spawning interpreters, you have either already been hit or you have a misconfiguration that makes the attack trivially easy. Fix the allowlist or start your IR process — either outcome is worth 20 minutes of your time right now.
