When the pentester is a fleet of AI agents: inside an autonomous vuln-hunting rig
We index exposed attacker directories for a living, and the tradecraft inside them is changing fast. On 37.60.248.174, a server at Contabo GmbH, we recovered the working tree of an offensive-security operation that is run almost entirely by AI agents, a self-built platform the operator calls "Brainstorm" that orchestrates a fleet of Claude Code workers to find, triage and confirm vulnerabilities at machine speed.
Key findings
- An operation run by agents, not a person. An orchestrator spawns named Claude Code workers ("germinal", "strategist", "Ask Code") that scan, read source and triage findings on their own, writing to a transcript the operator reads later.
- A home-built out-of-band collaborator confirms blind XXE / SSRF / SSTI / RCE / deserialization over DNS, HTTP, SMTP and LDAP, with per-payload tokens, durable and replayed on boot.
- Real offensive kit underneath: a nuclei template library, SAST taint "sink" configs (prototype-pollution and a universal catch-all), and responder for LLMNR/NBT-NS poisoning and relay, a capability that points past bug-bounty toward internal intrusion.
- Interchangeable model backend. The rig references both Anthropic and Kimi (Moonshot); these harnesses swap model providers by config, so provider-domain detection rots. We key on behavior and opsec leftovers instead.
- Secrets in the open: the exposed directory held five C2-style config files and a large set of harvested key-value credentials.
Timeline
Figure 1. Observed activity window and inferred operational phases.
How we found it
huntback promotes an open directory to a full harvest only when its listing scores malicious. This one did immediately: offensive scanning configs, a C2-style control service and harvested secrets in an exposed HTTP root is not a legitimate server. We recovered 365 files and classified them automatically. Our sensors place the host on 2026-09-30.
How the rig is wired
Figure 2. The operator points one orchestrator at a target; the agents do the rest.
A fleet of agents, not a person
The centre of the operation is an orchestration layer that spawns named Claude Code agents, "germinal", "strategist" and others, each given a role in the vulnerability-research loop. An install-codeintel.sh script wires the Serena LSP server in as an MCP tool so every spawned agent gets IDE-grade semantic code navigation (find_symbol, goto_definition, reference search) across 40+ languages, with ast-grep as a structural fallback. A companion assistant, "Ask Code", is described in its own prompt file as "a persistent agent with full read access to the VPS... and authenticated access to the Brainstorm HTTP API", whose job is to "mass-triage Code findings fast and correctly". The agents write to a transcript file rather than a human chat: the operator reads the results, not the conversation.
A home-built out-of-band collaborator
Confirming blind vulnerabilities needs an out-of-band listener, and rather than rent one the operator built their own. The recovered OOB Collaborator is a zero-dependency logger for blind XXE, SSRF (including JWT kid/jku/x5u abuse), SSTI, RCE and deserialization over DNS, HTTP, SMTP and LDAP. Every payload an agent emits is registered with a unique token, so any callback maps straight back to the request that caused it, durable, append-only, replayed on boot. This is the design of a commercial interaction server, rebuilt in-house for an automated pipeline.
Classic offensive tooling underneath
The AI layer sits on top of ordinary offensive kit: a library of nuclei templates, a stack of SAST "sink" configs (prototype-pollution, proto nested-assignment, a universal catch-all taint profile) for pulling dangerous call-sites out of source at scale, and responder, an LLMNR/NBT-NS poisoning and relay tool whose presence (MITRE T1557.001) points past external bug-bounty toward internal-network intrusion. The directory also held five C2-style config files and a large set of harvested secrets.
The model backend is interchangeable
One detail is worth pausing on: the rig is not tied to one AI provider. Its code references Anthropic and Kimi (Moonshot) backends side by side, with a dedicated kimi-usage-parser. That matches public reporting on these offensive harnesses (Hunt.io's SecFlow analysis; Unit 42's tracking of the "knaithe" operator), where one wrapper was pointed at DeepSeek, Qwen, GLM, Kimi or MiniMax just by swapping a model name and an API route. The lesson for defenders is concrete: detection keyed to one provider's API domains ages badly, because the operator changes the backend in a line of config. What survives a backend swap is the victim-side behavior and the opsec leftovers in exposed configs, autonomy flags such as an approval mode set to "yolo" or permissions set to bypass. huntback now flags both the set of backends a host references and those autonomy leftovers, provider-agnostically, so the signal does not rot when the model does.
Attribution and intent
The platform is self-built and the artifacts are English-language; we draw no national attribution from this host. The notable finding is structural, not geographic: a single human now operates at the scale of a team by delegating to agents, across whichever model backend is cheapest or least attributable that week.
MITRE ATT&CK
| Technique | Name | Observed via |
|---|---|---|
| T1595.002 | Active Scanning: Vulnerability Scanning | nuclei template library |
| T1588.002 | Obtain Capabilities: Tool | assembled responder, nuclei, Serena, agents |
| T1557.001 | Adversary-in-the-Middle: LLMNR/NBT-NS Poisoning & Relay | responder |
| T1071.004 | Application Layer Protocol: DNS | self-hosted OOB collaborator callbacks |
Why it matters
This is what "AI-assisted attacks" actually look like on the ground: not a model writing a novel exploit, but a human pointing a fleet of capable agents at a target and letting them scan, read code, fire payloads and triage callbacks in a loop. It compresses the slow parts of an intrusion, recon and finding-triage, from days to minutes. The gap between "exposed" and "exploited" is shrinking.
Indicators
Defending against it
- Assume faster, wider probing. Deception pays off precisely here: a decoy turns machine-speed scanning into high-confidence signal with no false positives, because any touch is hostile.
- Kill the internal-relay path: disable LLMNR/NBT-NS and enforce SMB signing so responder/relay has nothing to catch.
- Watch your own egress for OOB callbacks (unexpected DNS/LDAP/SMTP to one external token domain), the tell of blind-vuln confirmation.
We find rigs like this the same way we find everything else, by watching where attacks come from and pivoting on what the operator leaves exposed. Browse live finds or start free.
New attacker-infrastructure writeups and live indicators, straight to your inbox. No spam, unsubscribe anytime.
Hunt it yourself
Classify any IP, browse live attacker infrastructure, or deploy your own sensors. Free, no card.