Open-source AI agent harnesses · used in real attacks
| Attacker | What it did | ||||||
|---|---|---|---|---|---|---|---|
| 01 | PentestGPTautonomous pentest agent | CONFIRMED | UAT-10147Chinese-speaking, financially motivated | Installed on the attacker’s command server. Scanned web servers, ran PoC exploits, broke into at least one site. | 170,000URLs targeted | Yes | 2026-08-20 |
| 02 | PentAGImulti-agent pentest system | CONFIRMED | GTG-50020Russian-speaking criminalsGTG-50029French-speaking hacktivist | Named by Anthropic as a framework both groups used. One stole production AI API keys; the other dumped European political databases. | 12–26 GBdatabases taken | Yes | 2026-09-10 |
| 03 | OpenClawgeneral-purpose agent | CONFIRMED | “Recon” operatorfinancially motivatedUnnamed operatorSpanish-speaking | Ran a mass credential-harvesting operation. A second operator exploited Telegram Mini Apps and dumped their databases. | 23,800+secrets stolen | Yes | 2026-09-08 |
| 04 | Hermes Agenton DeepSeek, run over Telegram | CONFIRMED | knaithe / KnYuanChinese-speaking | Recon and exploit attempts against Langflow and n8n. The agent’s exploits failed; a Citrix exploit run by hand got in. | 0agent exploits landed | Human only | 2026-07-30 |
| 05 | oh-my-claudecodeadd-on layer for Claude Code | CONFIRMED | Zerofot | Chained with Codex and Claude Code to build a credential harvester that repaired itself when it broke. | 3agents chained | Tooling | 2026-08-31 |
| 06 | HexStrike AIwith Graphiti memory graph | OBSERVED | Suspected China-linked actor | Moved between recon tools on its own against a Japanese tech firm and an East Asian security company. | 2named targets | Not stated | 2026-05-11 |
| 07 | Strixmulti-agent pentest framework | OBSERVED | Suspected China-linked actorplus an attacker on hijacked AI servers | Validated vulnerabilities in the same campaign. A honeypot caught it aimed at a live French shop with no permission prompts. | 2,000+unattended steps | Not stated | 2026-05 · 06 |
| No documented attack yet, despite the coverage: | |||||||
PatternThe agents mostly do recon and vulnerability checks. When an exploit fails, a human finishes the job.
Evidence tiersConfirmed means a vendor, IR firm or CERT shows the harness in a real operation. Observed means it was aimed at real targets, but no breach is stated.
ScopePublished cases, not rates. Commercial agents like Claude Code, Codex and Gemini CLI are left out on purpose.