Skip to content
Security
Skill

/llm-redteam

LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling.

BOOST
From plugin
agentic-bug-hunter
5.3k19 skills9 agents41 commands3 hooks
Install
$ npx -y skills add awarexone/agentic-bug-hunter --skill llm-redteam --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/llm-redteam

Context preview

The summary Claude sees to decide when to auto-load this skill.

LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling.

SKILL.md

llm-redteam.SKILL.md
name: llm-redteam
description: LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling. Canonical OWASP LLM Top 10 + ASI01-ASI10 mapping. Use when a target exposes a chat/completions/assistant/copilot endpoint, an AI feature that consumes user text or documents, or any /v1/chat, /api/chat, /mcp surface.

LLM Red-Team

> Reflection is not exploitation. A model that *echoes* your payload is noise; a model that *acts* on it — leaks its system prompt, emits your canary, calls a tool, renders raw HTML downstream — is a bug. Prove the action, then chain it to concrete impact.

This is the first-class LLM-attack skill. The scanner behind it is `tools/llm_redteam.py` (canary-based corpus runner, `/llm-redteam`). For auditing an MCP *server*, use `skills/mcp-server-audit`; for attacking a *deployed agent* black-box, use `skills/agentic-app-audit`.

0. QUICK KILL CHECKLIST

Kill the lead (classify Informational, move on) when:

  • The model refuses and no canary / leak / tool-call ever lands. "It said something edgy" is not a finding.
  • System-prompt "leak" is just generic assistant boilerplate ("I am a helpful assistant"), not the target's actual instructions.
  • Prompt injection fires but reads/changes nothing the current user couldn't already access (no cross-tenant data, no tool with impact).
  • Output looks like XSS but the host app escapes it before rendering (check the actual downstream sink, not the chat response).
  • Exfil payload is a markdown image beacon but the client never fetches remote images.

1. ROUTING TABLE — signal on target -> what to run

| Signal | Run | |---|---| | Any chat/assistant endpoint | `tools/llm_redteam.py --url <endpoint> --field <json-field>` (full corpus) | | OpenAI-style API | `--template '{"messages":[{"role":"user","content":"{{PAYLOAD}}"}]}' --response-path choices.0.message.content` | | Only want one class | `--category jailbreak` (see `--list-categories`) | | Uploads/RAG/"summarize this doc" | indirect-injection + token-smuggling (invisible unicode in the doc) | | Agent has tools | excessive-agency category + `skills/agentic-app-audit` | | Confirm blind exfil | plant a canary host, correlate with `tools/oob_listener.py` |

2. ATTACK CLASSES (corpus categories)

The runner fires a canary per run (`RT_PWNED_xxxx`); a category "lands" when the canary / a real leak / a tool-call shows up in the response.

  • **prompt-injection** — direct instruction override, delimiter breaks.
  • **jailbreak** — DAN / developer-mode persona escape.
  • **system-prompt-leak** — repeat-above, keyword-anchor, scenario-escape.
  • **data-exfil** — markdown-image beacon / canary emission to an attacker host.
  • **indirect-injection** — payload framed as a retrieved document / tool result (the high-value one for RAG apps).
  • **guardrail-bypass** — base64 / split-string instruction smuggling.
  • **multi-turn-crescendo** — single-shot simulation of a gradual escalation jailbreak.
  • **cross-lingual** — override instruction in another language to dodge English-only filters.
  • **cipher-obfuscation** — leetspeak / ROT13-wrapped instruction.
  • **token-smuggling** — invisible "Sneaky Bits" unicode (U+2062/U+2064) carrying a hidden instruction (shared encoder with `tools/hai_payload_builder.py`).
  • **excessive-agency** — coerce an unsafe tool call / exfil (SSRF via "fetch URL", email send, payment).
  • **insecure-output** — get the model to emit raw HTML/script that a downstream sink renders (XSS via AI).

3. CANONICAL OWASP ASI01-ASI10 (single source of truth)

This table is the ONE authoritative ASI mapping for the whole toolkit — `skills/bug-bounty` and `skills/web2-vuln-classes` both defer to it. If you edit the taxonomy, edit it here.

| ID | Class | What to test | Chain to | |----|-------|--------------|----------| | ASI01 | Prompt Injection / Goal Hijack | Override objectives via direct or indirect injection | IDOR / exfil / tool abuse | | ASI02 | Tool Misuse | Attacker-controlled tool params ("fetch this URL", code tool) | SSRF / RCE | | ASI03 | Privilege Compromise | Agent uses broader perms / admin tokens than the user | Priv-esc / cross-tenant | | ASI04 | Supply Chain | Compromised plugin / MCP server / tool-output poisoning next agent | RCE / data theft | | ASI05 | Code Execution | Unsafe code-interpreter / sandbox escape | RCE | | ASI06 | Memory & Context Poisoning | Persistent RAG/memory corruption across sessions/users | Stored injection affecting all users | | ASI07 | Agent Communication | Inter-agent spoofing / IDOR (agent A reads agent B's context) | Cross-tenant disclosure | | ASI08 | Excessive Agency | Destructive action without confirmation; cascading failures | Funds/email/delete | | ASI09 | Insecure Output Handling | AI output rendered as XSS / SQLi / command injection downstream | XSS / injection | | ASI10 | Sensitive Information Disclosure | Leaks system prompt / keys / configs / user data | Secrets -> deeper access |

**Triage rule:** ASI alone = Informational. It is a bounty only when chained to IDOR / exfil / RCE / ATO with a working PoC.

4. CHAIN TO IMPACT (what makes it payable)

| Standalone | Chain | Result | |---|---|---| | Prompt injection | + reads another user's conversation/data (IDOR) | High | | Indirect injection (RAG doc) | + persists in memory -> hits every user | High/Critical | | Tool misuse "fetch URL" | + hits internal service / IMDS and returns data | SSRF (Medium/High) | | Insecure output | + the host renders it -> stored XSS -> session theft | High | | System-prompt leak | + the prompt contains an API key / internal URL | High | | Excessive agency | + unconfirmed destructive action (send funds, delete) | High/Critical |

5. CONFIRMATION DISCIPLINE (no false positives)

  • A category only counts when the **canary /
Read more
Ships withagentic-bug-hunter

AI-powered bug bounty hunting toolkit that works with or without subscription.

Get the whole plugin

Other skills on agentic-bug-hunter.