Skip to content
Security
Skill

/16-ai-llm-security

LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments

From plugin
claude-code-cybersecurity-skill
42122 skills
Install
$ npx -y skills add Masriyan/Claude-Code-CyberSecurity-Skill --skill 16-ai-llm-security --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/16-ai-llm-security

Context preview

The summary Claude sees to decide when to auto-load this skill.

LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments

SKILL.md

16-ai-llm-security.SKILL.md
name: AI & LLM Security
description: LLM and AI application security testing — prompt injection, jailbreak resistance, OWASP LLM Top 10 (2025), RAG and agent/tool-use security, model supply chain, and AI red teaming for authorized assessments
version: 3.1.0
author: Masriyan
tags: [cybersecurity, ai-security, llm, prompt-injection, owasp-llm, agent-security, rag, mlsecops, red-teaming]

AI & LLM Security

Purpose

Enable Claude to assess the security of AI/LLM-powered applications — chatbots, RAG pipelines, autonomous agents, and tool-using systems. Claude maps findings to the **OWASP Top 10 for LLM Applications (2025)** and the **MITRE ATLAS** adversarial-ML knowledge base, builds reproducible attack cases, and recommends concrete mitigations (input/output guardrails, least-privilege tool scopes, content provenance).

> **Authorization Required**: Only test AI systems you own or are explicitly authorized to assess. Prompt-injection and data-exfiltration testing against third-party AI services may violate their terms of service and local law. Confirm written scope before proceeding.

---

Activation Triggers

This skill activates when the user asks about:

  • Prompt injection (direct or indirect), jailbreaks, or system-prompt extraction
  • OWASP LLM Top 10, MITRE ATLAS, or AI/ML threat modeling
  • Securing a RAG pipeline, vector database, or retrieval layer
  • LLM agent / tool-use / function-calling security and confused-deputy risks
  • Guardrail, content-filter, or model output validation design
  • Sensitive-information disclosure or training-data leakage from a model
  • Model / ML supply chain security (model files, `pickle`, model registries)
  • AI red teaming, jailbreak corpora, or automated adversarial prompt generation
  • Securing MCP (Model Context Protocol) servers and tool integrations

---

Prerequisites

pip install requests pyyaml rich

**Optional enhanced capabilities:**

  • `garak` — LLM vulnerability scanner (NVIDIA)
  • `promptfoo` — prompt/red-team evaluation harness
  • API key for the target LLM endpoint (test environment only)
  • `modelscan` / `picklescan` — ML model file safety scanning

---

Core Capabilities

1. Threat Modeling (OWASP LLM Top 10 — 2025)

When asked to threat-model an AI application, map the system against each category and record exposure:

| ID | Risk | What to look for | |----|------|------------------| | LLM01 | Prompt Injection | Untrusted text reaching the prompt (direct & indirect via RAG/web/email) | | LLM02 | Sensitive Information Disclosure | PII/secrets in prompts, outputs, or training data; system-prompt leakage | | LLM03 | Supply Chain | Untrusted models, LoRA adapters, datasets, plugins, `pickle` deserialization | | LLM04 | Data & Model Poisoning | Tainted training/fine-tune/RAG data; backdoors | | LLM05 | Improper Output Handling | LLM output passed unsanitized to SQL, shell, browser (XSS), or `eval` | | LLM06 | Excessive Agency | Over-broad tool scopes, autonomous side effects, no human-in-the-loop | | LLM07 | System Prompt Leakage | Secrets/authz logic embedded in the system prompt | | LLM08 | Vector & Embedding Weaknesses | RAG access-control bypass, embedding inversion, cross-tenant leakage | | LLM09 | Misinformation | Hallucinations relied on for security/safety decisions | | LLM10 | Unbounded Consumption | Cost/DoS via token floods, model extraction, wallet-drain |

Produce a per-category table: **Exposure (Yes/No/Partial) → Evidence → Severity → Mitigation**.

2. Prompt Injection & Jailbreak Testing

**Direct injection** — user input that overrides instructions. Test families:

  • Instruction override ("ignore previous instructions and …")
  • Role-play / persona escape (DAN-style, hypothetical framing)
  • Encoding/obfuscation (Base64, ROT13, leetspeak, homoglyphs, zero-width chars)
  • Token smuggling and prompt-boundary confusion (fake delimiters, fake system tags)
  • Many-shot jailbreaking (long context of faux dialogue priming compliance)
  • Crescendo / multi-turn gradual escalation

**Indirect injection** — payload arrives via retrieved/processed content (web page, PDF, email, RAG doc, tool output). This is the highest-impact class for agents. Test that retrieved text **cannot** issue commands, exfiltrate context, or trigger tools.

For every test record: payload, channel (direct/indirect), goal (override / exfiltrate / tool-abuse), and result (blocked / partial / success). Use `scripts/prompt_injection_tester.py` to run a corpus and score outcomes.

**Refusal-quality note:** a single refusal is not a pass. Re-test the same goal across ≥3 phrasings and obfuscations before marking a control effective.

3. RAG & Vector Store Security

When reviewing a RAG pipeline: 1. **Access control at retrieval** — confirm the vector query is filtered by the *caller's* permissions, not just the app's. Test cross-tenant / cross-user document leakage. 2. **Indirect injection surface** — treat every ingested document as attacker-controlled; verify retrieved chunks are clearly delimited and never executed as instructions. 3. **Embedding inversion / membership** — sensitive source text may be partially reconstructable from embeddings; flag PII stored unencrypted in the vector DB. 4. **Chunk poisoning** — a single malicious document can dominate retrieval; check ranking/dedup and source allow-listing. 5. **Citation integrity** — outputs should cite retrieved sources so injected claims are traceable.

4. Agent & Tool-Use (Function Calling / MCP) Security

The agent is a **confused deputy**: it holds privileges the user may not. Review:

  • **Least-privilege tools** — each tool scoped to the minimum action; no broad `execute_shell`/`http_request` to arbitrary hosts
  • **Human-in-the-loop gates** on irreversible/outbound actions (payments, email send, file delete, deploy)
  • **Argument validation** — tool args are model-generated and untrusted; validate/allow-list server-side
  • **Injection → tool chain** — verify retrieved/indir
Read more
Ships withclaude-code-cybersecurity-skill

22 production-quality Claude Code Skills for cybersecurity professionals — covering offensive security, defensive operations, reverse engineering, threat hunting, threat intelligence, purple team / adversary emulation, CSOC automation, AI/LLM security,

Get the whole plugin
Stats
417
Stars
77
Forks
Active
Maintenance
Python
Language
MIT
License
8d ago
Last commit
6mo ago
Created

Repo: Masriyan/Claude-Code-CyberSecurity-Skill