Skip to content
Marketing
Skill

/agent-readiness-audit

Audit whether AI agents and AI crawlers can actually use a site — robots.txt rules per AI crawler token (OpenAI, Anthropic and Perplexity bots, Google-Extended, Applebot-Extended), Product/Offer/Organization/FAQ JSON-LD, whether main content is in the no-JavaScript server HTML,

BOOST
From plugin
digital-marketing-pro
846164 skills24 agents5 commands
Install
$ npx -y skills add indranilbanerjee/digital-marketing-pro --skill agent-readiness-audit --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/agent-readiness-audit

Context preview

The summary Claude sees to decide when to auto-load this skill.

Audit whether AI agents and AI crawlers can actually use a site — robots.txt rules per AI crawler token (OpenAI, Anthropic and Perplexity bots, Google-Extended, Applebot-Extended), Product/Offer/Organization/FAQ JSON-LD, whether main content is in the no-JavaScript server HTML,

SKILL.md

agent-readiness-audit.SKILL.md
name: agent-readiness-audit
description: "Audit whether AI agents and AI crawlers can actually use a site — robots.txt rules per AI crawler token (OpenAI, Anthropic and Perplexity bots, Google-Extended, Applebot-Extended), Product/Offer/Organization/FAQ JSON-LD, whether main content is in the no-JavaScript server HTML, Merchant Center feed completeness incl. native_commerce checkout eligibility and conversational attributes, an optional agentic-commerce feed, and an optional experimental WebMCP check. Triggers on \"/digital-marketing-pro:agent-readiness-audit\", \"can AI agents use our site\", \"are we blocking GPTBot or ClaudeBot\", \"is our product feed ready for AI Mode shopping\", \"run an agent-readiness check\". Runs agent-readiness-audit.py offline on exports (network only with --fetch) and never recommends llms.txt for Google."
argument-hint: "[site URL, or paths to robots.txt / HTML / feed exports]"

/digital-marketing-pro:agent-readiness-audit

Purpose

Answer one question with evidence: **can AI agents and AI crawlers use this site?** Concretely, the audit checks five things:

  • AI crawlers are allowed to reach the content.
  • The content is in the HTML the server sends, so it does not depend on JavaScript running.
  • Structured data says what the page is.
  • The product feed is complete enough for AI shopping surfaces.
  • Optionally, the site exposes agent tools (WebMCP).

Every check is deterministic and runs in `scripts/agent-readiness-audit.py`. The skill adds interpretation and the brand's policy decisions on top of the script's output; it never replaces that output with judgment.

What the primary sources say (checked 2026-10-04)

From Google's [AI optimization guide](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) (last updated 2026-07-10):

  • **No special files.** Under the heading "Mythbusting generative AI search", Google lists "LLMS.txt files and other 'special' markup". It says "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." **This skill never recommends llms.txt for Google.**
  • **Structured data:** "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add." The audit checks JSON-LD because it powers **rich results** (merchant listings, organization info), not because AI features require it.
  • **JavaScript:** "Google is able to process content within JavaScript as long as it isn't blocked." The no-JS check is a robustness test: content that is already in the server HTML can be read by every crawler and agent, whether or not it runs JavaScript.
  • **Agents:** browser agents may analyze screenshots, inspect the DOM structure, and interpret the accessibility tree. Hence the accessibility-basics check. Google names the Universal Commerce Protocol (UCP) as an emerging protocol.

AI crawler tokens (each verified on the vendor's own page)

| Token | Vendor | Role in the audit | What blocking it means (vendor's words, paraphrased) | Source | |---|---|---|---|---| | `OAI-SearchBot` | OpenAI | search: **must be allowed** | Site not surfaced in ChatGPT search features | [OpenAI bots](https://developers.openai.com/api/docs/bots) | | `ChatGPT-User` | OpenAI | user fetch: **must be allowed** | User-initiated visits. OpenAI: robots.txt "may not apply" to these | same | | `OAI-AdsBot` | OpenAI | ads: **must be allowed** if the brand runs ChatGPT Ads | Validates the safety of pages submitted as ChatGPT ads | same | | `GPTBot` | OpenAI | training: **policy** | Content excluded from foundation-model training | same | | `Claude-SearchBot` | Anthropic | search: **must be allowed** | Content not indexed for Claude's search results | [Anthropic](https://support.claude.com/en/articles/8896518) | | `Claude-User` | Anthropic | user fetch: **must be allowed** | Claude can't retrieve the page for a user's question. Anthropic honors robots.txt for all three of its bots | same | | `ClaudeBot` | Anthropic | training: **policy** | Future content excluded from training datasets | same | | `PerplexityBot` | Perplexity | search: **must be allowed** | Not surfaced or linked in Perplexity results (Perplexity says it is not used for model training) | [Perplexity bots](https://docs.perplexity.ai/guides/bots) | | `Perplexity-User` | Perplexity | user fetch: **must be allowed** | Perplexity says this fetcher "generally ignores robots.txt rules" | same | | `Google-Extended` | Google | training: **policy** | Controls use for Gemini training and grounding. It "does not impact a site's inclusion in Google Search nor is it used as a ranking signal", and it has no separate user-agent string | [Google crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers) | | `Applebot-Extended` | Apple | training: **policy** | Controls use in Apple foundation-model training. It does not crawl, and pages that disallow it "can still be included in search results" | [Apple](https://support.apple.com/en-us/119829) |

"Policy" means that blocking training crawlers is the brand's choice. The audit reports it as info unless the brand sets `--training-policy allow|block`, in which case a mismatch fails the check.

Inputs

All checks run **offline** on files the user exports. The script fetches over the network only with `--fetch`.

  • **robots.txt**: `--robots FILE`, or `--site URL --fetch`. Add `--path /products/` (repeatable) to test the paths that matter, beyond `/`.
  • **Server HTML**: `--html FILE` (repeatable). Save it as the server sends it, for example with `curl -L URL > page.html`, **not** "Save as" from a browser, which saves the post-JavaScript DOM. Add `--expect "Product name"` (repeatable) for text that must be in that HTML, such as a product name or price.
  • **Merchant Center feed export**: `--feed FILE` (TSV, CSV, or RSS/Atom XM
Read more
Ships withdigital-marketing-pro

Your agency just signed a 50-brand client. The previous agency left no playbook. Three brands are bleeding budget, two have stale positioning, one is launching in a regulated jurisdiction next month. Where do you start?

Get the whole plugin

Other skills on digital-marketing-pro.