Skip to content
AI & Agents
Skill

/extract

Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site. Use when the user

BOOST
From plugin
adobe-skills
190154 skills4 MCP
Install
$ npx -y skills add adobe/skills --skill extract --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/extract

Context preview

The summary Claude sees to decide when to auto-load this skill.

Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site. Use when the user

SKILL.md

extract.SKILL.md
name: extract
description: Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site. Use when the user wants to analyze an existing site's design, extract or reverse-engineer its design system or brand, capture design tokens from a live site, import a website as the starting point for a redesign, capture the current state before a migration, or invokes `$stardust extract` (`/stardust:extract` in Claude Code). Trigger phrases include "analyze this site", "extract the design tokens", "capture the brand", "crawl the site", "reverse engineer the design". Not for scraping page data or content for its own sake (it captures design evidence, not datasets), and not for the redesign itself — extraction is descriptive; direction and prototyping happen downstream.
license: Apache-2.0
compatibility: Requires Node 22+, Playwright with Chromium resolvable from the project, playwright-cli on PATH, and the impeccable skill (github.com/pbakaus/impeccable) installed alongside stardust.

stardust:extract

Crawl an existing website, parse each page, extract the brand surface, and produce a stardust-formatted snapshot of the current state under `stardust/current/`. The output describes what the site **is**; later sub-commands consume it to decide what it **should be**.

This skill is **descriptive**: it does not invent direction, it does not critique, and it does not modify the live site. It writes only under `stardust/current/` and updates `stardust/state.json`.

Inputs

  • `<url>` — required. The origin to crawl. Examples: `https://example.com`,

`https://example.com/shop`. A path narrows the same-origin crawl to that subtree.

  • `--cap <N>` — optional. Override the default 5-page cap. The cap

is intentionally small — a 5-page sample (home + four IA pillars/templates) is enough for cross-page brand aggregation, system-component detection, and the brand-review HTML; lift it (e.g. `--cap 25`) when a deeper crawl is genuinely needed.

  • `--all` — optional. Lift the cap entirely; extract every

discovered page after junk filtering. Equivalent to `--cap 0`. Use when the user spontaneously asks for a full crawl.

  • `--pages <slug,slug,...>` — optional. Restrict the crawl to specific

paths (slugs derived per `reference/ia-extraction.md`). Bypasses the cap.

  • `--refresh <slug>` — optional. Re-extract one page that already exists

in `state.json`.

  • `--single` — optional. Equivalent to `--cap 1`. Useful for testing.
  • `--wait <fast|medium|spec|auto>` — optional. Wait strategy per page.

Default `medium`. See `reference/playwright-recipe.md` § Wait modes.

  • `--no-junk-filter` — optional. Disable the default junk-page filter

in discovery (see `reference/ia-extraction.md` § Filtering).

  • `--no-consent-dismiss` — optional. Skip the pre-flight consent /

cookie banner dismissal (see `reference/playwright-recipe.md` § Pre-flight: consent dismissal). Use when the redesign scope includes the consent surface or the dismissal's side-effects (script activation that wouldn't otherwise run) must be avoided. Default is to dismiss, keeping screenshots, voice aggregation, and per-section style unpolluted by the banner.

  • `--dynamics` — optional, **migration-bound**. Record per-page reach

signals of the dynamic surface (data endpoints, forms, modal triggers, player ids) in each page JSON `dynamic` section and roll them up in `_crawl-log.json#dynamicSurface`. Set by `prepare-migration`, `replica` and `migrate`'s safety net; never by a bare extract, `uplift` or `audit` — dynamics is a migration concern. Depth and classification belong to the stardust `dynamics` skill.

  • `--concurrency <n>` — optional. Parallel browser contexts for the

per-page capture loop. Default 4; sane range 4–8. See § Concurrency.

  • `--brand-source <url>` — optional, repeatable. An additional

**same-brand** origin whose brand surface enriches the primary extraction (shallow capture: home + up to 2 nav-linked pages). See § Cross-site brand sources.

  • `--design-source <url>` — optional. Design-donor origin: its

design system is captured to `stardust/canon-source/` and becomes the fixed redesign target while the primary origin supplies content. See § Cross-site brand sources.

  • `--prep` — optional. Run in **migrate-prep mode**: lift the cap,

type each page, detect module candidates, capture typed content slots, emit the prep summary. See § Prep mode below. Typically invoked via the `prepare-migration` orchestrator skill rather than directly.

Setup

Run the master skill's setup procedure first (`skills/stardust/SKILL.md` § Setup): impeccable dep check, context loader, state read.

Additional checks for this sub-command:

1. **Playwright availability.** The extraction step needs a real browser. Detect Playwright in this order: a Playwright MCP server, then a project-importable `playwright` module. **The `npx playwright` probe is NOT sufficient** — it confirms the CLI (which resolves a global install) but the recipe and `scripts/crawl.mjs` do `import { chromium } from 'playwright'`, and ESM module resolution does **not** honour a global install or `NODE_PATH` — the import throws `ERR_MODULE_NOT_FOUND` even where `npx playwright --version` succeeds. Verify the module is import-resolvable from the project root (probe: `node -e "import('playwright').then(()=>process.exit(0))"`); if it isn't, run `npm i -D playwright pixelmatch pngjs cheerio --legacy-peer-deps` (devDependencies, never `--no-save` — #125) or use the Playwright MCP server, before crawling. The `--legacy-peer-deps` flag is required on `aem-boilerplate` targets (their pinned `eslint@8` makes a plain `npm i` exit `ERESOLVE` before playwright is even considered). Don't trust the CLI p

Read more
Ships withadobe-skills

Repository of Adobe skills for AI coding agents.

Get the whole plugin

Other skills on adobe-skills.