Skip to content
Development
Skill

/harvest-structured

Structured data extraction - tables, pricing, products, API endpoints with schema

From plugin
vibecosystem
532200 skills138 agents7 hooks
Install
$ npx -y skills add vibeeval/vibecosystem --skill harvest-structured --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/harvest-structured

Context preview

The summary Claude sees to decide when to auto-load this skill.

Structured data extraction - tables, pricing, products, API endpoints with schema

SKILL.md

harvest-structured.SKILL.md
name: harvest-structured
description: Structured data extraction - tables, pricing, products, API endpoints with schema
allowed-tools: [Bash, Read, Write, WebFetch]
keywords: [scrape, structured, data, extract, schema, table, pricing, product, json, csv]

Harvest Structured

Extract structured data from web pages using user-defined schemas. Turns messy HTML into clean JSON/CSV - pricing tables, product listings, API endpoint docs, comparison matrices.

Usage

/scrape <url> --schema "<field descriptions>"

Examples

# Extract pricing data
/scrape https://example.com/pricing --schema "plan_name, price, features[], cta_text"

# Extract product listings
/scrape https://store.example.com/products --schema "name, price, rating, reviews_count, image_url"

# Extract API endpoints
/scrape https://docs.api.com/reference --schema "method, path, description, parameters[], response_code"

Schema Definition

Define fields as comma-separated names. Use `[]` for arrays:

name            → Single text value
price           → Single value (auto-detects currency)
features[]      → Array of items
description     → Long text
url             → Auto-detects links
image_url       → Auto-detects image sources

How It Works

1. Fetch page content 2. Parse schema definition 3. Use CSS selectors or LLM extraction to match fields 4. Validate extracted data against schema 5. Output as JSON (default) or CSV

Output Format

JSON (default)

[
  {
    "plan_name": "Pro",
    "price": "$29/mo",
    "features": ["Unlimited projects", "Priority support", "API access"],
    "source_url": "https://example.com/pricing"
  }
]

CSV

plan_name,price,features,source_url
Pro,"$29/mo","Unlimited projects; Priority support; API access",https://example.com/pricing

Integration

  • **growth**: Competitor pricing extraction
  • **migrator**: Changelog/breaking changes extraction
  • **tech-radar**: Feature comparison across tools
  • **data-analyst**: Structured data for analysis

Rules

  • Only extract publicly visible data
  • Respect rate limits (1 req/sec)
  • Validate schema before extraction
  • Report confidence per field (high/medium/low)
  • Output includes source URL for every record
Read more
Ships withvibecosystem

Your AI software team. Built on Claude Code. vibecosystem turns Claude Code into a full AI software team — 138 specialized agents that plan, build, review, test, and learn from every mistake. No configuration needed — just install and code.

Get the whole plugin

Other skills on vibecosystem.