/s3-explore
Explore and query data on S3, Cloudflare R2, GCS, MinIO, or any S3-compatible storage. Use when the user mentions an s3://, r2://, gs://, or gcs:// URL, asks "what's in this bucket", wants to list remote files, preview remote Parquet/CSV/JSON, or query data on object storage
$ npx -y skills add duckdb/duckdb-skills --skill s3-explore --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/s3-explore
Context preview
The summary Claude sees to decide when to auto-load this skill.
Explore and query data on S3, Cloudflare R2, GCS, MinIO, or any S3-compatible storage. Use when the user mentions an s3://, r2://, gs://, or gcs:// URL, asks "what's in this bucket", wants to list remote files, preview remote Parquet/CSV/JSON, or query data on object storage
SKILL.md
s3-explore.SKILL.mdname: s3-explore
description: >
Explore and query data on S3, Cloudflare R2, GCS, MinIO, or any S3-compatible storage.
Use when the user mentions an s3://, r2://, gs://, or gcs:// URL, asks "what's in this bucket",
wants to list remote files, preview remote Parquet/CSV/JSON, or query data on object storage
without downloading it. Also triggers when the user wants to know the size, schema, or row count
of remote datasets.
argument-hint: <s3-url> [question about the data]
allowed-tools: Bash
You are helping the user explore data on remote object storage using DuckDB.
URL: `$0` Question: `${1:-list and describe what's there}`
Step 1 — Detect provider and set up credentials
Based on the URL or user context, prepend the appropriate secret configuration:
| Provider | URL patterns | Secret setup | |---|---|---| | **AWS S3** | `s3://` | `CREATE SECRET (TYPE S3, PROVIDER credential_chain);` | | **Cloudflare R2** | `r2://`, `s3://` with R2 endpoint | `CREATE SECRET (TYPE R2, PROVIDER credential_chain);` | | **GCS** | `gs://`, `gcs://` | `CREATE SECRET (TYPE GCS, PROVIDER credential_chain);` | | **MinIO / custom** | `s3://` with custom endpoint | `CREATE SECRET (TYPE S3, KEY_ID '...', SECRET '...', ENDPOINT '...', USE_SSL true);` |
For R2, if the user provides an account ID, the endpoint is `<account_id>.r2.cloudflarestorage.com`. R2 URLs like `r2://bucket/path` should be rewritten to `s3://bucket/path` with the R2 secret.
For public buckets (e.g., Overture Maps, AWS open data), no secret is needed — skip this step.
Always prepend:
LOAD httpfs;
Step 2 — Determine what the URL points to
If the URL looks like a **directory or bucket** (no file extension, or ends with `/`), list its contents with sizes:
duckdb -c "
LOAD httpfs;
<SECRET_SETUP>
SELECT filename, (size / 1024 / 1024)::DECIMAL(10,1) AS size_mb, last_modified
FROM read_blob('<URL>/*')
ORDER BY filename
LIMIT 50;
"Note: only select `filename`, `size`, `last_modified` — never select `content`, which would download the actual files.
If the URL points to a **specific file or glob pattern** (has a file extension or contains `*`), preview it:
duckdb -c "
LOAD httpfs;
<SECRET_SETUP>
DESCRIBE FROM '<URL>';
SELECT count(*) AS row_count FROM '<URL>';
FROM '<URL>' LIMIT 20;
"
For **Parquet files**, get row counts and sizes from metadata (no data download):
duckdb -c "
LOAD httpfs;
<SECRET_SETUP>
SELECT file_name,
sum(row_group_num_rows) AS total_rows,
(sum(row_group_compressed_bytes) / 1024 / 1024)::DECIMAL(10,1) AS compressed_mb
FROM parquet_metadata('<URL>')
GROUP BY file_name;
"Step 3 — Answer the question
Using the listing, schema, or sample data, answer:
`${1:-list and describe what's there}`
If the user asks an analytical question (e.g., "how many rows match X"), write and run the appropriate SQL query. DuckDB pushes predicates down into Parquet on S3, so filtering is efficient even on large remote datasets.
Error handling
- **`duckdb: command not found`** → delegate to `/duckdb-skills:install-duckdb`
- **Access denied / 403** → suggest the user check credentials: `aws configure`, environment variables, or provide explicit key/secret
- **Bucket not found / 404** → check the URL and region
- **Timeout on large listing** → suggest narrowing the glob pattern or adding a prefix
Read more
name: s3-explore description: > Explore and query data on S3, Cloudflare R2, GCS, MinIO, or any S3-compatible storage. Use when the user mentions an s3://, r2://, gs://, or gcs:// URL, asks "what's in this bucket", wants to list remote files, preview remote Parquet/CSV/JSON, or query data on object storage without downloading it. Also triggers when the user wants to know the size, schema, or row count of remote datasets. argument-hint: <s3-url> [question about the data] allowed-tools: Bash
You are helping the user explore data on remote object storage using DuckDB.
URL: `$0` Question: `${1:-list and describe what's there}`
Step 1 — Detect provider and set up credentials
Based on the URL or user context, prepend the appropriate secret configuration:
| Provider | URL patterns | Secret setup | |---|---|---| | **AWS S3** | `s3://` | `CREATE SECRET (TYPE S3, PROVIDER credential_chain);` | | **Cloudflare R2** | `r2://`, `s3://` with R2 endpoint | `CREATE SECRET (TYPE R2, PROVIDER credential_chain);` | | **GCS** | `gs://`, `gcs://` | `CREATE SECRET (TYPE GCS, PROVIDER credential_chain);` | | **MinIO / custom** | `s3://` with custom endpoint | `CREATE SECRET (TYPE S3, KEY_ID '...', SECRET '...', ENDPOINT '...', USE_SSL true);` |
For R2, if the user provides an account ID, the endpoint is `<account_id>.r2.cloudflarestorage.com`. R2 URLs like `r2://bucket/path` should be rewritten to `s3://bucket/path` with the R2 secret.
For public buckets (e.g., Overture Maps, AWS open data), no secret is needed — skip this step.
Always prepend:
LOAD httpfs;
Step 2 — Determine what the URL points to
If the URL looks like a **directory or bucket** (no file extension, or ends with `/`), list its contents with sizes:
duckdb -c "
LOAD httpfs;
<SECRET_SETUP>
SELECT filename, (size / 1024 / 1024)::DECIMAL(10,1) AS size_mb, last_modified
FROM read_blob('<URL>/*')
ORDER BY filename
LIMIT 50;
"Note: only select `filename`, `size`, `last_modified` — never select `content`, which would download the actual files.
If the URL points to a **specific file or glob pattern** (has a file extension or contains `*`), preview it:
duckdb -c " LOAD httpfs; <SECRET_SETUP> DESCRIBE FROM '<URL>'; SELECT count(*) AS row_count FROM '<URL>'; FROM '<URL>' LIMIT 20; "
For **Parquet files**, get row counts and sizes from metadata (no data download):
duckdb -c "
LOAD httpfs;
<SECRET_SETUP>
SELECT file_name,
sum(row_group_num_rows) AS total_rows,
(sum(row_group_compressed_bytes) / 1024 / 1024)::DECIMAL(10,1) AS compressed_mb
FROM parquet_metadata('<URL>')
GROUP BY file_name;
"Step 3 — Answer the question
Using the listing, schema, or sample data, answer:
`${1:-list and describe what's there}`
If the user asks an analytical question (e.g., "how many rows match X"), write and run the appropriate SQL query. DuckDB pushes predicates down into Parquet on S3, so filtering is efficient even on large remote datasets.
Error handling
- **`duckdb: command not found`** → delegate to `/duckdb-skills:install-duckdb`
- **Access denied / 403** → suggest the user check credentials: `aws configure`, environment variables, or provide explicit key/secret
- **Bucket not found / 404** → check the URL and region
- **Timeout on large listing** → suggest narrowing the glob pattern or adding a prefix
A Claude Code plugin that adds DuckDB-powered skills for data exploration and session memory.
Repo: duckdb/duckdb-skills
Other skills on duckdb-skills.
- /attach-db
Attach a DuckDB database file for use with /duckdb-skills:query. Explores the schema (tables, columns, row counts) and writes a SQL state file so subsequent queries can restore this session automatically via duckdb -init.
Open skill - /convert-file
Convert any data file to another format: CSV, Parquet, JSON, Excel, GeoJSON, and more. Use when the user says "convert to parquet", "save as xlsx", "export as JSON", "make this a CSV", "turn into parquet", or any variation of format-to-format conversion for data files. Also
Open skill - /duckdb-docs
Search DuckDB and DuckLake documentation and blog posts. Returns relevant doc chunks for a question or keyword using full-text search against a locally cached index.
Open skill - /install-duckdb
Install or update DuckDB extensions. Each argument is either a plain extension name (installs from core) or name@repo (e.g. magic@community). Pass --update to update extensions instead of installing.
Open skill - /query
Run SQL queries against the attached DuckDB database or ad-hoc against files. Accepts raw SQL or natural language questions. Uses DuckDB Friendly SQL idioms.
Open skill - /read-file
Read any data file (CSV, JSON, Parquet, Avro, Excel, spatial, SQLite) or remote URL (S3, HTTPS). Use when user references a data file, asks "what's in this file", or wants to preview/profile a dataset. Not for source code.
Open skill

