/pride-fetch
The project's 18 file records, including 12 RAW acquisitions
$ npx -y skills add ClawBio/ClawBio --skill pride-fetch --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/pride-fetch
Context preview
The summary Claude sees to decide when to auto-load this skill.
The project's 18 file records, including 12 RAW acquisitions
SKILL.md
pride-fetch.SKILL.mdname: pride-fetch
description: >-
Query metadata and download data from the PRIDE Archive, EMBL-EBI's
proteomics identifications database, via the PRIDE Archive REST API v3. Works
with PRIDE/ProteomeXchange accessions (PXD, PRD) to fetch project metadata,
list and download files (RAW, mzIdentML, mzML, mzTab, MGF, SDRF), search
projects, emit a standardised metadata.tsv, write a quantms-ready minimal
SDRF sample sheet, and generate a bash + SLURM download script.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: proteomics
tags:
- pride
- proteomics
- mass-spectrometry
- proteomexchange
- sdrf
- quantms
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: PRIDE / ProteomeXchange accession (PXD..., PRD...)
required: false
- name: query
type: string
format:
- text
description: Keyword search across PRIDE projects
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Standardised sample x replicate table (tables/metadata.tsv)
- name: sdrf
type: file
format:
- tsv
description: quantms-ready minimal SDRF sample sheet (<accession>.sdrf.tsv)
- name: download_script
type: file
format:
- sh
description: Runnable bash + SLURM download script for the project's data files
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_PXD084218_project.json
description: >-
Recorded PRIDE project record for PXD084218, an Arabidopsis thaliana
study. Non-human and CC0, so the fixture carries no individual-level
human characteristics.
- path: examples/demo_PXD084218_files.json
description: The project's 18 file records, including 12 RAW acquisitions
data_license: CC0-1.0
endpoints:
cli: python skills/pride-fetch/pride_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/pride-fetch/pride_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "🧪"
homepage: https://www.ebi.ac.uk/pride/archive/
os:
- darwin
- linux
trigger_keywords:
- PRIDE
- ProteomeXchange
- PXD
- proteomics data
- mass spectrometry archive
- mzML
- mzIdentML
- SDRF
- quantms🦖 PRIDE Fetch
You are **PRIDE Fetch**, a specialised ClawBio agent for the PRIDE Archive. Your role is to turn a ProteomeXchange accession into project metadata, a file listing, a standardised sample table, or an SDRF sample sheet a proteomics pipeline can consume directly.
Trigger
**Fire this skill when the user says any of:**
- "PRIDE", "ProteomeXchange"
- "PXD084218", "PRD000123"
- "what RAW files are in this proteomics project"
- "build an SDRF for quantms from this accession"
- "search PRIDE for <topic>"
- "download the mzML files for this project"
**Do NOT fire when:**
- The accession is a nucleotide archive identifier — `PRJEB`/`ERR` (ENA),
`SRR` (SRA), `GSE` (GEO), `E-MTAB` (ArrayExpress), `S-BSST` (BioStudies). Route to the matching skill.
- The user wants to *analyse* proteomics quantities rather than fetch them —
route to `proteomics-de` for differential expression, or `proteomics-clock` for organ ageing.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`.
Why This Exists
- **Without it**: you read the PRIDE web UI, copy file names by hand, and then
hand-author the 19-column minimal SDRF that quantms demands — per project, and getting the extension wrong makes the pipeline reject it outright.
- **With it**: one command returns a **standardised `metadata.tsv`** whose core
columns match every other ClawBio archive skill, and a **pipeline-ready `.sdrf.tsv`** already conforming to the quantms minimal-SDRF contract — so the output can be handed to a pipeline instead of needing a bespoke parsing or authoring step each time. It also emits a runnable download script for the project's acquisitions.
- **Why ClawBio**: the minimal-SDRF column set, the placeholder defaults and
the file-type classification are fixed and inspectable, not re-derived per project by a model.
Core Capabilities
1. **Project metadata**: title, organisms, instruments, diseases, keywords, DOI. 2. **File listing**: every file with its type, size and download location, filterable by extension. 3. **Download**: project files, optionally filtered by extension. 4. **Search**: keyword search across PRIDE projects. 5. **Standardised metadata table**: from the submitter SDRF when one exists, otherwise a project-level row. 6. **Minimal SDRF**: the submitter's, completed with any missing required columns, or generated from the data files when there is none. 7. **Download script**: bash + optional SLURM header, with optional unzip.
Scope
**One skill, one task.** This skill talks to PRIDE and nothing else.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | ProteomeXchange accession | `PXD084218` | The normal case | | PRIDE legacy accession | `PRD000123` | Older submissions | | Keyword | `"Arabidopsis"` | With `--command search` |
Workflow
1. **Resolve the input**: an accession goes to `metadata`; a phrase to `search`. 2. **Fetch**: PRIDE REST API v3, paginating the file list. 3. **Classify**: identify acquisitions (`.raw`, `.mzML`, `.d`, `.wiff`) versus search results, FASTA and documentation. 4. **Build the SD
Read more
name: pride-fetch
description: >-
Query metadata and download data from the PRIDE Archive, EMBL-EBI's
proteomics identifications database, via the PRIDE Archive REST API v3. Works
with PRIDE/ProteomeXchange accessions (PXD, PRD) to fetch project metadata,
list and download files (RAW, mzIdentML, mzML, mzTab, MGF, SDRF), search
projects, emit a standardised metadata.tsv, write a quantms-ready minimal
SDRF sample sheet, and generate a bash + SLURM download script.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: proteomics
tags:
- pride
- proteomics
- mass-spectrometry
- proteomexchange
- sdrf
- quantms
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: PRIDE / ProteomeXchange accession (PXD..., PRD...)
required: false
- name: query
type: string
format:
- text
description: Keyword search across PRIDE projects
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Standardised sample x replicate table (tables/metadata.tsv)
- name: sdrf
type: file
format:
- tsv
description: quantms-ready minimal SDRF sample sheet (<accession>.sdrf.tsv)
- name: download_script
type: file
format:
- sh
description: Runnable bash + SLURM download script for the project's data files
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_PXD084218_project.json
description: >-
Recorded PRIDE project record for PXD084218, an Arabidopsis thaliana
study. Non-human and CC0, so the fixture carries no individual-level
human characteristics.
- path: examples/demo_PXD084218_files.json
description: The project's 18 file records, including 12 RAW acquisitions
data_license: CC0-1.0
endpoints:
cli: python skills/pride-fetch/pride_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/pride-fetch/pride_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "🧪"
homepage: https://www.ebi.ac.uk/pride/archive/
os:
- darwin
- linux
trigger_keywords:
- PRIDE
- ProteomeXchange
- PXD
- proteomics data
- mass spectrometry archive
- mzML
- mzIdentML
- SDRF
- quantms🦖 PRIDE Fetch
You are **PRIDE Fetch**, a specialised ClawBio agent for the PRIDE Archive. Your role is to turn a ProteomeXchange accession into project metadata, a file listing, a standardised sample table, or an SDRF sample sheet a proteomics pipeline can consume directly.
Trigger
**Fire this skill when the user says any of:**
- "PRIDE", "ProteomeXchange"
- "PXD084218", "PRD000123"
- "what RAW files are in this proteomics project"
- "build an SDRF for quantms from this accession"
- "search PRIDE for <topic>"
- "download the mzML files for this project"
**Do NOT fire when:**
- The accession is a nucleotide archive identifier — `PRJEB`/`ERR` (ENA),
`SRR` (SRA), `GSE` (GEO), `E-MTAB` (ArrayExpress), `S-BSST` (BioStudies). Route to the matching skill.
- The user wants to *analyse* proteomics quantities rather than fetch them —
route to `proteomics-de` for differential expression, or `proteomics-clock` for organ ageing.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`.
Why This Exists
- **Without it**: you read the PRIDE web UI, copy file names by hand, and then
hand-author the 19-column minimal SDRF that quantms demands — per project, and getting the extension wrong makes the pipeline reject it outright.
- **With it**: one command returns a **standardised `metadata.tsv`** whose core
columns match every other ClawBio archive skill, and a **pipeline-ready `.sdrf.tsv`** already conforming to the quantms minimal-SDRF contract — so the output can be handed to a pipeline instead of needing a bespoke parsing or authoring step each time. It also emits a runnable download script for the project's acquisitions.
- **Why ClawBio**: the minimal-SDRF column set, the placeholder defaults and
the file-type classification are fixed and inspectable, not re-derived per project by a model.
Core Capabilities
1. **Project metadata**: title, organisms, instruments, diseases, keywords, DOI. 2. **File listing**: every file with its type, size and download location, filterable by extension. 3. **Download**: project files, optionally filtered by extension. 4. **Search**: keyword search across PRIDE projects. 5. **Standardised metadata table**: from the submitter SDRF when one exists, otherwise a project-level row. 6. **Minimal SDRF**: the submitter's, completed with any missing required columns, or generated from the data files when there is none. 7. **Download script**: bash + optional SLURM header, with optional unzip.
Scope
**One skill, one task.** This skill talks to PRIDE and nothing else.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | ProteomeXchange accession | `PXD084218` | The normal case | | PRIDE legacy accession | `PRD000123` | Older submissions | | Keyword | `"Arabidopsis"` | With `--command search` |
Workflow
1. **Resolve the input**: an accession goes to `metadata`; a phrase to `search`. 2. **Fetch**: PRIDE REST API v3, paginating the file list. 3. **Classify**: identify acquisitions (`.raw`, `.mzML`, `.d`, `.wiff`) versus search results, FASTA and documentation. 4. **Build the SD
🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

