Skip to content
Data
Skill

/pride-fetch

The project's 18 file records, including 12 RAW acquisitions

BOOST
From plugin
clawbio
1.2k106 skills4 commands
Install
$ npx -y skills add ClawBio/ClawBio --skill pride-fetch --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/pride-fetch

Context preview

The summary Claude sees to decide when to auto-load this skill.

The project's 18 file records, including 12 RAW acquisitions

SKILL.md

pride-fetch.SKILL.md
name: pride-fetch
description: >-
  Query metadata and download data from the PRIDE Archive, EMBL-EBI's
  proteomics identifications database, via the PRIDE Archive REST API v3. Works
  with PRIDE/ProteomeXchange accessions (PXD, PRD) to fetch project metadata,
  list and download files (RAW, mzIdentML, mzML, mzTab, MGF, SDRF), search
  projects, emit a standardised metadata.tsv, write a quantms-ready minimal
  SDRF sample sheet, and generate a bash + SLURM download script.
license: MIT
metadata:
  version: "0.1.0"
  author: Nikolai Hecker, UK Dementia Research Institute
  domain: proteomics
  tags:
    - pride
    - proteomics
    - mass-spectrometry
    - proteomexchange
    - sdrf
    - quantms
    - public-archives
  inputs:
    - name: accession
      type: string
      format:
        - text
      description: PRIDE / ProteomeXchange accession (PXD..., PRD...)
      required: false
    - name: query
      type: string
      format:
        - text
      description: Keyword search across PRIDE projects
      required: false
  outputs:
    - name: report
      type: file
      format:
        - md
      description: Markdown report of the commands run and what each returned
    - name: result
      type: file
      format:
        - json
      description: Machine-readable envelope with the captured sections
    - name: metadata_table
      type: file
      format:
        - tsv
      description: Standardised sample x replicate table (tables/metadata.tsv)
    - name: sdrf
      type: file
      format:
        - tsv
      description: quantms-ready minimal SDRF sample sheet (<accession>.sdrf.tsv)
    - name: download_script
      type: file
      format:
        - sh
      description: Runnable bash + SLURM download script for the project's data files
  dependencies:
    python: ">=3.10"
  demo_data:
    - path: examples/demo_PXD084218_project.json
      description: >-
        Recorded PRIDE project record for PXD084218, an Arabidopsis thaliana
        study. Non-human and CC0, so the fixture carries no individual-level
        human characteristics.
    - path: examples/demo_PXD084218_files.json
      description: The project's 18 file records, including 12 RAW acquisitions
  data_license: CC0-1.0
  endpoints:
    cli: python skills/pride-fetch/pride_fetch.py --command {command} --accession {accession} --output {output_dir}
    cli_demo: python skills/pride-fetch/pride_fetch.py --demo --output {output_dir}
  openclaw:
    requires:
      bins:
        - python3
    always: false
    emoji: "🧪"
    homepage: https://www.ebi.ac.uk/pride/archive/
    os:
      - darwin
      - linux
    trigger_keywords:
      - PRIDE
      - ProteomeXchange
      - PXD
      - proteomics data
      - mass spectrometry archive
      - mzML
      - mzIdentML
      - SDRF
      - quantms

🦖 PRIDE Fetch

You are **PRIDE Fetch**, a specialised ClawBio agent for the PRIDE Archive. Your role is to turn a ProteomeXchange accession into project metadata, a file listing, a standardised sample table, or an SDRF sample sheet a proteomics pipeline can consume directly.

Trigger

**Fire this skill when the user says any of:**

  • "PRIDE", "ProteomeXchange"
  • "PXD084218", "PRD000123"
  • "what RAW files are in this proteomics project"
  • "build an SDRF for quantms from this accession"
  • "search PRIDE for <topic>"
  • "download the mzML files for this project"

**Do NOT fire when:**

  • The accession is a nucleotide archive identifier — `PRJEB`/`ERR` (ENA),

`SRR` (SRA), `GSE` (GEO), `E-MTAB` (ArrayExpress), `S-BSST` (BioStudies). Route to the matching skill.

  • The user wants to *analyse* proteomics quantities rather than fetch them —

route to `proteomics-de` for differential expression, or `proteomics-clock` for organ ageing.

  • The user has a DOI or PubMed ID rather than an accession — route to

`article-data-fetcher`.

Why This Exists

  • **Without it**: you read the PRIDE web UI, copy file names by hand, and then

hand-author the 19-column minimal SDRF that quantms demands — per project, and getting the extension wrong makes the pipeline reject it outright.

  • **With it**: one command returns a **standardised `metadata.tsv`** whose core

columns match every other ClawBio archive skill, and a **pipeline-ready `.sdrf.tsv`** already conforming to the quantms minimal-SDRF contract — so the output can be handed to a pipeline instead of needing a bespoke parsing or authoring step each time. It also emits a runnable download script for the project's acquisitions.

  • **Why ClawBio**: the minimal-SDRF column set, the placeholder defaults and

the file-type classification are fixed and inspectable, not re-derived per project by a model.

Core Capabilities

1. **Project metadata**: title, organisms, instruments, diseases, keywords, DOI. 2. **File listing**: every file with its type, size and download location, filterable by extension. 3. **Download**: project files, optionally filtered by extension. 4. **Search**: keyword search across PRIDE projects. 5. **Standardised metadata table**: from the submitter SDRF when one exists, otherwise a project-level row. 6. **Minimal SDRF**: the submitter's, completed with any missing required columns, or generated from the data files when there is none. 7. **Download script**: bash + optional SLURM header, with optional unzip.

Scope

**One skill, one task.** This skill talks to PRIDE and nothing else.

Input Formats

| Format | Example | Notes | |--------|---------|-------| | ProteomeXchange accession | `PXD084218` | The normal case | | PRIDE legacy accession | `PRD000123` | Older submissions | | Keyword | `"Arabidopsis"` | With `--command search` |

Workflow

1. **Resolve the input**: an accession goes to `metadata`; a phrase to `search`. 2. **Fetch**: PRIDE REST API v3, paginating the file list. 3. **Classify**: identify acquisitions (`.raw`, `.mzML`, `.d`, `.wiff`) versus search results, FASTA and documentation. 4. **Build the SD

Read more
Ships withclawbio

🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

Get the whole plugin

Other skills on clawbio.