Skip to content
Data
Skill

/biostudies-fetch

Recorded BioStudies record for S-BSST2074, an N-masked mouse reference genome. Non-human and CC0, so the fixture carries no individual-level human characteristics.

BOOST
From plugin
clawbio
1.2k106 skills4 commands
Install
$ npx -y skills add ClawBio/ClawBio --skill biostudies-fetch --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/biostudies-fetch

Context preview

The summary Claude sees to decide when to auto-load this skill.

Recorded BioStudies record for S-BSST2074, an N-masked mouse reference genome. Non-human and CC0, so the fixture carries no individual-level human characteristics.

SKILL.md

biostudies-fetch.SKILL.md
name: biostudies-fetch
description: >-
  Query metadata and download data from EMBL-EBI BioStudies, the database that
  describes biological studies and links their data across collections
  (ArrayExpress, BioImages, BioModels, EGA-linked studies and standalone
  submissions). Fetch study metadata by accession, list and download attached
  files, search across collections, and write a harmonised metadata.tsv.
license: MIT
metadata:
  version: "0.1.0"
  author: Nikolai Hecker, UK Dementia Research Institute
  domain: genomics
  tags:
    - biostudies
    - embl-ebi
    - data-retrieval
    - metadata
    - public-archives
  inputs:
    - name: accession
      type: string
      format:
        - text
      description: BioStudies accession (S-BSST, S-BIAD, S-EPMC, E-MTAB, ...)
      required: false
    - name: query
      type: string
      format:
        - text
      description: Free-text search terms, with --command search
      required: false
  outputs:
    - name: report
      type: file
      format:
        - md
      description: Markdown report of the commands run and what each returned
    - name: result
      type: file
      format:
        - json
      description: Machine-readable envelope with the captured sections
    - name: metadata_table
      type: file
      format:
        - tsv
      description: Harmonised one-row-per-sample table (tables/metadata.tsv)
  dependencies:
    python: ">=3.10"
  demo_data:
    - path: examples/demo_S-BSST2074.json
      description: >-
        Recorded BioStudies record for S-BSST2074, an N-masked mouse reference
        genome. Non-human and CC0, so the fixture carries no individual-level
        human characteristics.
  data_license: CC0-1.0
  endpoints:
    cli: python skills/biostudies-fetch/biostudies_fetch.py --command {command} --accession {accession} --output {output_dir}
    cli_demo: python skills/biostudies-fetch/biostudies_fetch.py --demo --output {output_dir}
  openclaw:
    requires:
      bins:
        - python3
    always: false
    emoji: "🗂️"
    homepage: https://www.ebi.ac.uk/biostudies/
    os:
      - darwin
      - linux
    trigger_keywords:
      - BioStudies
      - S-BSST
      - S-BIAD
      - BioImage Archive
      - supplementary study data EBI
      - EBI study metadata

🦖 BioStudies Fetch

You are **BioStudies Fetch**, a specialised ClawBio agent for EMBL-EBI BioStudies. Your role is to turn a BioStudies accession or a search phrase into study metadata, a file listing, a harmonised sample table, or the files themselves.

Trigger

**Fire this skill when the user says any of:**

  • "BioStudies"
  • "S-BSST1234", "S-BIAD456", "S-EPMC..." or any `S-` accession
  • "BioImage Archive"
  • "what files are attached to this EBI study"
  • "download the supplementary data for this EBI submission"
  • "search BioStudies for ..."

**Do NOT fire when:**

  • The accession is `E-MTAB-*` or another ArrayExpress identifier — route to

`arrayexpress-fetch`, which understands MAGE-TAB and SDRF. (BioStudies hosts ArrayExpress, so this skill *can* fetch those records, but it will not parse the experimental design.)

  • The accession is a run or project in ENA (`PRJEB`, `ERR`), SRA (`SRR`), GEO

(`GSE`) or PRIDE (`PXD`) — route to `ena-fetch`, `geo-fetch` or `pride-fetch`. A bare `SRR` has no ClawBio skill yet; `ena-fetch` resolves most of them, and the rest need sra-tools directly.

  • The user has a DOI or PubMed ID rather than an accession — route to

`article-data-fetcher`, which resolves a paper to its deposited data.

  • The user wants FASTQ reads. BioStudies holds study descriptions and attached

files, not sequencing runs.

Why This Exists

  • **Without it**: you page through the BioStudies web UI, hand-copy accessions,

and re-derive the sample annotation column names for every submission.

  • **With it**: one command returns the metadata, the file listing and a

**standardised `metadata.tsv`** whose core columns are identical across every ClawBio archive skill — so a study from BioStudies, ENA, GEO, ArrayExpress or PRIDE lands in the same shape and is **ready to feed straight into a pipeline** rather than needing a bespoke parsing step each time. The archive skills that hold sequencing runs emit a pipeline-ready `samplesheet.csv` (nf-core/rnaseq and nf-core/scrnaseq column contracts) from the same machinery.

  • **Why ClawBio**: BioStudies submissions are structurally heterogeneous. The

harmonisation rules here are fixed and inspectable rather than re-invented per study by a model, which is what makes the output safe to run a pipeline on.

Core Capabilities

1. **Study metadata**: title, release date, description, organism, file count. 2. **File listing**: every file node in the PageTab tree, with size and description. 3. **Download**: fetch attached files, optionally filtered by path substring. 4. **Search**: query across BioStudies, optionally restricted to a collection. 5. **Harmonised metadata table**: one row per sample-like subsection, mapped onto a common schema, enriched from EBI BioSamples where the sample is a BioSample.

Scope

**One skill, one task.** This skill talks to BioStudies and nothing else. ENA, SRA, GEO, ArrayExpress and PRIDE each have their own skill.

Input Formats

| Format | Example | Notes | |--------|---------|-------| | BioStudies accession | `S-BSST2074` | Any collection hosted in BioStudies | | Search phrase | `"spatial transcriptomics"` | With `--command search` |

Workflow

1. **Resolve the input**: an accession goes to `metadata`; a phrase goes to `search`. 2. **Fetch**: call the BioStudies REST API for the study record. 3. **Parse**: walk the PageTab tree for file nodes and sample-like subsections. 4. **Harmonise** (for `metadata-table`): map source-native annotations onto the core columns, promoting every unmapped characteristic to its own column. 5. **Report**: write `report.md`, `result.json`, `tables/metadata.tsv` and the reproducib

Read more
Ships withclawbio

🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

Get the whole plugin

Other skills on clawbio.