/biostudies-fetch
Recorded BioStudies record for S-BSST2074, an N-masked mouse reference genome. Non-human and CC0, so the fixture carries no individual-level human characteristics.
$ npx -y skills add ClawBio/ClawBio --skill biostudies-fetch --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/biostudies-fetch
Context preview
The summary Claude sees to decide when to auto-load this skill.
Recorded BioStudies record for S-BSST2074, an N-masked mouse reference genome. Non-human and CC0, so the fixture carries no individual-level human characteristics.
SKILL.md
biostudies-fetch.SKILL.mdname: biostudies-fetch
description: >-
Query metadata and download data from EMBL-EBI BioStudies, the database that
describes biological studies and links their data across collections
(ArrayExpress, BioImages, BioModels, EGA-linked studies and standalone
submissions). Fetch study metadata by accession, list and download attached
files, search across collections, and write a harmonised metadata.tsv.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: genomics
tags:
- biostudies
- embl-ebi
- data-retrieval
- metadata
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: BioStudies accession (S-BSST, S-BIAD, S-EPMC, E-MTAB, ...)
required: false
- name: query
type: string
format:
- text
description: Free-text search terms, with --command search
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Harmonised one-row-per-sample table (tables/metadata.tsv)
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_S-BSST2074.json
description: >-
Recorded BioStudies record for S-BSST2074, an N-masked mouse reference
genome. Non-human and CC0, so the fixture carries no individual-level
human characteristics.
data_license: CC0-1.0
endpoints:
cli: python skills/biostudies-fetch/biostudies_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/biostudies-fetch/biostudies_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "🗂️"
homepage: https://www.ebi.ac.uk/biostudies/
os:
- darwin
- linux
trigger_keywords:
- BioStudies
- S-BSST
- S-BIAD
- BioImage Archive
- supplementary study data EBI
- EBI study metadata🦖 BioStudies Fetch
You are **BioStudies Fetch**, a specialised ClawBio agent for EMBL-EBI BioStudies. Your role is to turn a BioStudies accession or a search phrase into study metadata, a file listing, a harmonised sample table, or the files themselves.
Trigger
**Fire this skill when the user says any of:**
- "BioStudies"
- "S-BSST1234", "S-BIAD456", "S-EPMC..." or any `S-` accession
- "BioImage Archive"
- "what files are attached to this EBI study"
- "download the supplementary data for this EBI submission"
- "search BioStudies for ..."
**Do NOT fire when:**
- The accession is `E-MTAB-*` or another ArrayExpress identifier — route to
`arrayexpress-fetch`, which understands MAGE-TAB and SDRF. (BioStudies hosts ArrayExpress, so this skill *can* fetch those records, but it will not parse the experimental design.)
- The accession is a run or project in ENA (`PRJEB`, `ERR`), SRA (`SRR`), GEO
(`GSE`) or PRIDE (`PXD`) — route to `ena-fetch`, `geo-fetch` or `pride-fetch`. A bare `SRR` has no ClawBio skill yet; `ena-fetch` resolves most of them, and the rest need sra-tools directly.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`, which resolves a paper to its deposited data.
- The user wants FASTQ reads. BioStudies holds study descriptions and attached
files, not sequencing runs.
Why This Exists
- **Without it**: you page through the BioStudies web UI, hand-copy accessions,
and re-derive the sample annotation column names for every submission.
- **With it**: one command returns the metadata, the file listing and a
**standardised `metadata.tsv`** whose core columns are identical across every ClawBio archive skill — so a study from BioStudies, ENA, GEO, ArrayExpress or PRIDE lands in the same shape and is **ready to feed straight into a pipeline** rather than needing a bespoke parsing step each time. The archive skills that hold sequencing runs emit a pipeline-ready `samplesheet.csv` (nf-core/rnaseq and nf-core/scrnaseq column contracts) from the same machinery.
- **Why ClawBio**: BioStudies submissions are structurally heterogeneous. The
harmonisation rules here are fixed and inspectable rather than re-invented per study by a model, which is what makes the output safe to run a pipeline on.
Core Capabilities
1. **Study metadata**: title, release date, description, organism, file count. 2. **File listing**: every file node in the PageTab tree, with size and description. 3. **Download**: fetch attached files, optionally filtered by path substring. 4. **Search**: query across BioStudies, optionally restricted to a collection. 5. **Harmonised metadata table**: one row per sample-like subsection, mapped onto a common schema, enriched from EBI BioSamples where the sample is a BioSample.
Scope
**One skill, one task.** This skill talks to BioStudies and nothing else. ENA, SRA, GEO, ArrayExpress and PRIDE each have their own skill.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | BioStudies accession | `S-BSST2074` | Any collection hosted in BioStudies | | Search phrase | `"spatial transcriptomics"` | With `--command search` |
Workflow
1. **Resolve the input**: an accession goes to `metadata`; a phrase goes to `search`. 2. **Fetch**: call the BioStudies REST API for the study record. 3. **Parse**: walk the PageTab tree for file nodes and sample-like subsections. 4. **Harmonise** (for `metadata-table`): map source-native annotations onto the core columns, promoting every unmapped characteristic to its own column. 5. **Report**: write `report.md`, `result.json`, `tables/metadata.tsv` and the reproducib
Read more
name: biostudies-fetch
description: >-
Query metadata and download data from EMBL-EBI BioStudies, the database that
describes biological studies and links their data across collections
(ArrayExpress, BioImages, BioModels, EGA-linked studies and standalone
submissions). Fetch study metadata by accession, list and download attached
files, search across collections, and write a harmonised metadata.tsv.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: genomics
tags:
- biostudies
- embl-ebi
- data-retrieval
- metadata
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: BioStudies accession (S-BSST, S-BIAD, S-EPMC, E-MTAB, ...)
required: false
- name: query
type: string
format:
- text
description: Free-text search terms, with --command search
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Harmonised one-row-per-sample table (tables/metadata.tsv)
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_S-BSST2074.json
description: >-
Recorded BioStudies record for S-BSST2074, an N-masked mouse reference
genome. Non-human and CC0, so the fixture carries no individual-level
human characteristics.
data_license: CC0-1.0
endpoints:
cli: python skills/biostudies-fetch/biostudies_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/biostudies-fetch/biostudies_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "🗂️"
homepage: https://www.ebi.ac.uk/biostudies/
os:
- darwin
- linux
trigger_keywords:
- BioStudies
- S-BSST
- S-BIAD
- BioImage Archive
- supplementary study data EBI
- EBI study metadata🦖 BioStudies Fetch
You are **BioStudies Fetch**, a specialised ClawBio agent for EMBL-EBI BioStudies. Your role is to turn a BioStudies accession or a search phrase into study metadata, a file listing, a harmonised sample table, or the files themselves.
Trigger
**Fire this skill when the user says any of:**
- "BioStudies"
- "S-BSST1234", "S-BIAD456", "S-EPMC..." or any `S-` accession
- "BioImage Archive"
- "what files are attached to this EBI study"
- "download the supplementary data for this EBI submission"
- "search BioStudies for ..."
**Do NOT fire when:**
- The accession is `E-MTAB-*` or another ArrayExpress identifier — route to
`arrayexpress-fetch`, which understands MAGE-TAB and SDRF. (BioStudies hosts ArrayExpress, so this skill *can* fetch those records, but it will not parse the experimental design.)
- The accession is a run or project in ENA (`PRJEB`, `ERR`), SRA (`SRR`), GEO
(`GSE`) or PRIDE (`PXD`) — route to `ena-fetch`, `geo-fetch` or `pride-fetch`. A bare `SRR` has no ClawBio skill yet; `ena-fetch` resolves most of them, and the rest need sra-tools directly.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`, which resolves a paper to its deposited data.
- The user wants FASTQ reads. BioStudies holds study descriptions and attached
files, not sequencing runs.
Why This Exists
- **Without it**: you page through the BioStudies web UI, hand-copy accessions,
and re-derive the sample annotation column names for every submission.
- **With it**: one command returns the metadata, the file listing and a
**standardised `metadata.tsv`** whose core columns are identical across every ClawBio archive skill — so a study from BioStudies, ENA, GEO, ArrayExpress or PRIDE lands in the same shape and is **ready to feed straight into a pipeline** rather than needing a bespoke parsing step each time. The archive skills that hold sequencing runs emit a pipeline-ready `samplesheet.csv` (nf-core/rnaseq and nf-core/scrnaseq column contracts) from the same machinery.
- **Why ClawBio**: BioStudies submissions are structurally heterogeneous. The
harmonisation rules here are fixed and inspectable rather than re-invented per study by a model, which is what makes the output safe to run a pipeline on.
Core Capabilities
1. **Study metadata**: title, release date, description, organism, file count. 2. **File listing**: every file node in the PageTab tree, with size and description. 3. **Download**: fetch attached files, optionally filtered by path substring. 4. **Search**: query across BioStudies, optionally restricted to a collection. 5. **Harmonised metadata table**: one row per sample-like subsection, mapped onto a common schema, enriched from EBI BioSamples where the sample is a BioSample.
Scope
**One skill, one task.** This skill talks to BioStudies and nothing else. ENA, SRA, GEO, ArrayExpress and PRIDE each have their own skill.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | BioStudies accession | `S-BSST2074` | Any collection hosted in BioStudies | | Search phrase | `"spatial transcriptomics"` | With `--command search` |
Workflow
1. **Resolve the input**: an accession goes to `metadata`; a phrase goes to `search`. 2. **Fetch**: call the BioStudies REST API for the study record. 3. **Parse**: walk the PageTab tree for file nodes and sample-like subsections. 4. **Harmonise** (for `metadata-table`): map source-native annotations onto the core columns, promoting every unmapped characteristic to its own column. 5. **Report**: write `report.md`, `result.json`, `tables/metadata.tsv` and the reproducib
🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

