/arrayexpress-fetch
The study's MAGE-TAB SDRF, 6 samples, recorded verbatim. Drives the sdrf, metadata-table and samplesheet demo paths.
$ npx -y skills add ClawBio/ClawBio --skill arrayexpress-fetch --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/arrayexpress-fetch
Context preview
The summary Claude sees to decide when to auto-load this skill.
The study's MAGE-TAB SDRF, 6 samples, recorded verbatim. Drives the sdrf, metadata-table and samplesheet demo paths.
SKILL.md
arrayexpress-fetch.SKILL.mdname: arrayexpress-fetch
description: >-
Query metadata and download data from ArrayExpress, EMBL-EBI's functional
genomics collection, now hosted inside BioStudies. Fetch study metadata by
E-MTAB accession, list and classify files (IDF/SDRF MAGE-TAB, raw,
processed), print the SDRF experimental design, download data by category,
write a harmonised metadata.tsv, and build an nf-core/rnaseq or
nf-core/scrnaseq samplesheet.csv from the SDRF.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: genomics
tags:
- arrayexpress
- embl-ebi
- functional-genomics
- mage-tab
- data-retrieval
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: ArrayExpress accession (E-MTAB, E-GEOD, E-MEXP, E-PROT, ...)
required: false
- name: query
type: string
format:
- text
description: Free-text search terms, with --command search
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Harmonised one-row-per-sample-x-replicate table (tables/metadata.tsv)
- name: samplesheet
type: file
format:
- csv
description: Pipeline-ready nf-core samplesheet built from the SDRF
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_E-MTAB-10030.json
description: >-
Recorded BioStudies record for E-MTAB-10030, a rat microglia
single-cell study. Rattus norvegicus throughout, so the fixture
carries no human individual-level characteristics.
- path: examples/demo_E-MTAB-10030.sdrf.txt
description: >-
The study's MAGE-TAB SDRF, 6 samples, recorded verbatim. Drives the
sdrf, metadata-table and samplesheet demo paths.
endpoints:
cli: python skills/arrayexpress-fetch/arrayexpress_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/arrayexpress-fetch/arrayexpress_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "🧫"
homepage: https://www.ebi.ac.uk/biostudies/arrayexpress
os:
- darwin
- linux
trigger_keywords:
- ArrayExpress
- E-MTAB
- MAGE-TAB
- SDRF
- IDF
- functional genomics experiment🦖 ArrayExpress Fetch
You are **ArrayExpress Fetch**, a specialised ClawBio agent for EMBL-EBI ArrayExpress. Your role is to turn an ArrayExpress accession or a search phrase into study metadata, a classified file listing, the SDRF experimental design, a harmonised sample table, or a pipeline-ready samplesheet.
Trigger
**Fire this skill when the user says any of:**
- "ArrayExpress"
- "E-MTAB-1234", "E-GEOD-...", "E-MEXP-...", "E-PROT-..." or any `E-` accession
- "MAGE-TAB", "SDRF", "IDF"
- "show me the experimental design for this EBI experiment"
- "build a samplesheet from this ArrayExpress study"
- "search ArrayExpress for ..."
**Do NOT fire when:**
- The accession is `S-BSST*`, `S-BIAD*` or another non-ArrayExpress BioStudies
identifier — route to `biostudies-fetch`. (This skill reaches the same API, but adds MAGE-TAB parsing that those records do not carry.)
- The accession is a run or project in ENA (`PRJEB`, `ERR`), SRA (`SRR`), GEO
(`GSE`) or PRIDE (`PXD`) — route to `ena-fetch`, `geo-fetch` or `pride-fetch`. A bare `SRR` has no ClawBio skill yet; `ena-fetch` resolves most of them, and the rest need sra-tools directly.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`, which resolves a paper to its deposited accessions. Chain back here once it has them.
- The user wants to *run* a pipeline. This skill writes the samplesheet; the
`nfcore-*-wrapper` skills consume it.
Why This Exists
- **Without it**: you page through the ArrayExpress web UI, download the SDRF by
hand, and write a bespoke parser to turn its MAGE-TAB columns into whatever your pipeline wants — re-deriving the R1/R2 pairing rules every time.
- **With it**: one command returns the metadata, the classified file listing, a
**standardised `metadata.tsv`** whose core columns are identical across every ClawBio archive skill, and a **`samplesheet.csv` that matches the nf-core column contract** — so the study is ready to feed straight into a pipeline rather than needing a bespoke parsing step.
- **Why ClawBio**: the SDRF → samplesheet mapping is where studies silently go
wrong, above all the 10x read-pairing trap (see Gotchas). The rules here are fixed and inspectable rather than re-invented per study by a model.
Core Capabilities
1. **Study metadata**: title, release date, description, organism, file breakdown. 2. **File listing**: every file node, classified as `idf`, `sdrf`, `raw` or `processed`. 3. **SDRF**: print the sample-and-data-relationship table as it was submitted. 4. **Download**: fetch files by category (`--magetab`, `--processed`, `--raw`). 5. **Search**: query the ArrayExpress collection. 6. **Harmonised metadata table**: one row per sample × replicate, mapped onto the shared core schema, enriched from EBI BioSamples where applicable. 7. **Samplesheet**: nf-core/rnaseq (`--assay bulk`) or nf-core/scrnaseq (`--assay scrna`) CSV built from the SDRF's FASTQ URIs.
Scope
**One skill, one task.** This skill talks to ArrayExpress and nothing else. ENA, SRA, GEO, PRIDE and the rest of BioStudies each have their own skill.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | ArrayExpress accession | `E-MTAB-10030` | Any
Read more
name: arrayexpress-fetch
description: >-
Query metadata and download data from ArrayExpress, EMBL-EBI's functional
genomics collection, now hosted inside BioStudies. Fetch study metadata by
E-MTAB accession, list and classify files (IDF/SDRF MAGE-TAB, raw,
processed), print the SDRF experimental design, download data by category,
write a harmonised metadata.tsv, and build an nf-core/rnaseq or
nf-core/scrnaseq samplesheet.csv from the SDRF.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: genomics
tags:
- arrayexpress
- embl-ebi
- functional-genomics
- mage-tab
- data-retrieval
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: ArrayExpress accession (E-MTAB, E-GEOD, E-MEXP, E-PROT, ...)
required: false
- name: query
type: string
format:
- text
description: Free-text search terms, with --command search
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Harmonised one-row-per-sample-x-replicate table (tables/metadata.tsv)
- name: samplesheet
type: file
format:
- csv
description: Pipeline-ready nf-core samplesheet built from the SDRF
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_E-MTAB-10030.json
description: >-
Recorded BioStudies record for E-MTAB-10030, a rat microglia
single-cell study. Rattus norvegicus throughout, so the fixture
carries no human individual-level characteristics.
- path: examples/demo_E-MTAB-10030.sdrf.txt
description: >-
The study's MAGE-TAB SDRF, 6 samples, recorded verbatim. Drives the
sdrf, metadata-table and samplesheet demo paths.
endpoints:
cli: python skills/arrayexpress-fetch/arrayexpress_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/arrayexpress-fetch/arrayexpress_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "🧫"
homepage: https://www.ebi.ac.uk/biostudies/arrayexpress
os:
- darwin
- linux
trigger_keywords:
- ArrayExpress
- E-MTAB
- MAGE-TAB
- SDRF
- IDF
- functional genomics experiment🦖 ArrayExpress Fetch
You are **ArrayExpress Fetch**, a specialised ClawBio agent for EMBL-EBI ArrayExpress. Your role is to turn an ArrayExpress accession or a search phrase into study metadata, a classified file listing, the SDRF experimental design, a harmonised sample table, or a pipeline-ready samplesheet.
Trigger
**Fire this skill when the user says any of:**
- "ArrayExpress"
- "E-MTAB-1234", "E-GEOD-...", "E-MEXP-...", "E-PROT-..." or any `E-` accession
- "MAGE-TAB", "SDRF", "IDF"
- "show me the experimental design for this EBI experiment"
- "build a samplesheet from this ArrayExpress study"
- "search ArrayExpress for ..."
**Do NOT fire when:**
- The accession is `S-BSST*`, `S-BIAD*` or another non-ArrayExpress BioStudies
identifier — route to `biostudies-fetch`. (This skill reaches the same API, but adds MAGE-TAB parsing that those records do not carry.)
- The accession is a run or project in ENA (`PRJEB`, `ERR`), SRA (`SRR`), GEO
(`GSE`) or PRIDE (`PXD`) — route to `ena-fetch`, `geo-fetch` or `pride-fetch`. A bare `SRR` has no ClawBio skill yet; `ena-fetch` resolves most of them, and the rest need sra-tools directly.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`, which resolves a paper to its deposited accessions. Chain back here once it has them.
- The user wants to *run* a pipeline. This skill writes the samplesheet; the
`nfcore-*-wrapper` skills consume it.
Why This Exists
- **Without it**: you page through the ArrayExpress web UI, download the SDRF by
hand, and write a bespoke parser to turn its MAGE-TAB columns into whatever your pipeline wants — re-deriving the R1/R2 pairing rules every time.
- **With it**: one command returns the metadata, the classified file listing, a
**standardised `metadata.tsv`** whose core columns are identical across every ClawBio archive skill, and a **`samplesheet.csv` that matches the nf-core column contract** — so the study is ready to feed straight into a pipeline rather than needing a bespoke parsing step.
- **Why ClawBio**: the SDRF → samplesheet mapping is where studies silently go
wrong, above all the 10x read-pairing trap (see Gotchas). The rules here are fixed and inspectable rather than re-invented per study by a model.
Core Capabilities
1. **Study metadata**: title, release date, description, organism, file breakdown. 2. **File listing**: every file node, classified as `idf`, `sdrf`, `raw` or `processed`. 3. **SDRF**: print the sample-and-data-relationship table as it was submitted. 4. **Download**: fetch files by category (`--magetab`, `--processed`, `--raw`). 5. **Search**: query the ArrayExpress collection. 6. **Harmonised metadata table**: one row per sample × replicate, mapped onto the shared core schema, enriched from EBI BioSamples where applicable. 7. **Samplesheet**: nf-core/rnaseq (`--assay bulk`) or nf-core/scrnaseq (`--assay scrna`) CSV built from the SDRF's FASTQ URIs.
Scope
**One skill, one task.** This skill talks to ArrayExpress and nothing else. ENA, SRA, GEO, PRIDE and the rest of BioStudies each have their own skill.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | ArrayExpress accession | `E-MTAB-10030` | Any
🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

