Skip to content
Data
Skill

/arrayexpress-fetch

The study's MAGE-TAB SDRF, 6 samples, recorded verbatim. Drives the sdrf, metadata-table and samplesheet demo paths.

BOOST
From plugin
clawbio
1.2k106 skills4 commands
Install
$ npx -y skills add ClawBio/ClawBio --skill arrayexpress-fetch --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/arrayexpress-fetch

Context preview

The summary Claude sees to decide when to auto-load this skill.

The study's MAGE-TAB SDRF, 6 samples, recorded verbatim. Drives the sdrf, metadata-table and samplesheet demo paths.

SKILL.md

arrayexpress-fetch.SKILL.md
name: arrayexpress-fetch
description: >-
  Query metadata and download data from ArrayExpress, EMBL-EBI's functional
  genomics collection, now hosted inside BioStudies. Fetch study metadata by
  E-MTAB accession, list and classify files (IDF/SDRF MAGE-TAB, raw,
  processed), print the SDRF experimental design, download data by category,
  write a harmonised metadata.tsv, and build an nf-core/rnaseq or
  nf-core/scrnaseq samplesheet.csv from the SDRF.
license: MIT
metadata:
  version: "0.1.0"
  author: Nikolai Hecker, UK Dementia Research Institute
  domain: genomics
  tags:
    - arrayexpress
    - embl-ebi
    - functional-genomics
    - mage-tab
    - data-retrieval
    - public-archives
  inputs:
    - name: accession
      type: string
      format:
        - text
      description: ArrayExpress accession (E-MTAB, E-GEOD, E-MEXP, E-PROT, ...)
      required: false
    - name: query
      type: string
      format:
        - text
      description: Free-text search terms, with --command search
      required: false
  outputs:
    - name: report
      type: file
      format:
        - md
      description: Markdown report of the commands run and what each returned
    - name: result
      type: file
      format:
        - json
      description: Machine-readable envelope with the captured sections
    - name: metadata_table
      type: file
      format:
        - tsv
      description: Harmonised one-row-per-sample-x-replicate table (tables/metadata.tsv)
    - name: samplesheet
      type: file
      format:
        - csv
      description: Pipeline-ready nf-core samplesheet built from the SDRF
  dependencies:
    python: ">=3.10"
  demo_data:
    - path: examples/demo_E-MTAB-10030.json
      description: >-
        Recorded BioStudies record for E-MTAB-10030, a rat microglia
        single-cell study. Rattus norvegicus throughout, so the fixture
        carries no human individual-level characteristics.
    - path: examples/demo_E-MTAB-10030.sdrf.txt
      description: >-
        The study's MAGE-TAB SDRF, 6 samples, recorded verbatim. Drives the
        sdrf, metadata-table and samplesheet demo paths.
  endpoints:
    cli: python skills/arrayexpress-fetch/arrayexpress_fetch.py --command {command} --accession {accession} --output {output_dir}
    cli_demo: python skills/arrayexpress-fetch/arrayexpress_fetch.py --demo --output {output_dir}
  openclaw:
    requires:
      bins:
        - python3
    always: false
    emoji: "🧫"
    homepage: https://www.ebi.ac.uk/biostudies/arrayexpress
    os:
      - darwin
      - linux
    trigger_keywords:
      - ArrayExpress
      - E-MTAB
      - MAGE-TAB
      - SDRF
      - IDF
      - functional genomics experiment

🦖 ArrayExpress Fetch

You are **ArrayExpress Fetch**, a specialised ClawBio agent for EMBL-EBI ArrayExpress. Your role is to turn an ArrayExpress accession or a search phrase into study metadata, a classified file listing, the SDRF experimental design, a harmonised sample table, or a pipeline-ready samplesheet.

Trigger

**Fire this skill when the user says any of:**

  • "ArrayExpress"
  • "E-MTAB-1234", "E-GEOD-...", "E-MEXP-...", "E-PROT-..." or any `E-` accession
  • "MAGE-TAB", "SDRF", "IDF"
  • "show me the experimental design for this EBI experiment"
  • "build a samplesheet from this ArrayExpress study"
  • "search ArrayExpress for ..."

**Do NOT fire when:**

  • The accession is `S-BSST*`, `S-BIAD*` or another non-ArrayExpress BioStudies

identifier — route to `biostudies-fetch`. (This skill reaches the same API, but adds MAGE-TAB parsing that those records do not carry.)

  • The accession is a run or project in ENA (`PRJEB`, `ERR`), SRA (`SRR`), GEO

(`GSE`) or PRIDE (`PXD`) — route to `ena-fetch`, `geo-fetch` or `pride-fetch`. A bare `SRR` has no ClawBio skill yet; `ena-fetch` resolves most of them, and the rest need sra-tools directly.

  • The user has a DOI or PubMed ID rather than an accession — route to

`article-data-fetcher`, which resolves a paper to its deposited accessions. Chain back here once it has them.

  • The user wants to *run* a pipeline. This skill writes the samplesheet; the

`nfcore-*-wrapper` skills consume it.

Why This Exists

  • **Without it**: you page through the ArrayExpress web UI, download the SDRF by

hand, and write a bespoke parser to turn its MAGE-TAB columns into whatever your pipeline wants — re-deriving the R1/R2 pairing rules every time.

  • **With it**: one command returns the metadata, the classified file listing, a

**standardised `metadata.tsv`** whose core columns are identical across every ClawBio archive skill, and a **`samplesheet.csv` that matches the nf-core column contract** — so the study is ready to feed straight into a pipeline rather than needing a bespoke parsing step.

  • **Why ClawBio**: the SDRF → samplesheet mapping is where studies silently go

wrong, above all the 10x read-pairing trap (see Gotchas). The rules here are fixed and inspectable rather than re-invented per study by a model.

Core Capabilities

1. **Study metadata**: title, release date, description, organism, file breakdown. 2. **File listing**: every file node, classified as `idf`, `sdrf`, `raw` or `processed`. 3. **SDRF**: print the sample-and-data-relationship table as it was submitted. 4. **Download**: fetch files by category (`--magetab`, `--processed`, `--raw`). 5. **Search**: query the ArrayExpress collection. 6. **Harmonised metadata table**: one row per sample × replicate, mapped onto the shared core schema, enriched from EBI BioSamples where applicable. 7. **Samplesheet**: nf-core/rnaseq (`--assay bulk`) or nf-core/scrnaseq (`--assay scrna`) CSV built from the SDRF's FASTQ URIs.

Scope

**One skill, one task.** This skill talks to ArrayExpress and nothing else. ENA, SRA, GEO, PRIDE and the rest of BioStudies each have their own skill.

Input Formats

| Format | Example | Notes | |--------|---------|-------| | ArrayExpress accession | `E-MTAB-10030` | Any

Read more
Ships withclawbio

🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

Get the whole plugin

Other skills on clawbio.