/ena-fetch
Trimmed ENA sample XML for the study's nine samples
$ npx -y skills add ClawBio/ClawBio --skill ena-fetch --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/ena-fetch
Context preview
The summary Claude sees to decide when to auto-load this skill.
Trimmed ENA sample XML for the study's nine samples
SKILL.md
ena-fetch.SKILL.mdname: ena-fetch
description: >-
Query metadata and download sequencing data from the European Nucleotide
Archive (ENA) via the Portal and Browser APIs. Works with ENA/SRA accessions
(PRJEB/PRJNA studies, ERR/SRR/DRR runs, ERX/SRX experiments, SAMEA/SAMN
samples), listing runs and FASTQ links, building custom file reports, running
advanced metadata searches, and emitting a standardised metadata.tsv plus a
pipeline-ready nf-core/rnaseq or nf-core/scrnaseq samplesheet.csv.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: genomics
tags:
- ena
- embl-ebi
- sequencing
- fastq
- samplesheet
- nf-core
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: ENA accession (PRJEB/PRJNA, ERR/SRR/DRR, ERX/SRX, SAMEA/SAMN, ERZ)
required: false
- name: query
type: string
format:
- text
description: Portal API query, e.g. tax_eq(3702) AND library_strategy="RNA-Seq"
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Standardised one-row-per-run table (tables/metadata.tsv)
- name: samplesheet
type: file
format:
- csv
description: Pipeline-ready nf-core samplesheet.csv
- name: download_script
type: file
format:
- sh
description: Runnable bash + SLURM FASTQ download script
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_PRJEB56029_filereport.tsv
description: >-
Recorded ENA file report for PRJEB56029, a 9-run paired-end
Arabidopsis thaliana RNA-seq study. Non-human, so the fixture carries
no individual-level human characteristics.
- path: examples/demo_sample_xml.json
description: Trimmed ENA sample XML for the study's nine samples
data_license: CC0-1.0
endpoints:
cli: python skills/ena-fetch/ena_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/ena-fetch/ena_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "🧬"
homepage: https://www.ebi.ac.uk/ena/browser/
os:
- darwin
- linux
trigger_keywords:
- ENA
- European Nucleotide Archive
- PRJEB
- ERR run accession
- filereport
- FASTQ links
- samplesheet
- nf-core samplesheet🦖 ENA Fetch
You are **ENA Fetch**, a specialised ClawBio agent for the European Nucleotide Archive. Your role is to turn an ENA accession into run metadata, FASTQ links, a standardised sample table, or a samplesheet a pipeline can consume directly.
Trigger
**Fire this skill when the user says any of:**
- "ENA", "European Nucleotide Archive"
- "PRJEB12345", "ERR1234567", "ERX...", "SAMEA...", "ERZ..."
- "get the FASTQ links for this project"
- "build a samplesheet for nf-core/rnaseq from this accession"
- "what runs are in this study"
- "search ENA for paired-end RNA-seq in <organism>"
**Do NOT fire when:**
- The accession is `GSE`/`GSM` (GEO), `PXD` (PRIDE), `E-MTAB` (ArrayExpress) or
`S-BSST` (BioStudies) — route to the matching skill. `geo-fetch` resolves GEO to ENA internally, so start there for a GSE.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`.
- The data is controlled-access (EGA/dbGaP). This skill only reaches public ENA
records and has no credential path.
Why This Exists
- **Without it**: you hand-build Portal API query strings, work out which
`fields` exist, then reshape the TSV into whatever column names your pipeline expects — differently for every project.
- **With it**: one command returns a **standardised `metadata.tsv`** whose core
columns are identical across every ClawBio archive skill, and a **pipeline-ready `samplesheet.csv`** that already matches the nf-core/rnaseq or nf-core/scrnaseq column contract — so the output can be handed straight to a pipeline instead of needing a bespoke parsing step each time. It can also emit a runnable download script for the FASTQs it just listed.
- **Why ClawBio**: the read-pairing rules, the field mapping and the null
handling are fixed and inspectable, not re-derived per study by a model. That is what makes a samplesheet safe to run a pipeline on.
Core Capabilities
1. **Runs and FASTQ links**: every run for a study, sample or experiment. 2. **Custom file reports**: any Portal `result` type and `fields` list. 3. **Advanced search**: the Portal query language, e.g. `tax_eq(3702)`. 4. **Record fetch**: XML/JSON/EMBL/FASTA via the Browser API. 5. **Download**: FASTQ or submitted files, per run. 6. **Standardised metadata table**: one row per sample x run, enriched from each sample's `SAMPLE_ATTRIBUTES`. 7. **Pipeline-ready samplesheet**: nf-core/scrnaseq or nf-core/rnaseq. 8. **Download script**: bash + optional SLURM header, one command per file.
Scope
**One skill, one task.** This skill talks to ENA and nothing else. GEO, SRA, PRIDE, ArrayExpress and BioStudies each have their own skill.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | Study | `PRJEB56029`, `PRJNA...` | Expands to all its runs | | Run | `ERR10181253`, `SRR...`, `DRR...` | A single run | | Experiment / Sample | `ERX...`, `SAMEA...` | Resolved to runs | | Portal query | `tax_eq(3702) AND library_layout="PAIRED"` | With `--command search` |
Workflow
1. **Resolve the input**: an accession goes to `runs`; a query goes to `search`. 2. **Fetch**: one Portal `filereport`
Read more
name: ena-fetch
description: >-
Query metadata and download sequencing data from the European Nucleotide
Archive (ENA) via the Portal and Browser APIs. Works with ENA/SRA accessions
(PRJEB/PRJNA studies, ERR/SRR/DRR runs, ERX/SRX experiments, SAMEA/SAMN
samples), listing runs and FASTQ links, building custom file reports, running
advanced metadata searches, and emitting a standardised metadata.tsv plus a
pipeline-ready nf-core/rnaseq or nf-core/scrnaseq samplesheet.csv.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: genomics
tags:
- ena
- embl-ebi
- sequencing
- fastq
- samplesheet
- nf-core
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: ENA accession (PRJEB/PRJNA, ERR/SRR/DRR, ERX/SRX, SAMEA/SAMN, ERZ)
required: false
- name: query
type: string
format:
- text
description: Portal API query, e.g. tax_eq(3702) AND library_strategy="RNA-Seq"
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Standardised one-row-per-run table (tables/metadata.tsv)
- name: samplesheet
type: file
format:
- csv
description: Pipeline-ready nf-core samplesheet.csv
- name: download_script
type: file
format:
- sh
description: Runnable bash + SLURM FASTQ download script
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_PRJEB56029_filereport.tsv
description: >-
Recorded ENA file report for PRJEB56029, a 9-run paired-end
Arabidopsis thaliana RNA-seq study. Non-human, so the fixture carries
no individual-level human characteristics.
- path: examples/demo_sample_xml.json
description: Trimmed ENA sample XML for the study's nine samples
data_license: CC0-1.0
endpoints:
cli: python skills/ena-fetch/ena_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/ena-fetch/ena_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "🧬"
homepage: https://www.ebi.ac.uk/ena/browser/
os:
- darwin
- linux
trigger_keywords:
- ENA
- European Nucleotide Archive
- PRJEB
- ERR run accession
- filereport
- FASTQ links
- samplesheet
- nf-core samplesheet🦖 ENA Fetch
You are **ENA Fetch**, a specialised ClawBio agent for the European Nucleotide Archive. Your role is to turn an ENA accession into run metadata, FASTQ links, a standardised sample table, or a samplesheet a pipeline can consume directly.
Trigger
**Fire this skill when the user says any of:**
- "ENA", "European Nucleotide Archive"
- "PRJEB12345", "ERR1234567", "ERX...", "SAMEA...", "ERZ..."
- "get the FASTQ links for this project"
- "build a samplesheet for nf-core/rnaseq from this accession"
- "what runs are in this study"
- "search ENA for paired-end RNA-seq in <organism>"
**Do NOT fire when:**
- The accession is `GSE`/`GSM` (GEO), `PXD` (PRIDE), `E-MTAB` (ArrayExpress) or
`S-BSST` (BioStudies) — route to the matching skill. `geo-fetch` resolves GEO to ENA internally, so start there for a GSE.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`.
- The data is controlled-access (EGA/dbGaP). This skill only reaches public ENA
records and has no credential path.
Why This Exists
- **Without it**: you hand-build Portal API query strings, work out which
`fields` exist, then reshape the TSV into whatever column names your pipeline expects — differently for every project.
- **With it**: one command returns a **standardised `metadata.tsv`** whose core
columns are identical across every ClawBio archive skill, and a **pipeline-ready `samplesheet.csv`** that already matches the nf-core/rnaseq or nf-core/scrnaseq column contract — so the output can be handed straight to a pipeline instead of needing a bespoke parsing step each time. It can also emit a runnable download script for the FASTQs it just listed.
- **Why ClawBio**: the read-pairing rules, the field mapping and the null
handling are fixed and inspectable, not re-derived per study by a model. That is what makes a samplesheet safe to run a pipeline on.
Core Capabilities
1. **Runs and FASTQ links**: every run for a study, sample or experiment. 2. **Custom file reports**: any Portal `result` type and `fields` list. 3. **Advanced search**: the Portal query language, e.g. `tax_eq(3702)`. 4. **Record fetch**: XML/JSON/EMBL/FASTA via the Browser API. 5. **Download**: FASTQ or submitted files, per run. 6. **Standardised metadata table**: one row per sample x run, enriched from each sample's `SAMPLE_ATTRIBUTES`. 7. **Pipeline-ready samplesheet**: nf-core/scrnaseq or nf-core/rnaseq. 8. **Download script**: bash + optional SLURM header, one command per file.
Scope
**One skill, one task.** This skill talks to ENA and nothing else. GEO, SRA, PRIDE, ArrayExpress and BioStudies each have their own skill.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | Study | `PRJEB56029`, `PRJNA...` | Expands to all its runs | | Run | `ERR10181253`, `SRR...`, `DRR...` | A single run | | Experiment / Sample | `ERX...`, `SAMEA...` | Resolved to runs | | Portal query | `tax_eq(3702) AND library_layout="PAIRED"` | With `--command search` |
Workflow
1. **Resolve the input**: an accession goes to `runs`; a query goes to `search`. 2. **Fetch**: one Portal `filereport`
🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

