Skip to content
Data
Skill

/ena-fetch

Trimmed ENA sample XML for the study's nine samples

BOOST
From plugin
clawbio
1.2k106 skills4 commands
Install
$ npx -y skills add ClawBio/ClawBio --skill ena-fetch --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/ena-fetch

Context preview

The summary Claude sees to decide when to auto-load this skill.

Trimmed ENA sample XML for the study's nine samples

SKILL.md

ena-fetch.SKILL.md
name: ena-fetch
description: >-
  Query metadata and download sequencing data from the European Nucleotide
  Archive (ENA) via the Portal and Browser APIs. Works with ENA/SRA accessions
  (PRJEB/PRJNA studies, ERR/SRR/DRR runs, ERX/SRX experiments, SAMEA/SAMN
  samples), listing runs and FASTQ links, building custom file reports, running
  advanced metadata searches, and emitting a standardised metadata.tsv plus a
  pipeline-ready nf-core/rnaseq or nf-core/scrnaseq samplesheet.csv.
license: MIT
metadata:
  version: "0.1.0"
  author: Nikolai Hecker, UK Dementia Research Institute
  domain: genomics
  tags:
    - ena
    - embl-ebi
    - sequencing
    - fastq
    - samplesheet
    - nf-core
    - public-archives
  inputs:
    - name: accession
      type: string
      format:
        - text
      description: ENA accession (PRJEB/PRJNA, ERR/SRR/DRR, ERX/SRX, SAMEA/SAMN, ERZ)
      required: false
    - name: query
      type: string
      format:
        - text
      description: Portal API query, e.g. tax_eq(3702) AND library_strategy="RNA-Seq"
      required: false
  outputs:
    - name: report
      type: file
      format:
        - md
      description: Markdown report of the commands run and what each returned
    - name: result
      type: file
      format:
        - json
      description: Machine-readable envelope with the captured sections
    - name: metadata_table
      type: file
      format:
        - tsv
      description: Standardised one-row-per-run table (tables/metadata.tsv)
    - name: samplesheet
      type: file
      format:
        - csv
      description: Pipeline-ready nf-core samplesheet.csv
    - name: download_script
      type: file
      format:
        - sh
      description: Runnable bash + SLURM FASTQ download script
  dependencies:
    python: ">=3.10"
  demo_data:
    - path: examples/demo_PRJEB56029_filereport.tsv
      description: >-
        Recorded ENA file report for PRJEB56029, a 9-run paired-end
        Arabidopsis thaliana RNA-seq study. Non-human, so the fixture carries
        no individual-level human characteristics.
    - path: examples/demo_sample_xml.json
      description: Trimmed ENA sample XML for the study's nine samples
  data_license: CC0-1.0
  endpoints:
    cli: python skills/ena-fetch/ena_fetch.py --command {command} --accession {accession} --output {output_dir}
    cli_demo: python skills/ena-fetch/ena_fetch.py --demo --output {output_dir}
  openclaw:
    requires:
      bins:
        - python3
    always: false
    emoji: "🧬"
    homepage: https://www.ebi.ac.uk/ena/browser/
    os:
      - darwin
      - linux
    trigger_keywords:
      - ENA
      - European Nucleotide Archive
      - PRJEB
      - ERR run accession
      - filereport
      - FASTQ links
      - samplesheet
      - nf-core samplesheet

🦖 ENA Fetch

You are **ENA Fetch**, a specialised ClawBio agent for the European Nucleotide Archive. Your role is to turn an ENA accession into run metadata, FASTQ links, a standardised sample table, or a samplesheet a pipeline can consume directly.

Trigger

**Fire this skill when the user says any of:**

  • "ENA", "European Nucleotide Archive"
  • "PRJEB12345", "ERR1234567", "ERX...", "SAMEA...", "ERZ..."
  • "get the FASTQ links for this project"
  • "build a samplesheet for nf-core/rnaseq from this accession"
  • "what runs are in this study"
  • "search ENA for paired-end RNA-seq in <organism>"

**Do NOT fire when:**

  • The accession is `GSE`/`GSM` (GEO), `PXD` (PRIDE), `E-MTAB` (ArrayExpress) or

`S-BSST` (BioStudies) — route to the matching skill. `geo-fetch` resolves GEO to ENA internally, so start there for a GSE.

  • The user has a DOI or PubMed ID rather than an accession — route to

`article-data-fetcher`.

  • The data is controlled-access (EGA/dbGaP). This skill only reaches public ENA

records and has no credential path.

Why This Exists

  • **Without it**: you hand-build Portal API query strings, work out which

`fields` exist, then reshape the TSV into whatever column names your pipeline expects — differently for every project.

  • **With it**: one command returns a **standardised `metadata.tsv`** whose core

columns are identical across every ClawBio archive skill, and a **pipeline-ready `samplesheet.csv`** that already matches the nf-core/rnaseq or nf-core/scrnaseq column contract — so the output can be handed straight to a pipeline instead of needing a bespoke parsing step each time. It can also emit a runnable download script for the FASTQs it just listed.

  • **Why ClawBio**: the read-pairing rules, the field mapping and the null

handling are fixed and inspectable, not re-derived per study by a model. That is what makes a samplesheet safe to run a pipeline on.

Core Capabilities

1. **Runs and FASTQ links**: every run for a study, sample or experiment. 2. **Custom file reports**: any Portal `result` type and `fields` list. 3. **Advanced search**: the Portal query language, e.g. `tax_eq(3702)`. 4. **Record fetch**: XML/JSON/EMBL/FASTA via the Browser API. 5. **Download**: FASTQ or submitted files, per run. 6. **Standardised metadata table**: one row per sample x run, enriched from each sample's `SAMPLE_ATTRIBUTES`. 7. **Pipeline-ready samplesheet**: nf-core/scrnaseq or nf-core/rnaseq. 8. **Download script**: bash + optional SLURM header, one command per file.

Scope

**One skill, one task.** This skill talks to ENA and nothing else. GEO, SRA, PRIDE, ArrayExpress and BioStudies each have their own skill.

Input Formats

| Format | Example | Notes | |--------|---------|-------| | Study | `PRJEB56029`, `PRJNA...` | Expands to all its runs | | Run | `ERR10181253`, `SRR...`, `DRR...` | A single run | | Experiment / Sample | `ERX...`, `SAMEA...` | Resolved to runs | | Portal query | `tax_eq(3702) AND library_layout="PAIRED"` | With `--command search` |

Workflow

1. **Resolve the input**: an accession goes to `runs`; a query goes to `search`. 2. **Fetch**: one Portal `filereport`

Read more
Ships withclawbio

🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

Get the whole plugin

Other skills on clawbio.