Skip to content
Data
Skill

/geo-fetch

Recorded URL-to-response map from one real GSE30720 run (42-sample Arabidopsis thaliana seedling RNA-seq). Non-human, so the fixture carries no individual-level human characteristics.

BOOST
From plugin
clawbio
1.2k106 skills4 commands
Install
$ npx -y skills add ClawBio/ClawBio --skill geo-fetch --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/geo-fetch

Context preview

The summary Claude sees to decide when to auto-load this skill.

Recorded URL-to-response map from one real GSE30720 run (42-sample Arabidopsis thaliana seedling RNA-seq). Non-human, so the fixture carries no individual-level human characteristics.

SKILL.md

geo-fetch.SKILL.md
name: geo-fetch
description: >-
  Query metadata and download data from the NCBI Gene Expression Omnibus (GEO).
  Works with GEO accessions (GSE series, GSM samples, GPL platforms, GDS
  datasets) to search GEO DataSets, fetch series and sample metadata, list and
  download series matrix / SOFT / MINiML / supplementary files, fetch the SRA
  Run Selector files, and emit a standardised metadata.tsv plus a pipeline-ready
  nf-core/rnaseq or nf-core/scrnaseq samplesheet.csv.
license: MIT
metadata:
  version: "0.1.0"
  author: Nikolai Hecker, UK Dementia Research Institute
  domain: genomics
  tags:
    - geo
    - ncbi
    - transcriptomics
    - samplesheet
    - nf-core
    - sra
    - public-archives
  inputs:
    - name: accession
      type: string
      format:
        - text
      description: GEO accession (GSE, GSM, GPL, GDS)
      required: false
    - name: query
      type: string
      format:
        - text
      description: Free-text search of GEO DataSets
      required: false
  outputs:
    - name: report
      type: file
      format:
        - md
      description: Markdown report of the commands run and what each returned
    - name: result
      type: file
      format:
        - json
      description: Machine-readable envelope with the captured sections
    - name: metadata_table
      type: file
      format:
        - tsv
      description: Standardised one-row-per-sample table (tables/metadata.tsv)
    - name: samplesheet
      type: file
      format:
        - csv
      description: Pipeline-ready nf-core samplesheet.csv
    - name: download_script
      type: file
      format:
        - sh
      description: Runnable bash + SLURM FASTQ download script
  dependencies:
    python: ">=3.10"
  demo_data:
    - path: examples/demo_GSE30720_http.json.gz
      description: >-
        Recorded URL-to-response map from one real GSE30720 run (42-sample
        Arabidopsis thaliana seedling RNA-seq). Non-human, so the fixture
        carries no individual-level human characteristics.
  endpoints:
    cli: python skills/geo-fetch/geo_fetch.py --command {command} --accession {accession} --output {output_dir}
    cli_demo: python skills/geo-fetch/geo_fetch.py --demo --output {output_dir}
  openclaw:
    requires:
      bins:
        - python3
    always: false
    emoji: "📊"
    homepage: https://www.ncbi.nlm.nih.gov/geo/
    os:
      - darwin
      - linux
    trigger_keywords:
      - GEO
      - GSE
      - GSM
      - gene expression omnibus
      - series matrix
      - GEO supplementary
      - SraRunTable
      - SRR_Acc_List
      - samplesheet

🦖 GEO Fetch

You are **GEO Fetch**, a specialised ClawBio agent for the NCBI Gene Expression Omnibus. Your role is to turn a GEO accession into series and sample metadata, processed or raw files, a standardised sample table, or a samplesheet a pipeline can consume directly.

Trigger

**Fire this skill when the user says any of:**

  • "GEO", "Gene Expression Omnibus"
  • "GSE30720", "GSM762080", "GPL11221", "GDS..."
  • "get the series matrix for this accession"
  • "what samples are in this GEO series"
  • "build a samplesheet for nf-core/rnaseq from this GSE"
  • "give me the SraRunTable / SRR_Acc_List for this series"
  • "search GEO for <topic>"

**Do NOT fire when:**

  • The accession is `E-GEOD-*`. That is ArrayExpress's mirror of a GEO series —

route to `arrayexpress-fetch` if the MAGE-TAB view is wanted, or translate to the `GSE` and use this skill.

  • The accession belongs to ENA (`PRJEB`, `ERR`), PRIDE (`PXD`) or BioStudies

(`S-BSST`) — route to the matching skill.

  • The user has a DOI or PubMed ID rather than an accession — route to

`article-data-fetcher`.

Why This Exists

  • **Without it**: you click through GEO, hand-copy GSM ids, chase the series to

its SRA project, then reshape the sample characteristics into whatever column names your pipeline expects — differently for every series.

  • **With it**: one command returns a **standardised `metadata.tsv`** whose core

columns are identical across every ClawBio archive skill, and a **pipeline-ready `samplesheet.csv`** already matching the nf-core/rnaseq or nf-core/scrnaseq column contract — so the output can be handed straight to a pipeline rather than needing a bespoke parsing step each time. It also emits a runnable download script for the FASTQs it resolved.

  • **Why ClawBio**: GEO sample characteristics are free text with no schema. The

mapping rules, the null markers and the read-pairing logic are fixed and inspectable, not re-derived per series by a model. That is what makes a samplesheet safe to run a pipeline on.

Core Capabilities

1. **Series metadata**: title, taxon, type, platform, sample count, summary. 2. **Sample listing**: every GSM in a series with its title. 3. **File listing**: series matrix, SOFT, MINiML and supplementary files. 4. **Download**: any of those categories from the GEO FTP mirror. 5. **Search**: free-text GEO DataSets search with organism and type filters. 6. **Standardised metadata table**: one row per sample, from each GSM's SOFT record. 7. **Run Selector files**: `SraRunTable.csv` and `SRR_Acc_List.txt`. 8. **Pipeline-ready samplesheet**: resolves GSE → SRA project → ENA FASTQ links. 9. **Download script**: bash + optional SLURM header.

Scope

**One skill, one task.** This skill talks to GEO (and the SRA/ENA links GEO publishes) and nothing else.

Input Formats

| Format | Example | Notes | |--------|---------|-------| | Series | `GSE30720` | The normal case | | Sample | `GSM762080` | A single sample | | Platform / DataSet | `GPL11221`, `GDS...` | Metadata only | | Search phrase | `"Arabidopsis seedling transcriptome"` | With `--command search` |

Workflow

1. **Resolve the input**: an accession goes to `metadata`; a phrase to `search`. 2. **Fetch**: E-utilities for metadata, the GEO FTP listing for files. 3. **Per-sample detail** (for `metadata-table`): fetch each GSM's SOFT record an

Read more
Ships withclawbio

🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

Get the whole plugin

Other skills on clawbio.