/geo-fetch
Recorded URL-to-response map from one real GSE30720 run (42-sample Arabidopsis thaliana seedling RNA-seq). Non-human, so the fixture carries no individual-level human characteristics.
$ npx -y skills add ClawBio/ClawBio --skill geo-fetch --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/geo-fetch
Context preview
The summary Claude sees to decide when to auto-load this skill.
Recorded URL-to-response map from one real GSE30720 run (42-sample Arabidopsis thaliana seedling RNA-seq). Non-human, so the fixture carries no individual-level human characteristics.
SKILL.md
geo-fetch.SKILL.mdname: geo-fetch
description: >-
Query metadata and download data from the NCBI Gene Expression Omnibus (GEO).
Works with GEO accessions (GSE series, GSM samples, GPL platforms, GDS
datasets) to search GEO DataSets, fetch series and sample metadata, list and
download series matrix / SOFT / MINiML / supplementary files, fetch the SRA
Run Selector files, and emit a standardised metadata.tsv plus a pipeline-ready
nf-core/rnaseq or nf-core/scrnaseq samplesheet.csv.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: genomics
tags:
- geo
- ncbi
- transcriptomics
- samplesheet
- nf-core
- sra
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: GEO accession (GSE, GSM, GPL, GDS)
required: false
- name: query
type: string
format:
- text
description: Free-text search of GEO DataSets
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Standardised one-row-per-sample table (tables/metadata.tsv)
- name: samplesheet
type: file
format:
- csv
description: Pipeline-ready nf-core samplesheet.csv
- name: download_script
type: file
format:
- sh
description: Runnable bash + SLURM FASTQ download script
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_GSE30720_http.json.gz
description: >-
Recorded URL-to-response map from one real GSE30720 run (42-sample
Arabidopsis thaliana seedling RNA-seq). Non-human, so the fixture
carries no individual-level human characteristics.
endpoints:
cli: python skills/geo-fetch/geo_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/geo-fetch/geo_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "📊"
homepage: https://www.ncbi.nlm.nih.gov/geo/
os:
- darwin
- linux
trigger_keywords:
- GEO
- GSE
- GSM
- gene expression omnibus
- series matrix
- GEO supplementary
- SraRunTable
- SRR_Acc_List
- samplesheet🦖 GEO Fetch
You are **GEO Fetch**, a specialised ClawBio agent for the NCBI Gene Expression Omnibus. Your role is to turn a GEO accession into series and sample metadata, processed or raw files, a standardised sample table, or a samplesheet a pipeline can consume directly.
Trigger
**Fire this skill when the user says any of:**
- "GEO", "Gene Expression Omnibus"
- "GSE30720", "GSM762080", "GPL11221", "GDS..."
- "get the series matrix for this accession"
- "what samples are in this GEO series"
- "build a samplesheet for nf-core/rnaseq from this GSE"
- "give me the SraRunTable / SRR_Acc_List for this series"
- "search GEO for <topic>"
**Do NOT fire when:**
- The accession is `E-GEOD-*`. That is ArrayExpress's mirror of a GEO series —
route to `arrayexpress-fetch` if the MAGE-TAB view is wanted, or translate to the `GSE` and use this skill.
- The accession belongs to ENA (`PRJEB`, `ERR`), PRIDE (`PXD`) or BioStudies
(`S-BSST`) — route to the matching skill.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`.
Why This Exists
- **Without it**: you click through GEO, hand-copy GSM ids, chase the series to
its SRA project, then reshape the sample characteristics into whatever column names your pipeline expects — differently for every series.
- **With it**: one command returns a **standardised `metadata.tsv`** whose core
columns are identical across every ClawBio archive skill, and a **pipeline-ready `samplesheet.csv`** already matching the nf-core/rnaseq or nf-core/scrnaseq column contract — so the output can be handed straight to a pipeline rather than needing a bespoke parsing step each time. It also emits a runnable download script for the FASTQs it resolved.
- **Why ClawBio**: GEO sample characteristics are free text with no schema. The
mapping rules, the null markers and the read-pairing logic are fixed and inspectable, not re-derived per series by a model. That is what makes a samplesheet safe to run a pipeline on.
Core Capabilities
1. **Series metadata**: title, taxon, type, platform, sample count, summary. 2. **Sample listing**: every GSM in a series with its title. 3. **File listing**: series matrix, SOFT, MINiML and supplementary files. 4. **Download**: any of those categories from the GEO FTP mirror. 5. **Search**: free-text GEO DataSets search with organism and type filters. 6. **Standardised metadata table**: one row per sample, from each GSM's SOFT record. 7. **Run Selector files**: `SraRunTable.csv` and `SRR_Acc_List.txt`. 8. **Pipeline-ready samplesheet**: resolves GSE → SRA project → ENA FASTQ links. 9. **Download script**: bash + optional SLURM header.
Scope
**One skill, one task.** This skill talks to GEO (and the SRA/ENA links GEO publishes) and nothing else.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | Series | `GSE30720` | The normal case | | Sample | `GSM762080` | A single sample | | Platform / DataSet | `GPL11221`, `GDS...` | Metadata only | | Search phrase | `"Arabidopsis seedling transcriptome"` | With `--command search` |
Workflow
1. **Resolve the input**: an accession goes to `metadata`; a phrase to `search`. 2. **Fetch**: E-utilities for metadata, the GEO FTP listing for files. 3. **Per-sample detail** (for `metadata-table`): fetch each GSM's SOFT record an
Read more
name: geo-fetch
description: >-
Query metadata and download data from the NCBI Gene Expression Omnibus (GEO).
Works with GEO accessions (GSE series, GSM samples, GPL platforms, GDS
datasets) to search GEO DataSets, fetch series and sample metadata, list and
download series matrix / SOFT / MINiML / supplementary files, fetch the SRA
Run Selector files, and emit a standardised metadata.tsv plus a pipeline-ready
nf-core/rnaseq or nf-core/scrnaseq samplesheet.csv.
license: MIT
metadata:
version: "0.1.0"
author: Nikolai Hecker, UK Dementia Research Institute
domain: genomics
tags:
- geo
- ncbi
- transcriptomics
- samplesheet
- nf-core
- sra
- public-archives
inputs:
- name: accession
type: string
format:
- text
description: GEO accession (GSE, GSM, GPL, GDS)
required: false
- name: query
type: string
format:
- text
description: Free-text search of GEO DataSets
required: false
outputs:
- name: report
type: file
format:
- md
description: Markdown report of the commands run and what each returned
- name: result
type: file
format:
- json
description: Machine-readable envelope with the captured sections
- name: metadata_table
type: file
format:
- tsv
description: Standardised one-row-per-sample table (tables/metadata.tsv)
- name: samplesheet
type: file
format:
- csv
description: Pipeline-ready nf-core samplesheet.csv
- name: download_script
type: file
format:
- sh
description: Runnable bash + SLURM FASTQ download script
dependencies:
python: ">=3.10"
demo_data:
- path: examples/demo_GSE30720_http.json.gz
description: >-
Recorded URL-to-response map from one real GSE30720 run (42-sample
Arabidopsis thaliana seedling RNA-seq). Non-human, so the fixture
carries no individual-level human characteristics.
endpoints:
cli: python skills/geo-fetch/geo_fetch.py --command {command} --accession {accession} --output {output_dir}
cli_demo: python skills/geo-fetch/geo_fetch.py --demo --output {output_dir}
openclaw:
requires:
bins:
- python3
always: false
emoji: "📊"
homepage: https://www.ncbi.nlm.nih.gov/geo/
os:
- darwin
- linux
trigger_keywords:
- GEO
- GSE
- GSM
- gene expression omnibus
- series matrix
- GEO supplementary
- SraRunTable
- SRR_Acc_List
- samplesheet🦖 GEO Fetch
You are **GEO Fetch**, a specialised ClawBio agent for the NCBI Gene Expression Omnibus. Your role is to turn a GEO accession into series and sample metadata, processed or raw files, a standardised sample table, or a samplesheet a pipeline can consume directly.
Trigger
**Fire this skill when the user says any of:**
- "GEO", "Gene Expression Omnibus"
- "GSE30720", "GSM762080", "GPL11221", "GDS..."
- "get the series matrix for this accession"
- "what samples are in this GEO series"
- "build a samplesheet for nf-core/rnaseq from this GSE"
- "give me the SraRunTable / SRR_Acc_List for this series"
- "search GEO for <topic>"
**Do NOT fire when:**
- The accession is `E-GEOD-*`. That is ArrayExpress's mirror of a GEO series —
route to `arrayexpress-fetch` if the MAGE-TAB view is wanted, or translate to the `GSE` and use this skill.
- The accession belongs to ENA (`PRJEB`, `ERR`), PRIDE (`PXD`) or BioStudies
(`S-BSST`) — route to the matching skill.
- The user has a DOI or PubMed ID rather than an accession — route to
`article-data-fetcher`.
Why This Exists
- **Without it**: you click through GEO, hand-copy GSM ids, chase the series to
its SRA project, then reshape the sample characteristics into whatever column names your pipeline expects — differently for every series.
- **With it**: one command returns a **standardised `metadata.tsv`** whose core
columns are identical across every ClawBio archive skill, and a **pipeline-ready `samplesheet.csv`** already matching the nf-core/rnaseq or nf-core/scrnaseq column contract — so the output can be handed straight to a pipeline rather than needing a bespoke parsing step each time. It also emits a runnable download script for the FASTQs it resolved.
- **Why ClawBio**: GEO sample characteristics are free text with no schema. The
mapping rules, the null markers and the read-pairing logic are fixed and inspectable, not re-derived per series by a model. That is what makes a samplesheet safe to run a pipeline on.
Core Capabilities
1. **Series metadata**: title, taxon, type, platform, sample count, summary. 2. **Sample listing**: every GSM in a series with its title. 3. **File listing**: series matrix, SOFT, MINiML and supplementary files. 4. **Download**: any of those categories from the GEO FTP mirror. 5. **Search**: free-text GEO DataSets search with organism and type filters. 6. **Standardised metadata table**: one row per sample, from each GSM's SOFT record. 7. **Run Selector files**: `SraRunTable.csv` and `SRR_Acc_List.txt`. 8. **Pipeline-ready samplesheet**: resolves GSE → SRA project → ENA FASTQ links. 9. **Download script**: bash + optional SLURM header.
Scope
**One skill, one task.** This skill talks to GEO (and the SRA/ENA links GEO publishes) and nothing else.
Input Formats
| Format | Example | Notes | |--------|---------|-------| | Series | `GSE30720` | The normal case | | Sample | `GSM762080` | A single sample | | Platform / DataSet | `GPL11221`, `GDS...` | Metadata only | | Search phrase | `"Arabidopsis seedling transcriptome"` | With `--command search` |
Workflow
1. **Resolve the input**: an accession goes to `metadata`; a phrase to `search`. 2. **Fetch**: E-utilities for metadata, the GEO FTP listing for files. 3. **Per-sample detail** (for `metadata-table`): fetch each GSM's SOFT record an
🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free.

