Skip to content
Research
Skill

/ncbi_sequence_fetch

Retrieve protein and nucleotide sequences from NCBI databases using E-utilities. Supports direct accession lookup, CDS translation, gene+organism search, locus lookup, PubMed-linked sequences, patent protein extraction, and organism+length fallback search. Use when you need to

BOOST
From plugin
science-skills
3.2k40 skills
Install
$ npx -y skills add google-deepmind/science-skills --skill ncbi_sequence_fetch --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/ncbi_sequence_fetch

Context preview

The summary Claude sees to decide when to auto-load this skill.

Retrieve protein and nucleotide sequences from NCBI databases using E-utilities. Supports direct accession lookup, CDS translation, gene+organism search, locus lookup, PubMed-linked sequences, patent protein extraction, and organism+length fallback search. Use when you need to

SKILL.md

ncbi_sequence_fetch.SKILL.md
name: ncbi-sequence-fetch
description: >
  Retrieve protein and nucleotide sequences from NCBI databases using
  E-utilities. Supports direct accession lookup, CDS translation, gene+organism
  search, locus lookup, PubMed-linked sequences, patent protein extraction, and
  organism+length fallback search. Use when you need to fetch biological
  sequences by accession, gene name, locus tag, PubMed ID, or patent number.

NCBI Sequence Fetch

Prerequisites

1. **`uv`**: Read the `uv` skill and follow its Setup instructions to ensure `uv` is installed and on PATH. 2. **User Notification**: If .licenses/ncbi_sequence_fetch_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://www.ncbi.nlm.nih.gov/ and https://www.ncbi.nlm.nih.gov/home/about/policies/, then (2) create the file recording the notification text and timestamp. 3. **`.env` file**: Make sure the `.env` file exists in your home directory. Create one if it does not exist. 4. **`NCBI_API_KEY`** (optional): Raises the NCBI rate limit from 3 to 10 requests/second. The skill works without it, but a key is recommended if the user plans many queries or encounters a 429 error. You can register for a key for free at https://www.ncbi.nlm.nih.gov/account/settings/. You **MUST** use the safe credentials protocol in the `credentials` skill to check for and request this key if this skill looks relevant to the user's request.

Core Rules

  • **Use the Wrapper**: ALWAYS execute the provided helper scripts to query the

database rather than accessing the database directly. The scripts automatically enforce the required rate limit gracefully.

  • **API Key Support**: If the user provides an `NCBI_API_KEY` in their

environment, the query speed limits are automatically increased significantly.

  • **Notification**: If this skill is used, ensure this is mentioned in the

output.

Overview

Wraps NCBI's Entrez E-utilities (efetch, esearch, elink, esummary) for retrieving protein and nucleotide sequences. Provides 10 subcommands covering the full range of sequence retrieval workflows:

  • `fetch-protein` — Direct protein accession lookup (GenPept, RefSeq)
  • `fetch-nucleotide` — Direct nucleotide accession lookup
  • `cds-translate` — Fetch CDS and translate to protein (3 methods)
  • `search` — Free-text search of any NCBI database
  • `elink` — Follow cross-database links (PubMed→Protein, etc.)
  • `gene-protein` — Search protein by gene name + organism
  • `locus-protein` — Search protein by locus tag + organism
  • `pubmed-proteins` — Find proteins linked to a PubMed article
  • `patent-search` — Extract protein sequences from patents
  • `organism-length` — Last-resort search by organism + exact AA length

Utility Scripts

**`scripts/ncbi_fetch.py`** — Single script with subcommands.

All subcommands write structured JSON output. Use `--output FILE` to save to a file, or omit it to print to stdout. A human-readable summary is always printed to stdout.

1. Fetch Protein by Accession

Fetches protein FASTA from NCBI by accession (XP_, NP_, GenPept, etc.)

uv run scripts/ncbi_fetch.py fetch-protein XP_022033624 -o /tmp/result.json
uv run scripts/ncbi_fetch.py fetch-protein NP_001234567 ABC12345.1

2. Fetch Nucleotide by Accession

Fetches nucleotide FASTA from NCBI by accession.

uv run scripts/ncbi_fetch.py fetch-nucleotide MK034466 -o /tmp/result.json

3. CDS Translate

Fetches a CDS/nucleotide accession and translates to protein sequence. Tries three approaches in order: 1. NCBI's pre-translated CDS protein (`fasta_cds_aa`)

2. GenBank XML CDS annotation translations 3. Raw nucleotide → 6-frame ORF finding

uv run scripts/ncbi_fetch.py cds-translate MK034466 -o /tmp/result.json
uv run scripts/ncbi_fetch.py cds-translate HQ662330 --target-length 1043

If the accession is a **genomic record** (not mRNA/CDS), the tool will report `is_genomic: true` so you can fall back to a homology-based approach instead.

4. Search Any Database

Free-text search using Entrez query syntax. Supports all NCBI databases.

# Search protein database
uv run scripts/ncbi_fetch.py search "WRR4B[Gene Name] AND Arabidopsis[Organism]" \
  --database protein --retmax 5 --fetch-sequences

# Search nucleotide database
uv run scripts/ncbi_fetch.py search "Rz2[Gene Name] AND Beta vulgaris[Organism]" \
  --database nuccore --retmax 10

# Search with patent filter
uv run scripts/ncbi_fetch.py search "disease resistance AND Solanum[Organism] AND patent[Properties]" \
  --database protein --fetch-sequences

# Search by sequence length
uv run scripts/ncbi_fetch.py search '"Oryza sativa"[Organism] AND 1043[SLEN]' \
  --database protein --fetch-sequences --retmax 50

5. Cross-Database Links (elink)

Follow NCBI's cross-database links (e.g., PubMed article → linked proteins).

uv run scripts/ncbi_fetch.py elink 24896089 --dbfrom pubmed --db protein \
  --fetch-sequences -o /tmp/linked.json

6. Gene + Organism Search

Searches for protein sequences by gene name and organism. Searches NCBI Protein with `[Gene Name]` and `[Organism]` qualifiers.

uv run scripts/ncbi_fetch.py gene-protein WRR4B --organism "Arabidopsis thaliana"
uv run scripts/ncbi_fetch.py gene-protein Pikh-2 --organism "Oryza sativa" \
  --target-length 1043 -o /tmp/result.json

7. Locus Tag Search

Searches by locus tag in both NCBI Protein and Nuccore databases. Extracts CDS translations from GenBank XML when direct protein hits aren't available.

uv run scripts/ncbi_fetch.py locus-protein At1g56540 --organism "Arabidopsis thaliana"
uv run scripts/ncbi_fetch.py locus-protein Niben101Scf02422g02015.1 \
  --organism "Nicotiana benthamiana" -o /tmp/result.json

8. PubMed-Linked Proteins

Finds protein sequences linked to a PubMed article. Searches NC

Read more
Ships withscience-skills

A collection of agent skills for scientific research tasks, spanning genomics, structural biology, cheminformatics, literature search, and more.

Get the whole plugin

Other skills on science-skills.