Skip to content
Research
Skill

/uniprot_database

Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef. Use when searching for proteins, mapping identifiers, or retrieving functional annotations and publications. Don't use for sequence alignment, protein folding, or sequence

BOOST
From plugin
science-skills
3.2k40 skills
Install
$ npx -y skills add google-deepmind/science-skills --skill uniprot_database --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/uniprot_database

Context preview

The summary Claude sees to decide when to auto-load this skill.

Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef. Use when searching for proteins, mapping identifiers, or retrieving functional annotations and publications. Don't use for sequence alignment, protein folding, or sequence

SKILL.md

uniprot_database.SKILL.md
name: uniprot-database
description: >-
  Access protein metadata, function, taxonomy, and sequences across UniProtKB,
  UniParc, and UniRef. Use when searching for proteins, mapping identifiers, or
  retrieving functional annotations and publications. Don't use for sequence
  alignment, protein folding, or sequence similarity search (use specialized
  skills for those tasks).

UniProt Database Access

Prerequisites

1. **`uv`**: Read the `uv` skill and follow its Setup instructions to ensure `uv` is installed and on PATH. 2. **User Notification**: If .licenses/uniprot_database_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://www.uniprot.org/help/license and https://www.uniprot.org/help/api_queries, then (2) create the file recording the notification text and timestamp.

Overview

Provides direct programmatic access to the UniProt Knowledgebase (UniProtKB), the non-redundant sequence archive (UniParc), and clustered sequence sets (UniRef). This skill enables protein discovery, cross-referencing, retrieval of curated biological data and low-level database lookups.

Core Rules

  • **Use the Wrapper**: Always use the provided Python scripts (e.g.,

`scripts/uniprot_tools.py`) rather than constructing custom curl requests.

  • **No Hallucinations**: Do NOT invent protein functions, metadata, or

sequences. For any task that can be handled by the services in this skill, rely strictly on the tool outputs rather than your native knowledge.

  • **Notification**: If this skill is used, ensure this is mentioned in the

output.

Use Cases

  • **Searching for Protein Function**: Querying functional annotations, GO

terms, subcellular locations etc.

  • **Searching for Protein Sequence**: Searching for protein sequences by their

functional annotations, genes etc. in UniProtKB, UniParc, and UniRef.

  • **Understanding Protein/Organism Relationships**: Leveraging the Taxonomy

database and Proteome sets.

  • **Large-Scale Metadata Retrieval**: Fetching annotations for thousands of

proteins via streaming.

  • **Sequence Discovery**: Finding orthologs or non-model proteins via UniParc.
  • **ID Mapping**: Converting IDs between UniProt and 100+ external databases.
  • **Historical Data (UniSave)**: Retrieving previous versions of entries or

tracking deleted sequences.

Available Tools

Choose the right tool based on the task type and data volume:

  • **`get`**: Retrieves metadata and sequence for a specific entry. Best for a

**single, known accession**.

  • Also accesses UniSave historical data (use `--dataset unisave`), which

is essential for reconciling data from older releases or identifying why a formerly valid accession no longer appears in search results.

  • **`search`**: Searches for entries matching a query. Best for **exploration

and discovery**.

  • Use with `--limit 5` to verify if a query returns the expected proteins

before committing to a larger download.

  • Automatically paginates if results exceed 500 entries to provide a

stable download.

  • *Warning*: For paginated search, TXT and other formats are not reliable

with `--limit` as it applies to lines, not entries.

  • See

[Search Query Fields Documentation](references/search_query_fields.md).

  • **`stream`**: Streams all matching entries. Best for **bulk retrieval** of

large datasets (up to 10,000,000 entries).

  • Does NOT support `--limit`; always returns the full result set.
  • Use `search` with `--limit` if you need a subset.
  • **`count`**: Counts entries matching a query. Best for answering direct

count questions or for **initial estimation** before running a full `search` or `stream`.

  • **`sparql`**: Executes graph queries for complex discovery. Best for

counting, exact sequence matches, and multi-database queries.

  • See [SPARQL Examples](references/sparql_examples.md).
  • **`map`**: Converts IDs between UniProt and 100+ databases. Best for ID

mapping tasks.

  • See [ID Mapping Documentation](references/id_mapping_documentation.md).
  • **`search` vs. `map`**: Try `search` first before resorting to `map` if

not explicitly requested by the user. E.g., an external ID might be searchable in UniParc but fail to map to UniProtKB.

Workflows

Typical Protein Research Workflow

Copy this checklist and track progress:

  • [ ] Step 1: Identify target protein(s) and organism(s).
  • [ ] Step 2: Search UniProtKB for reviewed entries (`reviewed:true`).
  • [ ] Step 3: If no reviewed entries, search unreviewed or use UniParc for

sequence discovery.

  • [ ] Step 4: Map external IDs (e.g., Ensembl, PDB) to UniProt Accessions if

necessary.

  • [ ] Step 5: Retrieve functional metadata or sequence in desired format

(JSON, FASTA).

Handling Search Misses (e.g. Gene Search in Non-Model Organisms)

If a direct query (e.g., `gene:SYMBOL`) fails:

1. **Pivot to Protein Name**: Search for the common protein name (e.g., `protein_name:Alpha-crystallin A`). 2. **Use UniParc**: Search the UniParc dataset, which integrates sequences from across all of life, even if they aren't fully annotated in UniProtKB. 3. **Check Orthologs/Canonical**: Resolve the Human/Mouse ortholog first to find the correct naming/mnemonic.

Bulk Retrieval Priorities

> [!IMPORTANT] Always prefer **`stream`** or **`sparql`** for bulk data. > `search` is suitable for exploration; if results exceed 500 entries, it > automatically paginates to provide a stable download.

  • **Priority 0: `count`**: ALWAYS check the result count before running a

`search` or `stream`.

  • **Priority 1: `stream`**: The primary method for bulk data retrieval (up to

10M entries). Does NOT support `--limit`; always returns all results.

  • **Priority 2: `sparq
Read more
Ships withscience-skills

A collection of agent skills for scientific research tasks, spanning genomics, structural biology, cheminformatics, literature search, and more.

Get the whole plugin

Other skills on science-skills.