Skip to content
Research
Skill

/clinvar_database

Use when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls for human genomic variants.

BOOST
From plugin
science-skills
3.2k40 skills
Install
$ npx -y skills add google-deepmind/science-skills --skill clinvar_database --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/clinvar_database

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls for human genomic variants.

SKILL.md

clinvar_database.SKILL.md
name: clinvar-database
description: >
  Use when needing clinical significance, pathogenicity classifications (e.g.,
  Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard
  positive" benchmark controls for human genomic variants.

ClinVar Database

Prerequisites

1. **`uv`**: Read the `uv` skill and follow its Setup instructions to ensure `uv` is installed and on PATH. 2. **User Notification**: If .licenses/clinvar_database_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://www.ncbi.nlm.nih.gov/clinvar/, then (2) create the file recording the notification text and timestamp. 3. **`.env` file**: Make sure the `.env` file exists in your home directory. Create one if it does not exist. 4. **`NCBI_API_KEY`** (optional): Raises the NCBI rate limit from 3 to 10 requests/second. The skill works without it, but a key is recommended if the user plans many queries or encounters a 429 error. You can register for a key for free at https://www.ncbi.nlm.nih.gov/account/settings/. You **MUST** use the safe credentials protocol in the `credentials` skill to check for and request this key if this skill looks relevant to the user's request.

Overview

ClinVar is the primary consensus record for clinical classifications of human genomic variations. It provides the "clinical ground truth" for pathogenicity labels (Pathogenic, Likely Pathogenic, Benign, VUS) based on assertions from global laboratories.

When to Use

**Use when you need to:**

  • Find the current clinical significance and star rating (review status) for a

specific variant.

  • Fetch clinician notes, assertion criteria, or rationales for previous

clinical laboratory classifications.

  • Retrieve the preferred condition name and associated HPO terms for a

specific variant.

  • Find a list of variant controls (e.g., "Find all Pathogenic variants in the

HBB gene within 50bp of a signal").

  • Check for conflicting interpretations for a given variant and identify the

organizations submitting each classification.

**Do NOT use when you need to:**

  • Find specific allele frequencies in global populations (use **gnomAD**).
  • Describe the normal biological role of a protein and typical inheritance

patterns (use **OMIM**).

  • Predict mechanistic effects of novel mutations, like frameshifts or exon

skipping (use **AlphaGenome**).

  • Find recommended surveillance schedules for patients with a pathogenic

variant (use **GeneReviews**).

  • Generate or view 3D structural models of affected proteins (use **PDB /

AlphaFold**).

Quick Start

ClinVar queries are executed via a robust Python wrapper script to handle strict rate limiting and XML/JSON parsing.

Example: Search for BRCA1 variants

uv run scripts/clinvar_api.py search --query "BRCA1[gene]" --output results.json

Core Rules

  • **Retmax Constraint**: The search command defaults to `--retmax 200`. For

any "List all" or gene-wide request, you MUST explicitly set `--retmax` higher (e.g., 1000) to ensure data completeness.

  • **Use the Wrapper**: Prefer the wrapper script for standard queries. It

handles rate limiting, retries, and the complex XML parsing for you. If the script's parsed output does not contain the specific fields you need, you may modify the script or query the NCBI E-utilities API directly — but be aware that the raw XML schemas are complex and vary between record types.

  • If the rate limit is hit, the script will throw a clear error. You **MUST**

use the safe credentials protocol in the `credentials` skill to check for and request the `NCBI_API_KEY` to help the user add it to their `.env` file.

  • **Notification**: If this skill is used, ensure this is mentioned in the

output.

Utility Scripts

1. `count` — Count Matching Variants

**Purpose:** Check how many variants match a query without fetching IDs. Use to decide whether a full `search` is warranted.

*Arguments:*

  • `--query`: (Required) NCBI Entrez search query string.
  • `--output`: (Required) Output JSON file path.

*Example:* `uv run scripts/clinvar_api.py count \ --query "TP53[gene] AND \"uncertain significance\"[clinsig]" \ --output count.json` *Output:* `{"total_count": <int>}`

2. `search` — Search Variants

**Purpose:** Identify variants based on genomic location, gene symbols, or clinical attributes using NCBI Entrez search syntax. The search command **automatically paginates** through all matching results to ensure complete, deterministic retrieval.

# Fetch ALL matching variants (default behavior)
uv run scripts/clinvar_api.py search \
  --query "BRCA1[gene]" --output results.json

# Search by Chromosome and Position Range
uv run scripts/clinvar_api.py search \
  --query "11[chr] AND 5225000:5226000[chrpos]" --output results.json

# Combine terms using Entrez syntax
uv run scripts/clinvar_api.py search \
  --query "HBB[gene] AND pathogenic[clinsig]" --output results.json

# Cap results at 50
uv run scripts/clinvar_api.py search \
  --query "TP53[gene]" --retmax 50 --output results.json

*Arguments:*

  • `--query`: (Required) NCBI Entrez search query string.
  • `--retmax`: Maximum total number of variant IDs to return. **Default is 0,

which means "fetch all matching results."** Set to a positive integer to cap the result set.

  • `--page_size`: Number of IDs to fetch per API request (default: 500, max:

10000 per NCBI limits).

  • `--output`: (Required) Output JSON file path.

*Output:* A JSON object containing:

  • `total_count` — Total number of matching variants in ClinVar.
  • `fetched_count` — Number of IDs actually retrieved.
  • `variant_ids` — List of ClinVar Variation ID strings.

3. `summary` — Get Interpretation Summary

**Purpose:** Retrieve top-line clinical significance labels, star ratings (review

Read more
Ships withscience-skills

A collection of agent skills for scientific research tasks, spanning genomics, structural biology, cheminformatics, literature search, and more.

Get the whole plugin

Other skills on science-skills.