Skip to content
Development
Skill

/bionemo-msa-search-nim

Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. Use for homolog search, UniRef30/ColabFold env searches, A3M or FASTA alignments, paired MSA search for complexes, PDB70 structural templates, hosted NVIDIA API calls, or local

BOOST
From plugin
nvidia-skills
3.5k200 skills
Install
$ npx -y skills add NVIDIA/skills --skill bionemo-msa-search-nim --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/bionemo-msa-search-nim

Context preview

The summary Claude sees to decide when to auto-load this skill.

Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. Use for homolog search, UniRef30/ColabFold env searches, A3M or FASTA alignments, paired MSA search for complexes, PDB70 structural templates, hosted NVIDIA API calls, or local

SKILL.md

bionemo-msa-search-nim.SKILL.md
name: msa-search-nim
description: >
  Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. Use for homolog search, UniRef30/ColabFold env searches, A3M or FASTA alignments, paired MSA search for complexes, PDB70 structural templates, hosted NVIDIA API calls, or local Docker deployment. For local deployment, download the databases in parallel with aria2c and launch via NIM_MODEL_NAME (the recommended default fast path, ~14 min vs over 80 min for the built-in downloader); a plain docker run uses the slow built-in downloader.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
permissions:
  - env      # reads NGC_API_KEY/NVIDIA_API_KEY and local NIM setup variables
  - network  # hosted MSA requests and documented NGC/local NIM setup

MSA-Search NIM

Generate protein MSAs with GPU-accelerated MMSeqs2. Use this guide for first-pass hosted/local usage; load supplemental files only when needed:

  • `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
  • `references/science.md`: MSA purpose, pairing/templates, limits, handoffs.
  • `references/parameters.md`: database, pairing, depth, and template tuning.
  • `references/validation.md`: alignment, template, and artifact checks.
  • `references/examples.md`: compact hosted/local request patterns.

Choose Mode And Endpoint

Ask only when context is unclear:

> Hosted NVIDIA API or local Docker NIM?

  • Hosted standard MSA: `https://health.api.nvidia.com/v1/biology/colabfold/msa-search/predict`
  • Hosted paired MSA: `https://health.api.nvidia.com/v1/biology/colabfold/msa-search/paired/predict`
  • Local standard MSA: `http://localhost:8000/biology/colabfold/msa-search/predict`
  • Local paired MSA: `http://localhost:8000/biology/colabfold/msa-search/paired/predict`
  • Local templates: `http://localhost:8000/biology/colabfold/msa-search/structure-templates/predict`

Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for registry login, entitlement checks, and first-run model downloads; pass it into the container with `-e NGC_API_KEY`. Local inference requests use no auth header after readiness. Warm-cache key-free startup varies by image/version and should not be assumed. The hosted template path returned HTTP 404 in validation, so use local Docker for template search unless the hosted docs/service changes.

Local Docker

**Default local deployment = parallel download + `NIM_MODEL_NAME`.** The first recipe below is the one to use for real workflows. It downloads the database(s) with a range-parallel downloader (aria2c) and starts the NIM against those files — ~14 min for UniRef30 vs >80 min for the NIM's built-in downloader (measured, H100). Do **not** reach for the plain `docker run` (the "Fallback" subsection) unless you only want a `databases:pdb70` smoke test or you deliberately want the NIM to manage its own blob cache.

Local setup requires a GPU. Size the NVMe volume to the profile you pick (UniRef30 ~490 GB; full set ~1.4 TB). For setup answers, include env preflight, `docker login`, the parallel download, `NIM_MODEL_NAME` launch, readiness, and then no-auth local inference. Do not invent a cache default or drop the `NVIDIA_API_KEY` fallback.

# --- env preflight (do not drop the NVIDIA_API_KEY fallback) ---
set -a
[ -f .env ] && . ./.env
set +a
if [ -z "${NGC_API_KEY:-}" ] && [ -n "${NVIDIA_API_KEY:-}" ]; then
  export NGC_API_KEY="$NVIDIA_API_KEY"
fi
: "${NGC_API_KEY:?Set NGC_API_KEY or NVIDIA_API_KEY}"
: "${DB_DIR:=/data/fast-db}"          # where the parallel download lands

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

# --- 1) pick the DB version(s) you need (paired/complex work = uniref30 only) ---
DB_VERSION=uniref30_2302-m18v1
command -v aria2c >/dev/null || { echo "aria2c required; install it (e.g. apt-get install -y aria2) and re-run"; exit 1; }
mkdir -p "$DB_DIR"

# --- 2) parallel download from NGC (see "Parallel Download" section for the all-DB loop) ---
curl -fsS -H "Authorization: Bearer $NGC_API_KEY" \
  "https://api.ngc.nvidia.com/v2/org/nim/team/colabfold/models/msa-search/${DB_VERSION}/files" \
  -o /tmp/files.json
DB_DIR="$DB_DIR" python3 - <<'PY'
import json, os
d = json.load(open("/tmp/files.json")); dbdir = os.environ["DB_DIR"]; lines = []
for url, path in zip(d["urls"], d["filepath"]):
    lines += [url.strip(), f"  dir={dbdir}", f"  out={path}"]
open("/tmp/aria.in", "w").write("\n".join(lines) + "\n")
PY
aria2c -i /tmp/aria.in --max-concurrent-downloads=4 --max-connection-per-server=16 \
  --split=16 --min-split-size=1M --continue=true --file-allocation=none

# --- 3) launch the NIM against the downloaded files (skips the slow built-in download) ---
docker run -d --name msa-search --runtime=nvidia --gpus all \
  -e NGC_API_KEY \
  -e NIM_MODEL_NAME=/databases \
  -v "${DB_DIR}:/databases" \
  -p 8000:8000 \
  nvcr.io/nim/colabfold/msa-search:2

Readiness:

until curl -sf http://localhost:8000/v1/health/ready; do sleep 5; done

If the DB is already present in `$DB_DIR`, skip steps 1-2 — the launch alone is a ~20 s warm start. See "Parallel Download For Any Database Set" for the multi-database (`databases:all`) loop and the full rationale.

Fallback: Let The NIM Download Its Own Databases (slower)

Use this only for a quick `databases:pdb70` smoke test, or when you specifically want the NIM to manage its own blob cache. It uses the built-in downloader, which is slow on large profiles (UniRef30 stalled past 80 min in testing). Pin the smallest profile with `NIM_MODEL_PROFILE` (see "Faster Startup") so it does not fetch the full 1.4 TB.

: "${LOCAL_NIM_CACHE:?Set LOCAL_NIM_CACHE}"
mkdir -p "${LOCAL_NIM_CACHE}"; chmod 755 "${LOCAL_NIM_CACHE}"
docker run --rm --name ms
Read more
Ships withnvidia-skills

Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.

Get the whole plugin

Other skills on nvidia-skills.