Skip to content
Machine Learning
Skill

/nemotron-retrieval-recipes

Use when planning, debugging, tuning, evaluating, exporting, or deploying public Nemotron `embed`/`rerank` retrieval recipes.

BOOST
From plugin
nemotron
2.1k10 skills
Install
$ npx -y skills add nvidia-nemo/nemotron --skill nemotron-retrieval-recipes --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/nemotron-retrieval-recipes

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when planning, debugging, tuning, evaluating, exporting, or deploying public Nemotron `embed`/`rerank` retrieval recipes.

SKILL.md

nemotron-retrieval-recipes.SKILL.md
name: nemotron-retrieval-recipes
version: "0.2.0"
author: "NVIDIA Nemotron Team <noreply@nvidia.com>"
license: Apache-2.0
tags:
  - nemotron
  - retrieval
  - fine-tuning
  - embeddings
  - reranking
metadata:
  author: "NVIDIA Nemotron Team <noreply@nvidia.com>"
  tags:
    - nemotron
    - retrieval
    - fine-tuning
    - embeddings
    - reranking
tools:
  - Read
  - Bash
  - Search
description: Use when planning, debugging, tuning, evaluating, exporting, or deploying public Nemotron `embed`/`rerank` retrieval recipes.

Nemotron Retrieval Recipes

Invocation: `$nemotron-retrieval-recipes`.

Purpose

Use this skill to work with public Nemotron embedding and reranking retrieval recipes in a source checkout or installed package. Prefer the current checkout over memory, because the recipe CLI, configs, containers, and output paths are actively changing. Treat each recipe family as available only after its recipe directory and matching CLI files are present.

This is a public product skill, not contributor-only guidance. Its value over static docs is to make an agent route the user's retrieval failure to the right recipe family, reconcile docs with the current checkout, avoid accidental long-running launches, preserve secrets, and return concrete preview/execution/run-report commands.

Use it only for tasks tied to the public Nemotron `embed` or `rerank` recipe flow. If the request is unrelated retrieval theory, generic vector database selection, generic benchmark advice, or non-recipe Docker/Slurm/NIM troubleshooting, stop with a short scope note and do not inspect recipe files in that turn.

Security Notes

Use `Bash` for repo-scoped inspection, help, dry-run, and user-approved execution commands. Do not run API, GPU, Docker, Slurm, NIM, or other long-running work unless the user explicitly asks for it. Before Stage 0 SDG for either family, confirm the user's data-governance policy permits sending corpus content to the configured inference endpoints; otherwise use an approved private or air-gapped path. Never run broad environment dumps or commands that expose secret values. Prefer dotlist overrides and config review over editing recipe defaults.

Source Priority

Resolve conflicts in this order:

1. Current checkout recipe, CLI, config, and source files. 2. Bundled references in this skill. 3. User-provided docs or saved snippets. 4. Memory.

For runnable commands, treat the current checkout as authoritative. If a required recipe directory, CLI command, config, or env profile is missing, report the blocker instead of guessing.

Prerequisites

  • Repo environment: `uv sync --all-extras` or the smallest relevant extra documented by the checkout.
  • Stage 0 SDG: `NVIDIA_API_KEY`; never ask users to paste secret values.
  • Stages 1–3 GPU work: CUDA/NVIDIA driver availability and enough VRAM.
  • Stage 4 export: NeMo Export-Deploy container when using TensorRT. The default Nemotron 3 Embed profile intentionally skips export.
  • Stage 5 deploy: Docker. Default Nemotron 3 Embed can use the checked-in vLLM path with `backend=vllm`, or a compatible `NEMOTRON3_EMBED_NIM_IMAGE` with `backend=nim`; Llama Embed and rerank deployment may require NGC access and `NGC_API_KEY`.
  • Remote execution: root `env.toml` profile for `--run` or `--batch`; load `references/remote.md` when remote scheduling, logs, or GPU placement matter.

Instructions

1. Identify the recipe family.

  • Use `references/embed.md` for embedding, embed, bi-encoder, vector search, first-stage retrieval, low Recall@k, missing relevant documents, NIM embeddings, or `nemotron embed`.
  • Use `references/rerank.md` for rerank, reranker, cross-encoder, second-stage retrieval, acceptable recall but poor top-rank ordering, low nDCG with good Recall, or `nemotron rerank`.
  • Use both references only when the user asks about both families or asks which family to choose.

2. For `embed`, choose one model profile before composing stage commands.

  • Run `uv run nemotron embed info` when the requested model is unclear.
  • Use `-c default` for `nvidia/Nemotron-3-Embed-1B-BF16`.
  • Use `-c llama` for `nvidia/llama-nemotron-embed-1b-v2` and its export path.
  • Carry the selected profile and `artifact_root` through every stage; never combine artifacts from the two profiles.

3. Choose the model family to tune from the retrieval failure mode.

  • Prefer embedding fine-tuning when relevant documents are absent from the candidate set.
  • Prefer reranker fine-tuning when relevant documents are retrieved but ordered poorly near the top.
  • For production retrieval stacks, remember that these are complementary: embed first, rerank candidates second.

4. Identify the intent: plan a run, execute a stage, debug a failure, tune hyperparameters, interpret metrics, export/deploy a model, inspect configs, or propose dotlist overrides. 5. Inspect the current public surface before acting:

  • Recipe files: `src/nemotron/recipes/<embed|rerank>/`
  • CLI files: `src/nemotron/cli/commands/<embed|rerank>/`
  • Configs: `src/nemotron/recipes/<family>/stage*/config/<profile>.yaml`
  • Help and dry runs: `uv run nemotron <family> --help`, `uv run nemotron <family> <stage> -c <profile> -d`

Safe Workflow

1. Gather only context relevant to the task: recipe family, selected profile, corpus path, existing SDG/training/eval data, target stage range, artifact root, checkpoint path, execution mode, GPU IDs, and whether required secrets are configured. Never ask users to paste secret values. 2. Start with cheap checks before expensive work:

  • `uv run nemotron <family> --help`
  • `uv run nemotron <family> <stage> --help`
  • `uv run nemotron <family> <stage> -c <profile> -d`
  • `uv run nemotron <family> run -c <profile> -d --from <stage> --to <stage>`
  • `run --help` may omit inherited `-c` and `-d` options even though `run -c default -d ...` works; validate by running the dry-run when unsure.
  • In an already prepared checkout
Read more
Ships withnemotron

Open and efficient models for agentic AI. Training recipes, deployment guides, and use-case examples for the Nemotron family.

Get the whole plugin
Stats
2,137
Stars
433
Forks
Active
Maintenance
Jupyter Notebook
Language
Apache-2.0
License
1d ago
Last commit
1y ago
Created
18h ago
Added

Repo: nvidia-nemo/nemotron

Other skills on nemotron.