nvidia-skill-finder
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. Trigger on NVIDIA products, hardware, software,…
Finetune a pretrained KERMT encoder on a labeled CSV. Validate the checkpoint and data, prepare features, and run containerized training. Use a local checkpoint or optionally download a pinned Hugging Face model bundle using HF_TOKEN if configured. Write model bundles, prepared
$ npx -y skills add NVIDIA/skills --skill bionemo-kermt-finetune --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/bionemo-kermt-finetuneContext preview
The summary Claude sees to decide when to auto-load this skill.
Finetune a pretrained KERMT encoder on a labeled CSV. Validate the checkpoint and data, prepare features, and run containerized training. Use a local checkpoint or optionally download a pinned Hugging Face model bundle using HF_TOKEN if configured. Write model bundles, prepared
name: kermt-finetune description: Finetune a pretrained KERMT encoder on a labeled CSV. Validate the checkpoint and data, prepare features, and run containerized training. Use a local checkpoint or optionally download a pinned Hugging Face model bundle using HF_TOKEN if configured. Write model bundles, prepared data, logs, and trained models to user-selected host directories. license: Apache-2.0 compatibility: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron. metadata: owner: evax@nvidia.com classification: workflow-skill risk_tier: skill # Line/token budget: targets ~250 lines / ~3000 tokens — within the # 500-line / 5000-token cap for skill files.
Finetune a pretrained KERMT encoder on a user-supplied labeled CSV. The skill is the workflow orchestrator: validate ckpt, validate data, prepare data, launch the runner detached, return a run directory + container name.
Set `SKILL_DIR` to the absolute path of this installed skill directory. Export `KERMT_REPO` as the absolute path to the KERMT checkout used for model execution. The bundled container helper mounts that checkout at `/workspace` and this skill at `/skill` (read-only). Commands inside the container use `/skill/scripts/`; defaults are bundled in `config/`. See [Released models](references/released-models.md) for checkpoint bundle requirements.
The optional released-model branch reads `config/released_model.json` for the Hugging Face repository, pinned revision, and filenames. The bundled `scripts/fetch_released_model.py` downloads the model bundle over HTTPS into the host directory the user selects. Public models work without credentials; if `HF_TOKEN` is set, the container helper forwards it for Hugging Face authentication. Prepared data, logs, and workflow results go into the chosen run directory.
select one. For faster training on a multi-GPU host, pass `--num-gpus N` (N>1) to run data-parallel DDP across N GPUs — `--batch-size` is then per-GPU (effective global batch = batch_size × N).
works at smaller batch sizes — pass `--batch-size N` to override.
`kermt-setup` validates this up-front.
Required:
is a target.
Checkpoint (optional — defaults to the released model if omitted):
The validator refuses already-finetuned ckpts with a redirect to `kermt-infer`. **If omitted**, the skill offers to download the released pretrained hybrid model **nvidia/NV-KERMT-70M-v2** and finetune from it — see "Resolve & validate the checkpoint" (workflow step 3).
the interactive prompt (for non-interactive / agent runs). Mutually exclusive with `--ckpt`.
`$KERMT_REPO/models/NV-KERMT-70M-v2/`). An already-complete bundle there is reused, not re-downloaded.
Optional:
`regression` (from `defaults_finetune.json`). Drives loss, metric defaults, and head initialization. For classification tasks pass `--dataset-type classification`.
validator auto-detects numeric non-smiles columns and the skill confirms with the user before proceeding.
splits. Either pass both or pass neither (the skill auto-splits using the configured `--split-type`).
default `scaffold_balanced` from `defaults_finetune.json`.
from the train CSV. No `--val-csv` / `--test-csv` needed.
`--val-csv` + `--test-csv` (and, separately, per-fold index files — see `kermt/util/utils.split_data`). Use this when the dataset ships its own canonical split (e.g. `tests/data/Biogen_for_grover/scaffold/ balance/<endpoint>/{train,val,test}.csv`).
or any name `kermt.util.metrics.get_metric_func` accepts.
`--final-lr F` / `--warmup-epochs F` / `--weight-decay F` / `--dropout F` / `--bond-drop-rate F` / `--dist-coff F` / `--early-stop-epoch N` / `--seed N` — training-hyperparameter overrides. Anything not given is filled from `config/defaults_finetune.json`.
per-target FFN heads (default 0 = off; useful for heterogeneous multi-target finetunes). Both must be set together when N > 0.
each.
when `--num-gpus > 1`.
(single-process, unchanged). N>1 runs `main.py finetune` with `WORLD_SIZE=N` (one process per GPU); `--batch-size` is per-GPU.
`prepare_data.json` in `<dir>`. Useful when iterating on hyperparameters.
Let `$
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. Trigger on NVIDIA products, hardware, software,…
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and…
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras;…
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample…