nvidia-skill-finder
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not…
Fine-tune public CodonFM Encodon checkpoints on labeled coding-sequence or coding-variant data using LoRA, head-only, or full fine-tuning. Use when a user explicitly asks to fine-tune CodonFM or Encodon for regression or classification. Support generic public-v1 Encodon
$ npx -y skills add NVIDIA/skills --skill bionemo-codonfm-finetune --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/bionemo-codonfm-finetuneContext preview
The summary Claude sees to decide when to auto-load this skill.
Fine-tune public CodonFM Encodon checkpoints on labeled coding-sequence or coding-variant data using LoRA, head-only, or full fine-tuning. Use when a user explicitly asks to fine-tune CodonFM or Encodon for regression or classification. Support generic public-v1 Encodon
name: codonfm-finetune description: Fine-tune public CodonFM Encodon checkpoints on labeled coding-sequence or coding-variant data using LoRA, head-only, or full fine-tuning. Use when a user explicitly asks to fine-tune CodonFM or Encodon for regression or classification. Support generic public-v1 Encodon workflows only; reject Decodon, MissenseDataset, missense_synom_agg, and generation workflows. metadata: author: "NVIDIA BioNeMo <bionemofeedback@nvidia.com>"
Use `--pretrained_ckpt_path` for public v1. Do not substitute `--checkpoint_path`: the public runner does not forward that argument to the fine-tuning task.
Resolve the target label, dataset, checkpoint, and output directory from the request and available files. Reuse existing data and weights. For training, check the project's ML dependencies and a compatible NVIDIA GPU before launch. If a required resource is unavailable, complete the available data preparation and return the command with the missing prerequisite clearly identified. When training is requested and the prerequisites are met, execute it and check the resulting checkpoints and metrics. A request for preparation ends with the validated inputs and command.
Use the user's labeled dataset when provided. For a demonstration of sequence regression without a dataset, use the public human RiboNN translation-efficiency data below and state that choice. This is not a substitute for a user's intended assay or for labeled coding variants. If variant labels are missing, return the required schema and a command template promptly; do not search for labels or invent measured effects.
Default demonstration checkpoint: `nvidia/NV-CodonFM-Encodon-80M-v1`, revision `399ca9fe17b57941a7bebc6788033919b417413c`, file `NV-CodonFM-Encodon-80M-v1.safetensors` with sibling `config.json`. The [public weights](https://huggingface.co/nvidia/NV-CodonFM-Encodon-80M-v1) are about 307 MB. Download them when needed for the requested work; input preparation can record an intended checkpoint path. These are the original Encodon weights; the `-TE-` checkpoints use the separate TransformerEngine implementation.
For an unsupported Decodon or missense-aggregation request, inspect the public parser/model configuration, explain the missing feature, and finish. Do not implement the missing model or search private repositories.
Prepare a small public-data example with the standard-library helper [prepare_ribonn.py](scripts/prepare_ribonn.py), running from the repository root. Set `CODONFM_DATA_PATH` to the CSV you want to create:
python skills/codonfm-finetune/scripts/prepare_ribonn.py \
--output "$CODONFM_DATA_PATH"With an existing raw file, add `--input "$RIBONN_DATA_PATH"`. The default reads at most eight accepted rows per split; `--max-rows-per-split 0` processes the full input. Remote streaming has a time budget and no automatic retries; use a local file if it fails. The helper follows the CDS slicing in the [RiboNN notebook](../../notebooks/4-EnCodon-Downstream-Task-riboNN.ipynb):
and `value = mean_te` unchanged. Do not take another logarithm.
`test`. This is a demonstration holdout, not the notebook's cross-validation.
silently truncating labeled examples. Record counts and source in the adjacent `.metadata.json`. A small subset does not establish predictive performance.
The pinned dataset URL is in the helper; its source is [CenikLab/TE_classic_ML](https://github.com/CenikLab/TE_classic_ML/tree/main/data). The notebook extracts frozen Encodon embeddings and trains a random-forest regressor with fold-based cross-validation. This skill reuses its data source, CDS extraction, and target for a separate fine-tuning example; it does not reproduce the notebook's training procedure or results.
Accept only `encodon_80m`, `encodon_600m`, or `encodon_1b`.
Require `id`, `ref_seq`, `value`, and `split` columns. Extra columns are allowed. Map the user's columns to this loader schema; the RiboNN helper is only for RiboNN source data. Training needs `train` rows and, when validation is enabled, `val` rows. A `test` split is needed only for later evaluation. Labels in unused splits need not be populated. Regression targets must be finite numbers; classification targets must be integer class indices from zero through `num_classes - 1`. Use a downstream head for scalar targets.
Check sequence preparation with the user’s assay in mind. The loader converts uppercase RNA `U` to `T`, and the tokenizer uppercases bases. Ambiguous bases and incomplete codons can produce unknown tokens; overlength sequences are truncated. Review these cases rather than silently dropping user records. Choose batches and a training budget appropriate to the dataset; small training sets may be resampled by the loader.
Set `CODONFM_CHECKPOINT_PATH` to the checkpoint file and `CODONFM_RUN_DIR` to your chosen output directory. This example runs ten steps to check the workflow; choose the training budget for the actual dataset and task:
python -m src.runner finetune \
--exp_name property_finetune \
--model_name encodon_80m \
--pretrained_ckpt_path "$CODONFM_CHECKPOINT_PATH" \
--data_path "$CODONFM_DATA_PATH" \
--process_item codon_sequence \
--dataset_name CodonBertDataset \
--finetune_strategy lora \
--lora_alpOfficial, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not…
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration,…
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment,…
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use…
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip)…
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for…