nvidia-skill-finder
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not…
Validate coding-sequence CSVs, extract public CodonFM/Encodon embeddings, and choose checkpoints for downstream property modeling.
$ npx -y skills add NVIDIA/skills --skill bionemo-codonfm-embed --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/bionemo-codonfm-embedContext preview
The summary Claude sees to decide when to auto-load this skill.
Validate coding-sequence CSVs, extract public CodonFM/Encodon embeddings, and choose checkpoints for downstream property modeling.
name: codonfm-embed description: Validate coding-sequence CSVs, extract public CodonFM/Encodon embeddings, and choose checkpoints for downstream property modeling. license: Apache-2.0 metadata: author: "NVIDIA BioNeMo <bionemofeedback@nvidia.com>" tags: [biology, codonfm, embeddings]
Extract one frozen CLS vector per coding sequence with public Encodon v1. Support input validation, command preparation, extraction, and checkpoint selection for translation efficiency, expression, or mRNA stability modeling. Extraction does not automatically train a downstream regressor.
a compatible NVIDIA GPU, and local checkpoint weights. A metadata JSON is not a checkpoint. A `.safetensors` file needs its sibling `config.json`; `.ckpt` checkpoints are also supported by the public loader.
workspace, use supplied source artifacts; source paths below are relative to that checkout or source archive, not this skill directory.
Input source precedence: explicit user prompt arguments, then supplied files/checkpoint metadata, then inspected public runner defaults. Resolve conflicting model names and checkpoint metadata before execution. Supplied 80M metadata is useful for preparing an 80M command; it does not restrict an open-ended recommendation to that size.
Required for validation: a CSV. Required for extraction: the CSV, checkpoint, matching model name, and output directory. Optional: context length and batch size overrides. Checkpoint-selection questions can be answered without a CSV.
| Input | Requirement or default | | --- | --- | | Sequence CSV | Columns `id`, `ref_seq`, `value`, `split`; extra columns allowed | | `id` | Nonblank, unique IDs for unambiguous output association | | `ref_seq` | Coding sequence, uppercase DNA `A/C/G/T`, length divisible by three; public dataset converts uppercase `U` to `T` | | `value` | Numeric label; use `0.0` for new extraction-only data, preserve supplied labels | | `split` | Only exact `test` values enter extraction; blank/other values are excluded | | Checkpoint and model | Match weights/config to `encodon_80m`, `encodon_600m`, or `encodon_1b` | | Context length | Public runner default `2048` tokens, including CLS and SEP | | Output directory | A fresh run directory with an empty predictions directory |
1. **Choose the requested workflow.** For a checkpoint/performance question, read [checkpoint selection](references/checkpoint-selection.md) and answer from public benchmark evidence. For the strongest published downstream results, prefer the public **1B random-mask checkpoint** when resources allow; 80M is a demonstration or resource-constrained choice. A small labeled set alone does not establish that 80M frozen features are better. Do not download weights or inspect the entire source tree just to make a recommendation. 2. **Inspect supplied source only where needed.** Confirm runner/config, `src/data/codon_bert_dataset.py`, `src/data/preprocess/codon_sequence.py`, `src/inference/encodon.py`, or `src/utils/pred_writer.py` for the relevant behavior. Read ZIP members with `zipfile.ZipFile.namelist()` and `.read()`; source inspection does not need extraction. If a checkout is needed, use a new directory from `tempfile.mkdtemp()` or `mktemp -d`, without deleting or overwriting an existing directory. For Decodon support questions, inspect runner/config and model/inference modules, cite the inspected files, explain the missing public implementation, and finish there. 3. **Validate the CSV before running extraction.** Run the bundled checker below with the intended context length. Report per-row verdicts using CSV row numbers as well as IDs, since IDs can repeat. Separate excluded rows, invalid inputs, duplicate-ID warnings, and truncation. Propose fixes without silently rewriting supplied data. The checker is a preflight, not model execution or proof of biological CDS validity. 4. **Deliver the requested preparation or execution.** For preparation, return a complete command with resolved paths (or clearly identified prerequisites), the test-row count, validation findings, and the output contract below. Include all task/dataset/process flags in the final answer, even if already shown in a tool call. For extraction, reuse/download the chosen checkpoint when needed, execute once resources are ready, and verify the saved arrays. If resources are missing, finish preparation and state what is missing.
| Script | Purpose | Arguments | | --- | --- | --- | | [validate_inputs.py](scripts/validate_inputs.py) | Read-only CSV validation and per-row verdicts | Required CSV path; optional `--context-length` (default `2048`) |
Run the preflight with Python; `CODONFM_SKILL_DIR` is the directory containing this file:
python "$CODONFM_SKILL_DIR/scripts/validate_inputs.py" "$CODONFM_DATA_PATH" \
--context-length 2048The checker prints JSON. Exit `0` means no findings, `1` means row findings to review (including exclusions/warnings), and `2` means a file/schema error. Neither warnings nor exclusions imply that the public runner will crash.
The checker emits JSON with `total_rows`, `test_rows`, `excluded_rows`, `context_length`, `codon_limit`, `warnings`, and `rows`. Each row records its one-based data-row number (excluding the header), ID, split, verdict, issues, sequence/value validity, and retained/lost codons. A file/schema error emits `error` and `csv` instead. These are preflight findings, not generated embeddings.
Set `CODONFM_DATA_PATH` to the CSV, `CODONFM_CHECKPOINT_PATH` to the w
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not…
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration,…
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment,…
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use…
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip)…
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for…