Skip to content
Development
Skill

/bionemo-codonfm-embed

Validate coding-sequence CSVs, extract public CodonFM/Encodon embeddings, and choose checkpoints for downstream property modeling.

BOOST
From plugin
nvidia-skills
3.6k200 skills
Install
$ npx -y skills add NVIDIA/skills --skill bionemo-codonfm-embed --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/bionemo-codonfm-embed

Context preview

The summary Claude sees to decide when to auto-load this skill.

Validate coding-sequence CSVs, extract public CodonFM/Encodon embeddings, and choose checkpoints for downstream property modeling.

SKILL.md

bionemo-codonfm-embed.SKILL.md
name: codonfm-embed
description: Validate coding-sequence CSVs, extract public CodonFM/Encodon embeddings, and choose checkpoints for downstream property modeling.
license: Apache-2.0
metadata:
  author: "NVIDIA BioNeMo <bionemofeedback@nvidia.com>"
  tags: [biology, codonfm, embeddings]

Extract public Encodon embeddings

Purpose

Extract one frozen CLS vector per coding sequence with public Encodon v1. Support input validation, command preparation, extraction, and checkpoint selection for translation efficiency, expression, or mRNA stability modeling. Extraction does not automatically train a downstream regressor.

Prerequisites

  • Validation needs Python 3 standard library only; no GPU, weights, or API key.
  • Execution needs the public CodonFM checkout, its `requirements.txt` environment,

a compatible NVIDIA GPU, and local checkpoint weights. A metadata JSON is not a checkpoint. A `.safetensors` file needs its sibling `config.json`; `.ckpt` checkpoints are also supported by the public loader.

  • Run `python -m src.runner` from the CodonFM repository root. In an isolated

workspace, use supplied source artifacts; source paths below are relative to that checkout or source archive, not this skill directory.

Inputs

Input source precedence: explicit user prompt arguments, then supplied files/checkpoint metadata, then inspected public runner defaults. Resolve conflicting model names and checkpoint metadata before execution. Supplied 80M metadata is useful for preparing an 80M command; it does not restrict an open-ended recommendation to that size.

Required for validation: a CSV. Required for extraction: the CSV, checkpoint, matching model name, and output directory. Optional: context length and batch size overrides. Checkpoint-selection questions can be answered without a CSV.

| Input | Requirement or default | | --- | --- | | Sequence CSV | Columns `id`, `ref_seq`, `value`, `split`; extra columns allowed | | `id` | Nonblank, unique IDs for unambiguous output association | | `ref_seq` | Coding sequence, uppercase DNA `A/C/G/T`, length divisible by three; public dataset converts uppercase `U` to `T` | | `value` | Numeric label; use `0.0` for new extraction-only data, preserve supplied labels | | `split` | Only exact `test` values enter extraction; blank/other values are excluded | | Checkpoint and model | Match weights/config to `encodon_80m`, `encodon_600m`, or `encodon_1b` | | Context length | Public runner default `2048` tokens, including CLS and SEP | | Output directory | A fresh run directory with an empty predictions directory |

Instructions

1. **Choose the requested workflow.** For a checkpoint/performance question, read [checkpoint selection](references/checkpoint-selection.md) and answer from public benchmark evidence. For the strongest published downstream results, prefer the public **1B random-mask checkpoint** when resources allow; 80M is a demonstration or resource-constrained choice. A small labeled set alone does not establish that 80M frozen features are better. Do not download weights or inspect the entire source tree just to make a recommendation. 2. **Inspect supplied source only where needed.** Confirm runner/config, `src/data/codon_bert_dataset.py`, `src/data/preprocess/codon_sequence.py`, `src/inference/encodon.py`, or `src/utils/pred_writer.py` for the relevant behavior. Read ZIP members with `zipfile.ZipFile.namelist()` and `.read()`; source inspection does not need extraction. If a checkout is needed, use a new directory from `tempfile.mkdtemp()` or `mktemp -d`, without deleting or overwriting an existing directory. For Decodon support questions, inspect runner/config and model/inference modules, cite the inspected files, explain the missing public implementation, and finish there. 3. **Validate the CSV before running extraction.** Run the bundled checker below with the intended context length. Report per-row verdicts using CSV row numbers as well as IDs, since IDs can repeat. Separate excluded rows, invalid inputs, duplicate-ID warnings, and truncation. Propose fixes without silently rewriting supplied data. The checker is a preflight, not model execution or proof of biological CDS validity. 4. **Deliver the requested preparation or execution.** For preparation, return a complete command with resolved paths (or clearly identified prerequisites), the test-row count, validation findings, and the output contract below. Include all task/dataset/process flags in the final answer, even if already shown in a tool call. For extraction, reuse/download the chosen checkpoint when needed, execute once resources are ready, and verify the saved arrays. If resources are missing, finish preparation and state what is missing.

Available Scripts

| Script | Purpose | Arguments | | --- | --- | --- | | [validate_inputs.py](scripts/validate_inputs.py) | Read-only CSV validation and per-row verdicts | Required CSV path; optional `--context-length` (default `2048`) |

Run the preflight with Python; `CODONFM_SKILL_DIR` is the directory containing this file:

python "$CODONFM_SKILL_DIR/scripts/validate_inputs.py" "$CODONFM_DATA_PATH" \
    --context-length 2048

The checker prints JSON. Exit `0` means no findings, `1` means row findings to review (including exclusions/warnings), and `2` means a file/schema error. Neither warnings nor exclusions imply that the public runner will crash.

Output Format

The checker emits JSON with `total_rows`, `test_rows`, `excluded_rows`, `context_length`, `codon_limit`, `warnings`, and `rows`. Each row records its one-based data-row number (excluding the header), ID, split, verdict, issues, sequence/value validity, and retained/lost codons. A file/schema error emits `error` and `csv` instead. These are preflight findings, not generated embeddings.

Examples

Set `CODONFM_DATA_PATH` to the CSV, `CODONFM_CHECKPOINT_PATH` to the w

Read more
Ships withnvidia-skills

Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.

Get the whole plugin

Other skills on nvidia-skills.