Skip to content
Data
Skill

/geniml

Use Geniml for audited local genomic-interval workflows: validate BED and universe contracts, plan Region2Vec or scEmbed runs, inspect model/tokenizer compatibility, and assess consensus universes.

From plugin
k-dense-ai-scientific-agent-skills-2
45k165 skills
Install
$ npx -y skills add K-Dense-AI/scientific-agent-skills --skill geniml --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/geniml

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use Geniml for audited local genomic-interval workflows: validate BED and universe contracts, plan Region2Vec or scEmbed runs, inspect model/tokenizer compatibility, and assess consensus universes.

SKILL.md

geniml.SKILL.md
name: geniml
description: "Use Geniml for audited local genomic-interval workflows: validate BED and universe contracts, plan Region2Vec or scEmbed runs, inspect model/tokenizer compatibility, and assess consensus universes."
license: MIT
compatibility: Requires Python 3.10+ and uv. Guidance targets geniml 0.8.4 with gtars 0.9.2; ML workflows need the pinned ml extra and compatible native wheels. Bundled planners and inspectors are dependency-free, local-only, and make no network requests.
allowed-tools: Read Write Edit Bash Glob
metadata:
  version: "1.2"
  skill-author: "K-Dense Inc."
  upstream-version: "0.8.4"
  last-reviewed: "2026-07-23"

Geniml

Use Geniml for machine learning and statistical workflows over genomic interval sets. Treat coordinates, assemblies, token vocabularies, model artifacts, and sample grouping as explicit contracts. The bundled scripts validate or plan; they do not import Geniml, contact services, deserialize models, or execute training.

`Bash` is declared only for explicit, user-approved `uv`, Python, Geniml, Gtars, Git, and native CLI commands shown in this guide; bundled Python helpers do not spawn subprocesses. Example paths under `data/`, `refs/`, `work/`, and `models/` are user-provided project placeholders, not missing bundled files.

Verified release snapshot

  • Latest stable PyPI release on 2026-07-23: `geniml==0.8.4` (2026-01-14).
  • PyPI does not declare `Requires-Python`; its classifiers list Python

3.10-3.14. Prefer Python 3.11 or 3.12 where all native/ML wheels resolve.

  • `geniml==0.8.4` accepts `gtars>=0.2.5`; the verified base smoke used current

`gtars==0.9.2` (2026-06-17, Python >=3.10).

  • Extras are `ml` and `test`. The base install omits Torch, Gensim, Scanpy,

Hugging Face Hub, pyBigWig, and HMM dependencies.

  • Upstream documentation contains stale examples. Release source and installed

`--help` output take precedence where they conflict.

Install reproducibly

Use a project environment and commit its generated lockfile:

uv venv --python 3.12
uv pip install "geniml==0.8.4" "gtars==0.9.2"

For Region2Vec, scEmbed, evaluation, or universe methods needing ML libraries:

uv pip install "geniml[ml]==0.8.4" "gtars==0.9.2"

For a durable project, prefer:

uv add "geniml[ml]==0.8.4" "gtars==0.9.2"
uv lock

Do not install an unpinned Git branch. Record Python, OS/architecture, the resolved lockfile, and the PyPI artifact digest. Geniml itself is BSD-2-Clause; the `MIT` frontmatter value licenses this skill's content.

Start with the safety gate

Before importing Geniml or running an external binary:

1. Work only with explicit local regular files. Reject URLs, FIFOs, devices, and symlinks unless the user deliberately changes that policy. 2. Validate BED structure and the declared assembly against a trusted local chromosome-sizes file. 3. Bound file count, bytes, rows, workers, epochs, and output size. 4. Separate train/validation/test by patient, donor, biological replicate, or other independent unit—not by BED row or cell alone. 5. Inventory and checksum the universe, tokenizer, model, config, inputs, metadata manifest, and native binaries. 6. Obtain explicit approval before any BEDbase or Hugging Face download. Never infer approval from a model ID or BEDbase identifier. 7. Keep logs aggregate and bounded. BED filenames, sample IDs, phenotypes, labels, barcodes, and genomic intervals may be sensitive.

Coordinate and assembly contract

BED intervals are normally **0-based, half-open** `[start, end)`: start is included, end is excluded, and length is `end - start`. Do not mix them with 1-based closed coordinates from VCF/GFF or user-facing genome browsers.

For every corpus and artifact, record:

  • assembly and patch/accession where possible (for example GRCh38 versus

GRCh38.p14), plus the chromosome-sizes checksum;

  • contig naming convention (`chr1` versus `1`), alt/random/decoy policy, and

mitochondrial naming;

  • coordinate convention, sorting order, duplicate/overlap policy, and whether

BED strand is meaningful;

  • liftover tool, chain digest, source/target assemblies, unmapped fraction, and

post-liftover validation.

Reject negative coordinates, `end <= start`, integer overflow, unknown contigs, ends beyond contig length, malformed columns, mixed assemblies, and silent contig renaming. Sorting and normalization never repair an assembly mismatch. BED3 has no strand; when column 6 is present, preserve `+`, `-`, or `.` unless the assay contract says otherwise.

Run a bounded validation and normalization **plan** before analysis:

python skills/geniml/scripts/bed_validator.py \
  --input data/peaks.bed \
  --assembly GRCh38 \
  --chrom-sizes refs/GRCh38.chrom.sizes

The validator reports proposed actions but never rewrites the BED file.

Current API map

Region and tokenizer I/O

Prefer Gtars for new interval/tokenizer code:

from gtars.models import Region, RegionSet
from gtars.tokenizers import Tokenizer

regions = RegionSet("data/peaks.bed")
tokenizer = Tokenizer.from_bed("refs/universe.bed")
encoded = tokenizer(regions)
input_ids = encoded["input_ids"]

`RegionSet` and `Tokenizer` also accept remote inputs in some constructors; this skill permits local paths only unless network access is explicitly approved. `geniml.io.RegionSet(regions, backed=False)` remains available as a legacy Python implementation; backed sets are iterable but not indexable. `geniml.io.Region` uses `stop`, while `gtars.models.Region` uses `end`.

With gtars 0.9.2, seven special tokens are added to a BED vocabulary. Therefore `len(tokenizer)` is not simply the number of universe rows. Preserve universe row order and the exact special-token map.

Region2Vec

The modern class lives at a concrete module path:

from geniml.region2vec.main import Region2VecExModel
from geniml.region2vec.utils import Region2VecDataset
from gtars.tokenizers im
Read more
Ships withk-dense-ai-scientific-agent-skills-2

🔔 Claude Scientific Skills is now Scientific Agent Skills. Same skills, broader compatibility — now works with any AI agent that supports the open Agent Skills standard, not just Claude.

Get the whole plugin
Stats
44,851
Stars
4,066
Forks
Active
Maintenance
Python
Language
MIT
License
8h ago
Last commit
10mo ago
Created

Repo: K-Dense-AI/scientific-agent-skills