Skip to content
Research
Skill

/scvi-tools

Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression. Reach for this skill to integrate scRNA-seq batches, embed cells for clustering, transfer annotations

BOOST
From plugin
open-science
5.5k25 skills
Install
$ npx -y skills add aipoch/open-science --skill scvi-tools --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/scvi-tools

Context preview

The summary Claude sees to decide when to auto-load this skill.

Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression. Reach for this skill to integrate scRNA-seq batches, embed cells for clustering, transfer annotations

SKILL.md

scvi-tools.SKILL.md
name: scvi-tools
description: >
  Probabilistic single-cell RNA-seq with scvi-tools — scVI for a
  batch-corrected latent space, scANVI for semi-supervised label transfer,
  and Bayesian differential expression. Reach for this skill to integrate
  scRNA-seq batches, embed cells for clustering, transfer annotations from a
  reference onto a query, or score differentially expressed genes per cluster.
  For spatial deconvolution / mapping use the cell2location, DestVI, or
  Tangram methods instead.
license: Apache-2.0
requirements: [gpu]
metadata:
  display-name: scvi-tools

scvi-tools — scVI / scANVI

scvi-tools (Gayoso et al. 2022, github.com/scverse/scvi-tools, BSD-3-Clause) wraps a family of deep generative models for single-cell omics. The scRNA-seq core is **scVI** (unsupervised batch-corrected latent embedding) and **scANVI** (scVI + a classifier head for semi-supervised cell-type label transfer). Both expect **raw integer UMI counts** and emit a low-dimensional `X_scVI` / `X_scANVI` that drops into the scanpy neighbors → leiden → umap pipeline.

Setup (any agent, no API key)

This is a **pure skill** — `kernel.py` is deterministic Python and _you_ (the base model) do all the reasoning. There is no `host` runtime and no LLM API. The helpers are `prepare_scvi_counts` for input validation and `h5ad_safe_obs` for obs/var frames that `anndata.write_h5ad()` can serialize. Load them once per session in a Python cell:

exec(open("scvi-tools/kernel.py", encoding="utf-8").read())   # path to this skill's kernel.py

Nothing auto-loads it outside Claude Science. Then call the helpers directly. If a helper raises `NameError`, you haven't exec'd kernel.py.

Dependencies: `pip install scvi-tools scanpy anndata`. Training needs a CUDA-capable GPU — see [Remote compute](#remote-compute-rent-a-gpu) to fall out to a rented GPU when you don't have one locally.

How to run

scVI — batch-corrected latent space

`prepare_scvi_counts(adata)` checks and preserves existing `counts`. If absent, it checks `.X` before copying it to `counts`. Values must be finite, nonnegative, and integer-valued, with at least one positive count. Invalid existing counts raise instead of being replaced by `.X`; sparse matrices stay sparse.

These numerical checks cannot prove raw-count provenance. Check the dataset documentation or preprocessing history; do not round or exponentiate transformed values to make them pass. If the history is unclear, resolve it before training. The helper does not use `.raw.X`, which may be normalized or have a different gene axis. Use an in-memory AnnData object; materialize backed data or copy views only within the available memory budget.

import scanpy as sc
import scvi

adata = sc.read_h5ad("dataset.h5ad")
# Verify count provenance from the input documentation or preprocessing history.
# Preserve counts if present; otherwise validate .X before copying it to counts.
counts_record = prepare_scvi_counts(adata)
print(counts_record)       # numerical checks only; does not certify provenance
adata.X = adata.layers["counts"].copy()        # derive plotting/HVG data from verified counts
sc.pp.normalize_total(adata); sc.pp.log1p(adata) # optional, for HVG / plotting only
sc.pp.highly_variable_genes(adata, n_top_genes=2000, batch_key="batch", subset=True)

scvi.model.SCVI.setup_anndata(adata, layer="counts", batch_key="batch")
model = scvi.model.SCVI(adata, n_latent=30)
model.train(max_epochs=200, early_stopping=True, accelerator="gpu", devices=1)

adata.obsm["X_scVI"] = model.get_latent_representation()
adata.layers["scvi_normalized"] = model.get_normalized_expression(library_size=1e4)

scANVI — label transfer from a partially-annotated reference

lvae = scvi.model.SCANVI.from_scvi_model(
    model, labels_key="cell_type", unlabeled_category="Unknown",
)
lvae.train(max_epochs=20, n_samples_per_label=100, accelerator="gpu", devices=1)

adata.obsm["X_scANVI"] = lvae.get_latent_representation()
adata.obs["pred_cell_type"] = lvae.predict()

`accelerator="gpu", devices=1` is the PyTorch-Lightning spelling; the legacy `use_gpu=` kwarg was **removed** in scvi-tools 1.x and now raises `TypeError`.

Differential expression

de = model.differential_expression(
    groupby="leiden", group1="3",   # group2=None → vs. all other cells
    mode="change", delta=0.25,
)
top = de.sort_values("proba_de", ascending=False).head(50)

For one-vs-rest leave `group2` out — `"rest"` is scanpy's `rank_genes_groups` convention, not scvi-tools'; here `group2` is a literal category name and `"rest"` would match zero cells.

scvi-tools ≥1.4 defaults to `mode="vanilla"`, whose result columns are exactly:

['proba_m1', 'proba_m2', 'bayes_factor', 'scale1', 'scale2', 'raw_mean1',
 'raw_mean2', 'non_zeros_proportion1', 'non_zeros_proportion2',
 'raw_normalized_mean1', 'raw_normalized_mean2', 'comparison', 'group1',
 'group2']

— no `lfc_*`, no `proba_de`, no `is_de_fdr_*`. **Pass `mode="change"`** to get `lfc_mean` / `lfc_median` / `proba_de` / `is_de_fdr_0.05`. Sort on `proba_de` (or on `bayes_factor` if you deliberately stayed in vanilla mode).

Output format

| Key | What | | --------------------------------- | ------------------------------------------------------ | | `adata.obsm["X_scVI"]` | `n_cells × n_latent` batch-corrected embedding | | `adata.obsm["X_scANVI"]` | label-aware embedding (better separates known classes) | | `adata.obs["pred_cell_type"]` | scANVI predicted label per cell | | `adata.layers["scvi_normalized"]` | decoded expression, library-size normalized | | DE dataframe | per-gene `lfc_*` / `proba_de` (with `mode="change"`) |

Remote compute (rent a GPU)

An A100-class GPU is recommended for >50k cells. Training is a plain Python script (`pipeline.py`) t

Read more
Ships withopen-science

The open-source AI research workbench for scientific research and agent workflows. Local-first, model-agnostic desktop app with extensible skills, MCP tools and connectors, Python/R execution and traceable artifacts for reproducible research on macOS, Windows and Linux.

Get the whole plugin
Stats
5,469
Stars
506
Forks
Active
Maintenance
TypeScript
Language
Apache-2.0
License
23m ago
Last commit
3mo ago
Created
13h ago
Added

Repo: aipoch/open-science

Other skills on open-science.