alphafold2
Predict protein structure for monomers and multimers with AlphaFold2 via the ColabFold runner…
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. 2026, github.com/Biohub/esm). Single-sequence and MSA modes; protein, DNA, RNA, ligand (CCD/SMILES), modified residues. FoldBench Ab-Ag 50-55%, PPI 70-77% DockQ-pass. Also covers the ESMC-{300M,600M,6B} protein
$ npx -y skills add aipoch/open-science --skill esmfold2 --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/esmfold2Context preview
The summary Claude sees to decide when to auto-load this skill.
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. 2026, github.com/Biohub/esm). Single-sequence and MSA modes; protein, DNA, RNA, ligand (CCD/SMILES), modified residues. FoldBench Ab-Ag 50-55%, PPI 70-77% DockQ-pass. Also covers the ESMC-{300M,600M,6B} protein
name: esmfold2
description: >
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. 2026,
github.com/Biohub/esm). Single-sequence and MSA modes; protein, DNA, RNA,
ligand (CCD/SMILES), modified residues. FoldBench Ab-Ag 50-55%, PPI 70-77%
DockQ-pass. Also covers the ESMC-{300M,600M,6B} protein language models from
the same release: masked-LM logits, hidden states, mutation scoring, contact
prediction, and the SAE interpretability head. MIT-licensed weights on
HuggingFace org `biohub`. Use this skill when: (1) Predicting complex
structures with single-sequence input, (2) Validating designed binders with
ESMFold2-Fast, (3) Running ESMFold2 with MSA input, (4) Getting ESMC
embeddings or per-residue mutation scores, (5) Choosing kernel backend and
sampling-step settings for paper-faithful throughput.
license: Apache-2.0
category: biomodels
requirements: [gpu]
metadata:
display-name: ESMFold2
# SKILL.md body: "**License:** MIT (code github.com/Biohub/esm + weights HF
# `biohub/*`)"
# github.com/Biohub/esm/blob/main/LICENSE.md: MIT (© 2026 Chan Zuckerberg
# Biohub, Inc.). verified 2026-06-30
third_party:
- kind: weights
name: ESMFold2 / ESMC
provider: Biohub
license: MIT
terms_url: https://github.com/Biohub/esm/blob/main/LICENSE.mdAll-atom diffusion co-folding from the Biohub ESM release (2026). ESMFold2 = 48 pair layers with MSA support; ESMFold2-Fast = 24 layers, single-sequence only, ~1.7x faster.
**License:** MIT (code github.com/Biohub/esm + weights HF `biohub/*`). **Paper:** "Language Modeling Materializes a World Model of Protein Biology" (2026).
CUDA 12.x GPU (H100/A100-class); Python **3.12 only**. Fresh venv; needs egress to HF Hub, GitHub, PyPI:
pip install --no-cache-dir uv
uv venv --python 3.12 /work/venv && source /work/venv/bin/activate
uv pip install \
"torch>=2.5,<2.8" einops "biotite>=1.0" rdkit msgpack-numpy biopython \
scikit-learn brotli attrs pandas cloudpathlib httpx tenacity zstd pydssp \
pygtrie accelerate huggingface_hub safetensors "numpy<3" networkx \
sentencepiece tokenizers regex packaging filelock pyyaml typing_extensions \
"transformers @ git+https://github.com/Biohub/transformers.git@3a8956fb4d4ea16b0ec8e71deef2c2909b6a5cbf"
uv pip install --no-deps "esm @ git+https://github.com/Biohub/esm.git@f652b471"
# OPTIONAL — only affects ESMC attention; trunk speedup comes from set_kernel_backend("fused")
uv pip install ninja packaging wheel setuptools
MAX_JOBS=8 uv pip install --no-deps --no-build-isolation "flash-attn<3"
# Do NOT install transformer-engine — RuntimeError (not ImportError) on import
# slips ESMC's guard and kills ESMFold2Model import.The bundled `esmfold2_gpu` Modal env (remote-compute-modal skill) is the canonical, version-pinned recipe.
**Gotchas:**
from esm.models.esmfold2 import (
ESMFold2InputBuilder, StructurePredictionInput,
ProteinInput, DNAInput, RNAInput, LigandInput, Modification,
)
from transformers.models.esmfold2.modeling_esmfold2 import ESMFold2Model
model = ESMFold2Model.from_pretrained("biohub/ESMFold2").cuda().eval()
# or "biohub/ESMFold2-Fast" (24 layers, no MSA, ~1.7x faster)
# or "biohub/ESMFold2-Experimental{,-Fast}{,-Cutoff2025}" (4 design-critic models)
spi = StructurePredictionInput(sequences=[
ProteinInput(id="A", sequence=target_seq),
ProteinInput(id="B", sequence=binder_seq),
# DNAInput(id="C", sequence="ACGT", modifications=[Modification(position=5, ccd="C36")]),
# RNAInput(id="D", sequence="ACGU"),
# LigandInput(id="L", ccd=["SAH"]), # or smiles="..."
])
# Homodimer: ProteinInput(id=["A","B"], sequence=seq)
results = ESMFold2InputBuilder().fold(
model, spi,
num_loops=10, # paper FoldBench eval: 10; 20-loop variant: 20
num_sampling_steps=68, # paper eval: 68 (truncated EDM)
num_diffusion_samples=5, # paper eval: 5/seed
seed=0,
)
# fold() returns list[Prediction], one per diffusion sample. Each carries
# .plddt [L], .ptm, .iptm, .pae [L,L], .pair_chains_iptm, .complex.to_mmcif().
# Rank by ipTM for complexes / mean pLDDT for monomers:
best = max(results, key=lambda r: float(r.iptm if r.iptm is not None
else r.plddt.mean()))
open("pred.cif", "w").write(best.complex.to_mmcif())**Paper-faithful FoldBench settings:** 10 loops, 68 sampling steps, 25 seeds x 5 diffusion samples; rank by ipTM (complexes) or pLDDT (monomers); MSA mode adds `msa_depth=1024` with 10% column masking and ESMC dropout 0.3.
| repo | size | pair layers | MSA | use | | ----------------------------------------------------------------- | ------------------------- | ----------- | --- | ---------------------- | | `ESMFold2` | 0.94 GB + ccd.pkl 0.42 GB | 48 | yes | full eval | | `ESMFold2-Fast` | 0.76 GB | 24 | no | fast single-seq | | `ESMFold2-Experimental{,-Fast}` | 0.90 / 0.72 GB | 48 / 24 | — | design search (Alg 11) | | `ESMFold2-Experimental{,-Fast}-Cutoff2025` | 0.90 / 0.72 GB | — | — | design search + critic | | `ESMFold2-Experimental-Fast-base{300M,600M,6B}-step{250k..1500k}` | — | — | — |
The open-source AI research workbench for scientific research and agent workflows. Local-first, model-agnostic desktop app with extensible skills, MCP tools and connectors, Python/R execution and traceable artifacts for reproducible research on macOS, Windows and Linux.
Repo: aipoch/open-science
Predict protein structure for monomers and multimers with AlphaFold2 via the ColabFold runner…
Structure prediction for protein, nucleic-acid, and small-molecule complexes with Boltz-2…
Predict genome-wide functional tracks (RNA-seq, CAGE, DNase, ChIP) from DNA sequence with…
Structure prediction for protein, nucleic-acid, and small-molecule complexes with the Chai-1…
Prepare reproducible setup instructions and validate a user-managed named software…