sciagent-skill-creator
Scaffold a new SciAgent-Skills entry. Picks pipeline/toolkit/database/guide template, creates skills/{category}/{name}/SKILL.md with valid frontmatter, appends…
Genomic interval ops on BED/BAM/GFF/VCF. Find overlaps, merge intervals, compute coverage, extract FASTA, find nearest features. Core for ChIP-seq peak annotation, region filtering, genome arithmetic. Use tabix for indexed single-region queries; use deeptools for normalized
$ npx -y skills add jaechang-hits/SciAgent-Skills --skill bedtools-genomic-intervals --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/bedtools-genomic-intervalsContext preview
The summary Claude sees to decide when to auto-load this skill.
Genomic interval ops on BED/BAM/GFF/VCF. Find overlaps, merge intervals, compute coverage, extract FASTA, find nearest features. Core for ChIP-seq peak annotation, region filtering, genome arithmetic. Use tabix for indexed single-region queries; use deeptools for normalized
name: "bedtools-genomic-intervals" description: "Genomic interval ops on BED/BAM/GFF/VCF. Find overlaps, merge intervals, compute coverage, extract FASTA, find nearest features. Core for ChIP-seq peak annotation, region filtering, genome arithmetic. Use tabix for indexed single-region queries; use deeptools for normalized bigWig coverage." license: "GPL-2.0"
bedtools is the standard toolkit for operating on genomic intervals in BED, BAM, GFF, and VCF formats. It solves the core problem of genome arithmetic: finding overlaps between feature sets, computing coverage, extracting sequences, merging adjacent regions, and annotating features with nearest neighbors. bedtools operates on sorted coordinate lists and runs at C speed, making it practical for whole-genome analyses.
> **Check before installing**: The tool may already be available in the current environment (e.g., inside a `pixi` / `conda` env). Run `command -v bedtools` first and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool via `pixi run bedtools` rather than bare `bedtools`.
# Bioconda (recommended) conda install -c bioconda bedtools # Homebrew (macOS) brew install bedtools # Verify bedtools --version # bedtools v2.31.0 # Create genome file from FASTA index samtools faidx reference.fa cut -f1,2 reference.fa.fai > genome.txt # chr → size table
# Find peaks overlapping genes, then merge overlapping peaks bedtools intersect -a peaks.bed -b genes.bed -wa -wb > peaks_with_genes.bed bedtools merge -i peaks.bed > merged_peaks.bed bedtools coverage -a genes.bed -b reads.bam > gene_coverage.bed
Find regions that overlap between two feature sets.
# Basic intersection: output overlapping regions bedtools intersect -a peaks.bed -b genes.bed # Report original A and B features for each overlap bedtools intersect -a peaks.bed -b genes.bed -wa -wb # Count B overlaps per A feature (adds column) bedtools intersect -a peaks.bed -b genes.bed -c # Output: chr1 1000 2000 peak1 gene_count # Peaks with ANY overlap (report each peak once) bedtools intersect -a peaks.bed -b genes.bed -u # Peaks with NO overlap in B (invert filter) bedtools intersect -a peaks.bed -b blacklist.bed -v
# Require reciprocal 50% overlap both ways
bedtools intersect -a exp1.bed -b exp2.bed -f 0.5 -F 0.5 -r
# Same-strand intersections only
bedtools intersect -a peaks.bed -b genes.bed -s
# Multiple database files with overlap counts per file
bedtools intersect -a query.bed -b enhancers.bed promoters.bed \
-names enh prom -C
# Memory-efficient mode for pre-sorted large files
bedtools intersect -a sorted_peaks.bed -b sorted_genes.bed -sortedCombine overlapping intervals and perform set operations.
# Merge overlapping and adjacent intervals sort -k1,1 -k2,2n peaks.bed | bedtools merge -i stdin # Merge intervals within 500 bp of each other bedtools merge -i peaks.bed -d 500 # Merge and count original features bedtools merge -i peaks.bed -c 1 -o count # Output: chr1 1000 5000 3 (3 original peaks merged) # Merge and collapse feature names bedtools merge -i peaks.bed -c 4 -o collapse -delim ";" # Output: chr1 1000 5000 peak1;peak2;peak3
# Subtract B from A (remove covered bases) bedtools subtract -a peaks.bed -b blacklist.bed # Remove entire A feature if ANY B overlap bedtools subtract -a peaks.bed -b exclusion.bed -A # Find genomic gaps (complement of covered regions) bedtools complement -i merged.bed -g genome.txt
Calculate depth and breadth of read coverage over features.
# Coverage stats per feature (count, bases covered, % covered) bedtools coverage -a target_genes.bed -b aligned.bam # Output: chr start end gene n_overlapping_reads bases_covered feature_len fraction_covered # Per-base depth within each feature bedtools coverage -a targets.bed -b aligned.bam -d # Output: chr start end name position depth # Coverage histogram per feature bedtools coverage -a features.bed -b aligned.bam -hist
# Genome-wide BEDGRAPH (coverage per bin) bedtools genomecov -ibam aligned.bam -bg -o coverage.bedgraph # Include zero-coverage regions (for whole-genome coverage) bedtools genomecov -ibam aligned.bam -bga > full_coverage.bedgraph # Per-base depth for whole genome bedtools genomecov -ibam aligned.bam -d > depth.txt # Scaled BEDGRAPH (RPM normalization: total=50M reads → scale=1/50) bedtools genomecov -ibam aligned.bam -bg -scale 0.00000002 > rpm.bedgraph # Strand-specific coverage tracks bedtools genomecov -ibam rnaseq.bam -bg -strand + > forward.bedgraph bedtools genomecov -ibam rnaseq.bam -bg -strand - > reverse.bedgraph
Turn your AI coding agent into a life sciences expert — 199 bioinformatics skills for Claude Code covering RNA-seq, single-cell analysis, genomics, proteomics, drug discovery, and more. Boosted BixBench from 65% to 92%. Open source.
Scaffold a new SciAgent-Skills entry. Picks pipeline/toolkit/database/guide template, creates skills/{category}/{name}/SKILL.md with valid frontmatter, appends…
Bayesian modeling with PyMC 5: priors, likelihood, NUTS/ADVI sampling, diagnostics (R-hat, ESS), LOO/WAIC comparison, prediction. Hierarchical, logistic, GP…
Time-to-event modeling with scikit-survival: Cox PH (elastic net), Random Survival Forests, Boosting, SVMs for censored data. C-index, Brier, time-dependent…
Guided statistical analysis: test choice, assumption checks, effect sizes, power, APA reporting. Pick tests, verify assumptions, or format results for…
Python statistical modeling: regression (OLS, WLS, GLM), discrete (Logit, Poisson, NegBin), time series (ARIMA, SARIMAX, VAR), with rigorous inference,…
DL cell/nucleus segmentation for fluorescence and brightfield microscopy. Pre-trained models (cyto3, nuclei, tissuenet) and a generalist flow-based algorithm…