/bio-consensus-sequences
Generate consensus FASTA sequences by applying VCF variants to a reference using bcftools consensus. Use when creating sample-specific reference sequences or reconstructing haplotypes.
$ npx -y skills add FreedomIntelligence/OpenClaw-Medical-Skills --skill bio-consensus-sequences --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/bio-consensus-sequences
Context preview
The summary Claude sees to decide when to auto-load this skill.
Generate consensus FASTA sequences by applying VCF variants to a reference using bcftools consensus. Use when creating sample-specific reference sequences or reconstructing haplotypes.
SKILL.md
bio-consensus-sequences.SKILL.mdname: bio-consensus-sequences
description: Generate consensus FASTA sequences by applying VCF variants to a reference using bcftools consensus. Use when creating sample-specific reference sequences or reconstructing haplotypes.
tool_type: cli
primary_tool: bcftools
Version Compatibility
Reference examples tested with: BioPython 1.83+, bcftools 1.19+, bedtools 2.31+, minimap2 2.26+, samtools 1.19+
Before using code patterns, verify installed versions match. If versions differ:
- Python: `pip show <package>` then `help(module.function)` to check signatures
- CLI: `<tool> --version` then `<tool> --help` to confirm flags
If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Consensus Sequences
**"Generate a consensus sequence from my VCF"** → Apply called variants to a reference FASTA, producing a sample-specific genome with optional haplotype selection and low-coverage masking.
- CLI: `bcftools consensus -f reference.fa input.vcf.gz`
- Python: `cyvcf2` + `Bio.SeqIO` for simple SNP-only cases
Basic Usage
Generate Consensus
bcftools consensus -f reference.fa input.vcf.gz > consensus.fa
Specify Sample
bcftools consensus -f reference.fa -s sample1 input.vcf.gz > sample1.fa
Output to File
bcftools consensus -f reference.fa -o consensus.fa input.vcf.gz
Haplotype Selection
First Haplotype Only
bcftools consensus -f reference.fa -H 1 input.vcf.gz > haplotype1.fa
Second Haplotype Only
bcftools consensus -f reference.fa -H 2 input.vcf.gz > haplotype2.fa
Haplotype Options
| Option | Description | |--------|-------------| | `-H 1` | First haplotype | | `-H 2` | Second haplotype | | `-H A` | Apply all ALT alleles | | `-H R` | Apply REF alleles where heterozygous | | `-I` | Apply IUPAC ambiguity codes (separate flag) |
IUPAC Codes for Heterozygous Sites
bcftools consensus -f reference.fa -I input.vcf.gz > consensus_iupac.fa
Heterozygous sites encoded with IUPAC ambiguity codes:
- A/G → R
- C/T → Y
- A/C → M
- G/T → K
- A/T → W
- C/G → S
Missing Data Handling
Mark Missing as N
bcftools consensus -f reference.fa -M N input.vcf.gz > consensus.fa
Mark Low Coverage as N
Using a mask BED file:
# Create mask from depth
samtools depth input.bam | awk '$3<10 {print $1"\t"$2-1"\t"$2}' > low_coverage.bed
# Apply mask
bcftools consensus -f reference.fa -m low_coverage.bed input.vcf.gz > consensus.faMask Options
| Option | Description | |--------|-------------| | `-m FILE` | Mask regions in BED file with N | | `-M CHAR` | Character for masked regions (default N) |
Region Selection
Specific Region
bcftools consensus -f reference.fa -r chr1:1000-2000 input.vcf.gz > region.fa
Multiple Regions
Use with BED file to extract multiple regions.
Chain Files
Generate Chain File
bcftools consensus -f reference.fa -c chain.txt input.vcf.gz > consensus.fa
Chain files map coordinates between reference and consensus:
- Useful for liftover of annotations
- Required when indels change sequence length
Chain File Format
chain score ref_name ref_size ref_strand ref_start ref_end query_name query_size query_strand query_start query_end id
Sample-Specific Consensus
For Each Sample
for sample in $(bcftools query -l input.vcf.gz); do
bcftools consensus -f reference.fa -s "$sample" input.vcf.gz > "${sample}.fa"
doneBoth Haplotypes
sample="sample1"
bcftools consensus -f reference.fa -s "$sample" -H 1 input.vcf.gz > "${sample}_hap1.fa"
bcftools consensus -f reference.fa -s "$sample" -H 2 input.vcf.gz > "${sample}_hap2.fa"Filtering Before Consensus
PASS Variants Only
bcftools view -f PASS input.vcf.gz | \
bcftools consensus -f reference.fa > consensus.faHigh-Quality Variants Only
bcftools filter -i 'QUAL>=30 && INFO/DP>=10' input.vcf.gz | \
bcftools consensus -f reference.fa > consensus.faSNPs Only
bcftools view -v snps input.vcf.gz | \
bcftools consensus -f reference.fa > consensus_snps.faSequence Naming
Default Naming
Output uses reference sequence names.
Custom Prefix
bcftools consensus -f reference.fa -p "sample1_" input.vcf.gz > consensus.fa
Sequences named: `sample1_chr1`, `sample1_chr2`, etc.
Common Workflows
**Goal:** Generate consensus sequences for downstream analyses like phylogenetics, viral surveillance, or gene-level comparison.
**Approach:** Filter variants to high-quality calls, apply per-sample consensus generation, mask low-coverage regions with N, then combine for multi-sample workflows.
Phylogenetic Analysis Preparation
# For each sample, generate consensus
mkdir -p consensus
for sample in $(bcftools query -l cohort.vcf.gz); do
bcftools view -s "$sample" cohort.vcf.gz | \
bcftools view -c 1 | \
bcftools consensus -f reference.fa > "consensus/${sample}.fa"
done
# Combine for alignment
cat consensus/*.fa > all_samples.faViral Genome Assembly
# Apply high-quality variants only
bcftools filter -i 'QUAL>=30 && INFO/DP>=20' variants.vcf.gz | \
bcftools view -f PASS | \
bcftools consensus -f reference.fa -M N > consensus.faGene-Specific Consensus
# Extract gene region
bcftools consensus -f reference.fa -r chr1:1000000-1010000 \
-s sample1 variants.vcf.gz > gene.faMasked Low-Coverage Regions
# Create mask from coverage
samtools depth -a input.bam | \
awk '$3<5 {print $1"\t"$2-1"\t"$2}' | \
bedtools merge > low_coverage.bed
# Generate consensus with mask
bcftools consensus -f reference.fa -m low_coverage.bed \
variants.vcf.gz > consensus.faVerify Consensus
Check Differences
# Align consensus t
Read more
name: bio-consensus-sequences description: Generate consensus FASTA sequences by applying VCF variants to a reference using bcftools consensus. Use when creating sample-specific reference sequences or reconstructing haplotypes. tool_type: cli primary_tool: bcftools
Version Compatibility
Reference examples tested with: BioPython 1.83+, bcftools 1.19+, bedtools 2.31+, minimap2 2.26+, samtools 1.19+
Before using code patterns, verify installed versions match. If versions differ:
- Python: `pip show <package>` then `help(module.function)` to check signatures
- CLI: `<tool> --version` then `<tool> --help` to confirm flags
If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Consensus Sequences
**"Generate a consensus sequence from my VCF"** → Apply called variants to a reference FASTA, producing a sample-specific genome with optional haplotype selection and low-coverage masking.
- CLI: `bcftools consensus -f reference.fa input.vcf.gz`
- Python: `cyvcf2` + `Bio.SeqIO` for simple SNP-only cases
Basic Usage
Generate Consensus
bcftools consensus -f reference.fa input.vcf.gz > consensus.fa
Specify Sample
bcftools consensus -f reference.fa -s sample1 input.vcf.gz > sample1.fa
Output to File
bcftools consensus -f reference.fa -o consensus.fa input.vcf.gz
Haplotype Selection
First Haplotype Only
bcftools consensus -f reference.fa -H 1 input.vcf.gz > haplotype1.fa
Second Haplotype Only
bcftools consensus -f reference.fa -H 2 input.vcf.gz > haplotype2.fa
Haplotype Options
| Option | Description | |--------|-------------| | `-H 1` | First haplotype | | `-H 2` | Second haplotype | | `-H A` | Apply all ALT alleles | | `-H R` | Apply REF alleles where heterozygous | | `-I` | Apply IUPAC ambiguity codes (separate flag) |
IUPAC Codes for Heterozygous Sites
bcftools consensus -f reference.fa -I input.vcf.gz > consensus_iupac.fa
Heterozygous sites encoded with IUPAC ambiguity codes:
- A/G → R
- C/T → Y
- A/C → M
- G/T → K
- A/T → W
- C/G → S
Missing Data Handling
Mark Missing as N
bcftools consensus -f reference.fa -M N input.vcf.gz > consensus.fa
Mark Low Coverage as N
Using a mask BED file:
# Create mask from depth
samtools depth input.bam | awk '$3<10 {print $1"\t"$2-1"\t"$2}' > low_coverage.bed
# Apply mask
bcftools consensus -f reference.fa -m low_coverage.bed input.vcf.gz > consensus.faMask Options
| Option | Description | |--------|-------------| | `-m FILE` | Mask regions in BED file with N | | `-M CHAR` | Character for masked regions (default N) |
Region Selection
Specific Region
bcftools consensus -f reference.fa -r chr1:1000-2000 input.vcf.gz > region.fa
Multiple Regions
Use with BED file to extract multiple regions.
Chain Files
Generate Chain File
bcftools consensus -f reference.fa -c chain.txt input.vcf.gz > consensus.fa
Chain files map coordinates between reference and consensus:
- Useful for liftover of annotations
- Required when indels change sequence length
Chain File Format
chain score ref_name ref_size ref_strand ref_start ref_end query_name query_size query_strand query_start query_end id
Sample-Specific Consensus
For Each Sample
for sample in $(bcftools query -l input.vcf.gz); do
bcftools consensus -f reference.fa -s "$sample" input.vcf.gz > "${sample}.fa"
doneBoth Haplotypes
sample="sample1"
bcftools consensus -f reference.fa -s "$sample" -H 1 input.vcf.gz > "${sample}_hap1.fa"
bcftools consensus -f reference.fa -s "$sample" -H 2 input.vcf.gz > "${sample}_hap2.fa"Filtering Before Consensus
PASS Variants Only
bcftools view -f PASS input.vcf.gz | \
bcftools consensus -f reference.fa > consensus.faHigh-Quality Variants Only
bcftools filter -i 'QUAL>=30 && INFO/DP>=10' input.vcf.gz | \
bcftools consensus -f reference.fa > consensus.faSNPs Only
bcftools view -v snps input.vcf.gz | \
bcftools consensus -f reference.fa > consensus_snps.faSequence Naming
Default Naming
Output uses reference sequence names.
Custom Prefix
bcftools consensus -f reference.fa -p "sample1_" input.vcf.gz > consensus.fa
Sequences named: `sample1_chr1`, `sample1_chr2`, etc.
Common Workflows
**Goal:** Generate consensus sequences for downstream analyses like phylogenetics, viral surveillance, or gene-level comparison.
**Approach:** Filter variants to high-quality calls, apply per-sample consensus generation, mask low-coverage regions with N, then combine for multi-sample workflows.
Phylogenetic Analysis Preparation
# For each sample, generate consensus
mkdir -p consensus
for sample in $(bcftools query -l cohort.vcf.gz); do
bcftools view -s "$sample" cohort.vcf.gz | \
bcftools view -c 1 | \
bcftools consensus -f reference.fa > "consensus/${sample}.fa"
done
# Combine for alignment
cat consensus/*.fa > all_samples.faViral Genome Assembly
# Apply high-quality variants only
bcftools filter -i 'QUAL>=30 && INFO/DP>=20' variants.vcf.gz | \
bcftools view -f PASS | \
bcftools consensus -f reference.fa -M N > consensus.faGene-Specific Consensus
# Extract gene region
bcftools consensus -f reference.fa -r chr1:1000000-1010000 \
-s sample1 variants.vcf.gz > gene.faMasked Low-Coverage Regions
# Create mask from coverage
samtools depth -a input.bam | \
awk '$3<5 {print $1"\t"$2-1"\t"$2}' | \
bedtools merge > low_coverage.bed
# Generate consensus with mask
bcftools consensus -f reference.fa -m low_coverage.bed \
variants.vcf.gz > consensus.faVerify Consensus
Check Differences
# Align consensus t
The largest open-source medical AI skill library for OpenClaw.
Other skills on openclaw-medical-skills.
adaptyv
Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding…
adhd-daily-planner
Time-blind friendly planning, executive function support, and daily structure for ADHD brains. Specializes in realistic time estimation, dopamine-aware task…
aeon
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection,…
agent-browser
Browse the web for any task — research topics, read articles, interact with web apps, fill forms, take screenshots, extract data, and test web pages. Use…

