/bio-alignment-msa-parsing
Parse and analyze multiple sequence alignments using Biopython. Extract sequences, identify conserved regions, analyze gaps, work with annotations, and manipulate alignment data for downstream analysis. Use when parsing or manipulating multiple sequence alignments.
$ npx -y skills add FreedomIntelligence/OpenClaw-Medical-Skills --skill bio-alignment-msa-parsing --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/bio-alignment-msa-parsing
Context preview
The summary Claude sees to decide when to auto-load this skill.
Parse and analyze multiple sequence alignments using Biopython. Extract sequences, identify conserved regions, analyze gaps, work with annotations, and manipulate alignment data for downstream analysis. Use when parsing or manipulating multiple sequence alignments.
SKILL.md
bio-alignment-msa-parsing.SKILL.mdname: bio-alignment-msa-parsing
description: Parse and analyze multiple sequence alignments using Biopython. Extract sequences, identify conserved regions, analyze gaps, work with annotations, and manipulate alignment data for downstream analysis. Use when parsing or manipulating multiple sequence alignments.
tool_type: python
primary_tool: Bio.AlignIO
Version Compatibility
Reference examples tested with: BioPython 1.83+
Before using code patterns, verify installed versions match. If versions differ:
- Python: `pip show <package>` then `help(module.function)` to check signatures
If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
MSA Parsing and Analysis
Parse multiple sequence alignments to extract information, analyze content, and prepare for downstream analysis.
Required Import
**Goal:** Load modules for parsing, analyzing, and manipulating multiple sequence alignments.
**Approach:** Import AlignIO for reading, Counter for column analysis, and alignment classes for constructing modified alignments.
from Bio import AlignIO
from Bio.Align import MultipleSeqAlignment
from Bio.SeqRecord import SeqRecord
from Bio.Seq import Seq
from collections import Counter
Loading Alignments
**Goal:** Read an MSA file and inspect its dimensions.
**Approach:** Use `AlignIO.read()` specifying the file and format.
from Bio import AlignIO
alignment = AlignIO.read('alignment.fasta', 'fasta')
print(f'{len(alignment)} sequences, {alignment.get_alignment_length()} columns')Extracting Sequence Information
Get All Sequence IDs
seq_ids = [record.id for record in alignment]
Get Sequences as Strings
sequences = [str(record.seq) for record in alignment]
Get Sequence by ID
def get_sequence_by_id(alignment, seq_id):
for record in alignment:
if record.id == seq_id:
return record
return None
target = get_sequence_by_id(alignment, 'species_A')Access Descriptions and Annotations
for record in alignment:
print(f'ID: {record.id}')
print(f'Description: {record.description}')
print(f'Annotations: {record.annotations}')Column-wise Analysis
**Goal:** Analyze alignment content column by column to assess composition, conservation, and variability.
**Approach:** Use column indexing (`alignment[:, idx]`) and Counter to examine character frequencies at each position.
Get Single Column
column_5 = alignment[:, 5] # Returns string of characters at position 5
print(column_5) # e.g., 'AAAGA'
Iterate Over Columns
for col_idx in range(alignment.get_alignment_length()):
column = alignment[:, col_idx]
print(f'Column {col_idx}: {column}')Count Characters in Column
from collections import Counter
def column_composition(alignment, col_idx):
column = alignment[:, col_idx]
return Counter(column)
counts = column_composition(alignment, 0)
print(counts) # Counter({'A': 3, 'G': 1, '-': 1})Find Conserved Positions
def find_conserved_positions(alignment, threshold=1.0):
conserved = []
for col_idx in range(alignment.get_alignment_length()):
column = alignment[:, col_idx]
counts = Counter(column)
most_common_char, most_common_count = counts.most_common(1)[0]
if most_common_char != '-':
conservation = most_common_count / len(alignment)
if conservation >= threshold:
conserved.append((col_idx, most_common_char))
return conserved
fully_conserved = find_conserved_positions(alignment, threshold=1.0)
mostly_conserved = find_conserved_positions(alignment, threshold=0.8)Gap Analysis
**Goal:** Quantify gap distribution across sequences and columns to identify problematic regions or sequences.
**Approach:** Count gap characters per sequence and per column, then identify positions exceeding a gap fraction threshold.
Count Gaps Per Sequence
gap_counts = [(record.id, str(record.seq).count('-')) for record in alignment]
for seq_id, gaps in gap_counts:
print(f'{seq_id}: {gaps} gaps')Count Gaps Per Column
def gaps_per_column(alignment):
return [alignment[:, i].count('-') for i in range(alignment.get_alignment_length())]
gap_profile = gaps_per_column(alignment)Find Gappy Columns
def find_gappy_columns(alignment, threshold=0.5):
gappy = []
num_seqs = len(alignment)
for col_idx in range(alignment.get_alignment_length()):
column = alignment[:, col_idx]
gap_fraction = column.count('-') / num_seqs
if gap_fraction >= threshold:
gappy.append(col_idx)
return gappy
columns_to_remove = find_gappy_columns(alignment, threshold=0.5)Remove Gappy Columns
def remove_gappy_columns(alignment, threshold=0.5):
num_seqs = len(alignment)
keep_columns = []
for col_idx in range(alignment.get_alignment_length()):
column = alignment[:, col_idx]
gap_fraction = column.count('-') / num_seqs
if gap_fraction < threshold:
keep_columns.append(col_idx)
new_records = []
for record in alignment:
new_seq = ''.join(str(record.seq)[i] for i in keep_columns)
new_records.append(SeqRecord(Seq(new_seq), id=record.id, description=record.description))
return MultipleSeqAlignment(new_records)
cleaned = remove_gappy_columns(alignment, threshold=0.5)Consensus Sequence
**"Get consensus sequence"** → Derive a single representative sequence from an MSA based on majority-rule voting at each column.
**Goal:** Generate a consensus sequence from the alignment using a frequency threshold.
**Approach:** At each column, select the most common non-gap character if it exceeds the threshold; otherwise mark as ambiguous.
Simp
Read more
name: bio-alignment-msa-parsing description: Parse and analyze multiple sequence alignments using Biopython. Extract sequences, identify conserved regions, analyze gaps, work with annotations, and manipulate alignment data for downstream analysis. Use when parsing or manipulating multiple sequence alignments. tool_type: python primary_tool: Bio.AlignIO
Version Compatibility
Reference examples tested with: BioPython 1.83+
Before using code patterns, verify installed versions match. If versions differ:
- Python: `pip show <package>` then `help(module.function)` to check signatures
If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
MSA Parsing and Analysis
Parse multiple sequence alignments to extract information, analyze content, and prepare for downstream analysis.
Required Import
**Goal:** Load modules for parsing, analyzing, and manipulating multiple sequence alignments.
**Approach:** Import AlignIO for reading, Counter for column analysis, and alignment classes for constructing modified alignments.
from Bio import AlignIO from Bio.Align import MultipleSeqAlignment from Bio.SeqRecord import SeqRecord from Bio.Seq import Seq from collections import Counter
Loading Alignments
**Goal:** Read an MSA file and inspect its dimensions.
**Approach:** Use `AlignIO.read()` specifying the file and format.
from Bio import AlignIO
alignment = AlignIO.read('alignment.fasta', 'fasta')
print(f'{len(alignment)} sequences, {alignment.get_alignment_length()} columns')Extracting Sequence Information
Get All Sequence IDs
seq_ids = [record.id for record in alignment]
Get Sequences as Strings
sequences = [str(record.seq) for record in alignment]
Get Sequence by ID
def get_sequence_by_id(alignment, seq_id):
for record in alignment:
if record.id == seq_id:
return record
return None
target = get_sequence_by_id(alignment, 'species_A')Access Descriptions and Annotations
for record in alignment:
print(f'ID: {record.id}')
print(f'Description: {record.description}')
print(f'Annotations: {record.annotations}')Column-wise Analysis
**Goal:** Analyze alignment content column by column to assess composition, conservation, and variability.
**Approach:** Use column indexing (`alignment[:, idx]`) and Counter to examine character frequencies at each position.
Get Single Column
column_5 = alignment[:, 5] # Returns string of characters at position 5 print(column_5) # e.g., 'AAAGA'
Iterate Over Columns
for col_idx in range(alignment.get_alignment_length()):
column = alignment[:, col_idx]
print(f'Column {col_idx}: {column}')Count Characters in Column
from collections import Counter
def column_composition(alignment, col_idx):
column = alignment[:, col_idx]
return Counter(column)
counts = column_composition(alignment, 0)
print(counts) # Counter({'A': 3, 'G': 1, '-': 1})Find Conserved Positions
def find_conserved_positions(alignment, threshold=1.0):
conserved = []
for col_idx in range(alignment.get_alignment_length()):
column = alignment[:, col_idx]
counts = Counter(column)
most_common_char, most_common_count = counts.most_common(1)[0]
if most_common_char != '-':
conservation = most_common_count / len(alignment)
if conservation >= threshold:
conserved.append((col_idx, most_common_char))
return conserved
fully_conserved = find_conserved_positions(alignment, threshold=1.0)
mostly_conserved = find_conserved_positions(alignment, threshold=0.8)Gap Analysis
**Goal:** Quantify gap distribution across sequences and columns to identify problematic regions or sequences.
**Approach:** Count gap characters per sequence and per column, then identify positions exceeding a gap fraction threshold.
Count Gaps Per Sequence
gap_counts = [(record.id, str(record.seq).count('-')) for record in alignment]
for seq_id, gaps in gap_counts:
print(f'{seq_id}: {gaps} gaps')Count Gaps Per Column
def gaps_per_column(alignment):
return [alignment[:, i].count('-') for i in range(alignment.get_alignment_length())]
gap_profile = gaps_per_column(alignment)Find Gappy Columns
def find_gappy_columns(alignment, threshold=0.5):
gappy = []
num_seqs = len(alignment)
for col_idx in range(alignment.get_alignment_length()):
column = alignment[:, col_idx]
gap_fraction = column.count('-') / num_seqs
if gap_fraction >= threshold:
gappy.append(col_idx)
return gappy
columns_to_remove = find_gappy_columns(alignment, threshold=0.5)Remove Gappy Columns
def remove_gappy_columns(alignment, threshold=0.5):
num_seqs = len(alignment)
keep_columns = []
for col_idx in range(alignment.get_alignment_length()):
column = alignment[:, col_idx]
gap_fraction = column.count('-') / num_seqs
if gap_fraction < threshold:
keep_columns.append(col_idx)
new_records = []
for record in alignment:
new_seq = ''.join(str(record.seq)[i] for i in keep_columns)
new_records.append(SeqRecord(Seq(new_seq), id=record.id, description=record.description))
return MultipleSeqAlignment(new_records)
cleaned = remove_gappy_columns(alignment, threshold=0.5)Consensus Sequence
**"Get consensus sequence"** → Derive a single representative sequence from an MSA based on majority-rule voting at each column.
**Goal:** Generate a consensus sequence from the alignment using a frequency threshold.
**Approach:** At each column, select the most common non-gap character if it exceeds the threshold; otherwise mark as ambiguous.
Simp
The largest open-source medical AI skill library for OpenClaw.
Other skills on openclaw-medical-skills.
- /aav-vector-design-agent
<!--
Open skill - /adaptyv
Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use
Open skill - /adhd-daily-planner
Time-blind friendly planning, executive function support, and daily structure for ADHD brains. Specializes in realistic time estimation, dopamine-aware task design, and building systems that
Open skill - /aeon
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations
Open skill - /agent-browser
Browse the web for any task — research topics, read articles, interact with web apps, fill forms, take screenshots, extract data, and test web pages. Use whenever a browser would be useful, not just when the user explicitly asks.
Open skill - /agentd-drug-discovery
<!--
Open skill

