Skip to content

/pipeline-hic

Execute ENCODE Hi-C pipeline from FASTQ to contact matrices and loop calls. Child of pipeline-guide. Provides Nextflow execution with Docker and cloud deployment. Use when processing Hi-C data, generating contact matrices, calling loops or TADs. Trigger on: Hi-C pipeline,

From plugin
2994 skills7 agents10 commands
shell
$ npx -y skills add ammawla/encode-toolkit --skill pipeline-hic --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.
  • You can call itInvoke it directly when you want it.
  • Slash command/pipeline-hic
How auto-invocation works

Context preview

The summary Claude sees to decide when to auto-load this skill.

Execute ENCODE Hi-C pipeline from FASTQ to contact matrices and loop calls. Child of pipeline-guide. Provides Nextflow execution with Docker and cloud deployment. Use when processing Hi-C data, generating contact matrices, calling loops or TADs. Trigger on: Hi-C pipeline,

SKILL.md

pipeline-hic.SKILL.md
name: pipeline-hic
description: "Execute ENCODE Hi-C pipeline from FASTQ to contact matrices and loop calls. Child of pipeline-guide. Provides Nextflow execution with Docker and cloud deployment. Use when processing Hi-C data, generating contact matrices, calling loops or TADs. Trigger on: Hi-C pipeline, chromatin conformation, contact matrix, loop calling, TAD detection, Juicer, HiCCUPS, 3D genome."

ENCODE Hi-C Pipeline: FASTQ to Contact Matrices and Loops

When to Use

  • User wants to run a Hi-C processing pipeline from FASTQ to contact matrices and loop calls
  • User asks about "Hi-C pipeline", "contact matrix", "loop calling", "Juicer", "HiCCUPS", or "TAD detection"
  • User needs to process Hi-C data for 3D genome structure analysis
  • Example queries: "process my Hi-C FASTQs", "generate contact matrices from Hi-C", "call chromatin loops with HiCCUPS"

Execute the ENCODE Hi-C pipeline for chromatin conformation capture data, producing multi-resolution contact matrices, loop calls, and compartment annotations.

Pipeline Overview

FASTQ -> Trim -> BWA (per-mate) -> pairtools parse -> dedup -> .pairs
                                                                 |
                                                    +------------+------------+
                                                    |                         |
                                              Juicer pre -> .hic        cooler -> .mcool
                                                    |                         |
                                              HiCCUPS loops              Compartments

ENCODE Repository

  • **GitHub**: `ENCODE-DCC/hic-pipeline`
  • **Container**: `encodedcc/hic-pipeline`
  • **WDL**: Available for Cromwell execution
  • **This skill**: Nextflow DSL2 reimplementation for portability

Core Tools and Versions

| Tool | Version | Purpose | Citation | |------|---------|---------|----------| | BWA-MEM | 0.7.17 | Alignment (per-mate) | Li & Durbin 2009 | | pairtools | 1.0.3 | Pair classification, dedup | Open2C | | Juicer tools | 2.20.00 | .hic generation, HiCCUPS | Durand et al. 2016 | | cooler | 0.9.3 | .cool/.mcool generation | Abdennur & Mirny 2020 | | samtools | 1.19 | BAM operations | Li et al. 2009 | | FastQC | 0.12.1 | Read quality | Andrews (Babraham) | | MultiQC | 1.21 | Aggregated QC | Ewels et al. 2016 |

Key Literature

1. **Rao et al. 2014** - "A 3D Map of the Human Genome at Kilobase Resolution Reveals Principles of Chromatin Looping" (Cell, ~5,000 citations) DOI: 10.1016/j.cell.2014.11.021

2. **Lieberman-Aiden et al. 2009** - "Comprehensive Mapping of Long-Range Interactions Reveals Folding Principles of the Human Genome" (Science, ~6,000 citations) DOI: 10.1126/science.1181369

3. **Durand et al. 2016** - "Juicer Provides a One-Click System for Analyzing Loop-Resolution Hi-C Experiments" (Cell Systems, ~2,000 citations) DOI: 10.1016/j.cels.2016.07.002

4. **Abdennur & Mirny 2020** - "Cooler: scalable storage for Hi-C data and other genomically labeled arrays" (Bioinformatics) DOI: 10.1093/bioinformatics/btz540

5. **Amemiya et al. 2019** - "The ENCODE Blacklist" (Scientific Reports, ~1,372 citations) DOI: 10.1038/s41598-019-45839-z

Execution

Quick Start (Local)

nextflow run main.nf \
    -profile local \
    --reads '/data/fastq/*_R{1,2}.fastq.gz' \
    --bwa_index '/ref/bwa_index/genome.fa' \
    --chrom_sizes '/ref/hg38.chrom.sizes' \
    --restriction_site 'GATC' \
    --outdir results/ \
    -resume

SLURM HPC

nextflow run main.nf \
    -profile slurm \
    --reads '/data/fastq/*_R{1,2}.fastq.gz' \
    --bwa_index '/ref/bwa_index/genome.fa' \
    --chrom_sizes '/ref/hg38.chrom.sizes' \
    --restriction_site 'GATC' \
    --outdir results/ \
    -resume

Cloud (GCP / AWS)

nextflow run main.nf \
    -profile gcp \
    --reads 'gs://bucket/fastq/*_R{1,2}.fastq.gz' \
    --bwa_index 'gs://bucket/ref/genome.fa' \
    --chrom_sizes 'gs://bucket/ref/hg38.chrom.sizes' \
    --restriction_site 'GATC' \
    --outdir 'gs://bucket/results/' \
    -resume

Resource Requirements

| Step | CPUs | RAM | Time (2B contacts) | |------|------|-----|---------------------| | BWA alignment | 8 | 16 GB | 4-6 hours | | pairtools parse | 4 | 8 GB | 2-3 hours | | pairtools dedup | 4 | 16 GB | 1-2 hours | | Juicer pre + hic | 4 | 64 GB | 2-4 hours | | HiCCUPS | 4 | 16 GB (+ GPU optional) | 1-2 hours | | **Total** | **8** | **64 GB** | **8-16 hours** |

Pipeline Parameters

| Parameter | Default | Description | |-----------|---------|-------------| | `--reads` | required | Glob pattern to paired FASTQ files | | `--bwa_index` | required | Path to BWA genome index (.fa with .bwt etc.) | | `--chrom_sizes` | required | Chromosome sizes file | | `--restriction_site` | `GATC` | Restriction enzyme site (GATC for MboI/DpnII) | | `--outdir` | `./results` | Output directory | | `--resolutions` | `1000,5000,10000,25000,50000,100000,250000,500000,1000000` | Matrix resolutions | | `--min_mapq` | `30` | Minimum MAPQ for pair filtering | | `--assembly` | `hg38` | Genome assembly name for .hic header |

Output Files

results/
  fastqc/                         # Raw read quality
  alignment/
    {sample}.R1.bam               # Per-mate alignments
    {sample}.R2.bam
  pairs/
    {sample}.pairs.gz             # Classified, deduplicated pairs
    {sample}.dedup_stats.txt      # Duplication metrics
    {sample}.pair_stats.txt       # Pair type classification
  matrices/
    {sample}.hic                  # Juicer .hic file (primary output)
    {sample}.mcool                # Cooler multi-resolution matrix
  loops/
    {sample}.hiccups_loops.bedpe  # Called loops (HiCCUPS)
  qc/
    {sample}.contact_stats.txt    # Contact statistics
  multiqc/
    multiqc_report.html

.hic File Format

The .hic format (Juicer) stores multi-resolution contact matrices with normalization vectors. Can be visua

Read more
Read it on GitHub ↗

Showing the first part of this file.

Ships withencode-toolkit

Search ENCODE, cross-reference 14 databases, run 7 analysis pipelines, and generate publication-ready methods — all from natural language in Claude Code.

Get the whole plugin, auto-invoked
Stats
29
Stars
0
Views
5
Forks
Active
Maintenance
Python
Language
AGPL-3.0
License
8d ago
Last commit
4mo ago
Created

Repo: ammawla/encode-toolkit

Other skills on encode-toolkit.