Metapipeline-DNA: A comprehensive germline and somatic genomics Nextflow pipeline

Mar 1, 2026·
Yash Patel
Chenghao Zhu
Chenghao Zhu
,
Takafumi N. Yamaguchi
,
Nicholas K. Wang
,
Nicholas Wiltsie
,
Nicole Zeltser
,
Alfredo E. Gonzalez
,
Helena Winata
,
Yu Pan
,
Mohammed Faizal Eeman Mootor
,
Timothy Sanders
,
Sorel T. Fitz-Gibbon
,
Cyriac Kandoth
,
Julie Livingstone
,
Lydia Liu
,
Benjamin Carlin
,
Aaron Holmes
,
Jieun Oh
,
John M. Sahrmann
,
Shu Tao
,
Stefan E. Eng
,
Rupert Hugh-White
,
Kiarod Pashminehazar
,
Arpi Beshlikyan
,
Madison Jordan
,
Selina Wu
,
Mao Tian
,
Jaron Arbet
,
Beth K. Neilsen
,
Roni Haas
,
Yuan Zhe Bugh
,
Gina Kim
,
Joseph Salmingo
,
Wenshu Zhang
,
Aakarsh Anand
,
Edward Hwang
,
Anna Neiman-Golden
,
Philippa Steinberg
,
Wenyan Zhao
,
Prateek Anand
,
Raag Agrawal
,
Brandon L. Tsai
,
Paul C. Boutros
· 0 min read
Abstract
The price, quality, and throughput of DNA sequencing continue to improve. Algorithmic innovations have allowed inference of a growing range of features from DNA sequencing data, quantifying nuclear, mitochondrial, and evolutionary aspects of both germline and somatic genomes. To automate analyses of the full range of genomic characteristics, we created an extensible Nextflow metapipeline called metapipeline-DNA. It analyzes targeted and whole-genome sequencing data from raw reads through preprocessing, feature detection by multiple algorithms, quality control, and data-visualization. Each step can be run independently and is supported by robust software engineering including automated failure-recovery, granular testing, and consistent verifications of inputs, outputs, and parameters. Metapipeline-DNA is cloud-compatible and highly configurable, with options to subset and optimize each analysis. Metapipeline-DNA facilitates high-scale, comprehensive analysis of DNA sequencing data, and is open-source under the GPLv2 license.
Type
Publication
Cell Reports Methods
publications
Chenghao Zhu
Authors
Research Assistant Professor
Chenghao Zhu is a Research Assistant Professor in the NCI-designated Cancer Center at Sanford Burnham Prebys. His research focuses on developing computational methods and software for proteogenomics and applying proteogenomics to cancer diagnosis, prognosis, and clinico-epidemiologic questions.