Metapipeline-DNA: A comprehensive germline and somatic genomics Nextflow pipeline
Mar 1, 2026·
,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,·
0 min read
Yash Patel
Chenghao Zhu
Takafumi N. Yamaguchi
Nicholas K. Wang
Nicholas Wiltsie
Nicole Zeltser
Alfredo E. Gonzalez
Helena Winata
Yu Pan
Mohammed Faizal Eeman Mootor
Timothy Sanders
Sorel T. Fitz-Gibbon
Cyriac Kandoth
Julie Livingstone
Lydia Liu
Benjamin Carlin
Aaron Holmes
Jieun Oh
John M. Sahrmann
Shu Tao
Stefan E. Eng
Rupert Hugh-White
Kiarod Pashminehazar
Arpi Beshlikyan
Madison Jordan
Selina Wu
Mao Tian
Jaron Arbet
Beth K. Neilsen
Roni Haas
Yuan Zhe Bugh
Gina Kim
Joseph Salmingo
Wenshu Zhang
Aakarsh Anand
Edward Hwang
Anna Neiman-Golden
Philippa Steinberg
Wenyan Zhao
Prateek Anand
Raag Agrawal
Brandon L. Tsai
Paul C. Boutros
Abstract
The price, quality, and throughput of DNA sequencing continue to improve. Algorithmic innovations have allowed inference of a growing range of features from DNA sequencing data, quantifying nuclear, mitochondrial, and evolutionary aspects of both germline and somatic genomes. To automate analyses of the full range of genomic characteristics, we created an extensible Nextflow metapipeline called metapipeline-DNA. It analyzes targeted and whole-genome sequencing data from raw reads through preprocessing, feature detection by multiple algorithms, quality control, and data-visualization. Each step can be run independently and is supported by robust software engineering including automated failure-recovery, granular testing, and consistent verifications of inputs, outputs, and parameters. Metapipeline-DNA is cloud-compatible and highly configurable, with options to subset and optimize each analysis. Metapipeline-DNA facilitates high-scale, comprehensive analysis of DNA sequencing data, and is open-source under the GPLv2 license.
Type
Publication
Cell Reports Methods

Authors
Research Assistant Professor
Chenghao Zhu is a Research Assistant Professor in the NCI-designated Cancer
Center at Sanford Burnham Prebys. His research focuses on developing
computational methods and software for proteogenomics and applying
proteogenomics to cancer diagnosis, prognosis, and clinico-epidemiologic
questions.