Scholay

学术搜索 · AI 审稿 · LaTeX 协作

The Open Pediatric Cancer Project

作者:Zhuangzhuang Geng, Eric Wafula, Ryan J. Corbett, Yuanchao Zhang, Run Jin, Krutika S. Gaonkar, Sangeeta Shukla, Komal S. Rathi, Dave Hill, Aditya Lahiri, Daniel P. Miller, Alex Sickler, Kelsey Keith, Christopher Blackden, Antonia Chroni, Miguel Brown, Adam Kraya, Kaylyn L. Clark, Brian R. Rood, Adam Resnick, Nicholas Van Kuren, John M. Maris, Alvin Farrel, Mateusz P. Koptyra, Gerri Trooskin, Noel Coleman, Yuankun Zhu, Stephanie Stefankiewicz, Zied Abdullaev, Asif Chinwalla, Mariarita Santi, Ammar S. Naqvi, Jennifer L. Mason, Carl Koschmann, Xiaoyan Huang, Sharon J. Diskin, Kenneth Aldape, Bailey Farrow, Weiping Ma, Bo Zhang, Brian Ennis, Sarah K. Tasian, Saksham Phul, Matthew R. Lueder, Chuwei Zhong, Joseph M. Dybas, Pei Wang, Deanne Taylor, Jo Lynne Rokita · 发表于:bioRxiv (Cold Spring Harbor Laboratory) · 年份:2024 · DOI:10.1101/2024.07.09.599086 · 被引用次数:8 · 研究领域:Glioma Diagnosis and Treatment、Neuroblastoma Research and Treatments、Genomics and Rare Diseases

Background: In 2019, the Open Pediatric Brain Tumor Atlas (OpenPBTA) was created as a global, collaborative open-science initiative to genomically characterize 1,074 pediatric brain tumors and 22 patient-derived cell lines. Here, we present an extension of the OpenPBTA called the Open Pediatric Cancer (OpenPedCan) Project, a harmonized open-source multi-omic dataset from 6,112 pediatric cancer patients with 7,096 tumor events across more than 100 histologies. Combined with RNA-Seq from the Genotype-Tissue Expression (GTEx) and The Cancer Genome Atlas (TCGA), OpenPedCan contains nearly 48,000 total biospecimens (24,002 tumor and 23,893 normal specimens). Findings: We utilized Gabriella Miller Kids First (GMKF) workflows to harmonize WGS, WXS, RNA-seq, and Targeted Sequencing datasets to include somatic SNVs, InDels, CNVs, SVs, RNA expression, fusions, and splice variants. We integrated summarized CPTAC whole cell proteomics and phospho-proteomics data, miRNA-Seq data, and have developed a methylation array harmonization workflow to include m-values, beta-vales, and copy number calls. OpenPedCan contains reproducible, dockerized workflows in GitHub, CAVATICA, and Amazon Web Services (AWS) to deliver harmonized and processed data from over 60 scalable modules which can be leveraged both locally and on AWS. The processed data are released in a versioned manner and accessible through CAVATICA or AWS S3 download (from GitHub), and queryable through PedcBioPortal and the NCI's pedia...