RegTools: Integrated analysis of genomic and transcriptomic data for the discovery of splice-associated variants in cancer
Citations Over Time
Abstract
Abstract Somatic mutations within non-coding regions and even exons may have unidentified regulatory consequences that are often overlooked in analysis workflows. Here we present RegTools ( www.regtools.org ), a computationally efficient, free, and open-source software package designed to integrate somatic variants from genomic data with splice junctions from bulk or single cell transcriptomic data to identify variants that may cause aberrant splicing. RegTools was applied to over 9,000 tumor samples with both tumor DNA and RNA sequence data. We discovered 235,778 events where a splice-associated variant significantly increased the splicing of a particular junction, across 158,200 unique variants and 131,212 unique junctions. To characterize these somatic variants and their associated splice isoforms, we annotated them with the Variant Effect Predictor (VEP), SpliceAI, and Genotype-Tissue Expression (GTEx) junction counts and compared our results to other tools that integrate genomic and transcriptomic data. While many events were corroborated by the aforementioned tools, the flexibility of RegTools also allowed us to identify novel splice-associated variants and previously unreported patterns of splicing disruption in known cancer drivers, such as TP53, CDKN2A , and B2M , as well as in genes not previously considered cancer-relevant.
Related Papers
- → EST comparison indicates 38% of human mRNAs contain possible alternative splice forms(2000)301 cited
- → Molecular cloning and tissue distribution of a short form chicken leptin receptor mRNA(2006)55 cited
- → Alternative Splice Variants Encoding Unstable Protein Domains Exist in the Human Brain(2004)19 cited
- → TAPAS: tools to assist the targeted protein quantification of human alternative splice variants(2014)1 cited
- → Comparison of splice sites in mammals and chicken(2004)94 cited