scieee AI-readable full text Open interactive document viewer

Driver Fusions and Their Implications in the Development and Treatment of Human Cancers

Gao, Qingsong,Liang, Wen-Wei,Foltz, Steven M,Nykter, Matti

Full text

Resource Driver Fusions and Their Implications in the Development and Treatment of Human Cancers Graphical Abstract Highlights dHighly recurrent fusions were found in prostate, bladder, breast, and lung cancers dExpression increased in oncogene fusions but decreased in tumor suppressor genes dThyroid carcinoma showed significantly higher rates of kinase fusions dTumors with fusion events tend to have lower mutational burden Authors Qingsong Gao, Wen-Wei Liang, Steven M. Foltz, ..., Matti Nykter, Ilya Shmulevich, Li Ding Correspondence [email protected] In Brief Gao et al. analyze a 9,624 sample TCGA cohort with 33 cancer types to detect gene fusion events. They provide a landscape of fusion events detected, relate fusions to gene expression, focus on kinase fusion structures, examine mutually exclusive mutation and fusion patterns, and highlight fusion druggability. Gao et al., 2018, Cell Reports 23, 227–238 April 3, 2018 ª2018 The Authors. https://doi.org/10.1016/j.celrep.2018.03.050 Cell Reports Resource Driver Fusions and Their Implications in the Development and Treatment of Human Cancers Qingsong Gao, 1,2,13 Wen-Wei Liang, 1,2,13 Steven M. Foltz, 1,2,13 Gnanavel Mutharasu, 3 Reyka G. Jayasinghe, 1,2 Song Cao, 1,2 Wen-Wei Liao, 1,2 Sheila M. Reynolds, 4 Matthew A. Wyczalkowski, 1,2 Lijun Yao, 1,2 Lihua Yu, 5 Sam Q. Sun, 1,2 The Fusion Analysis Working Group, The Cancer Genome Atlas Research Network, Ken Chen, 6 Alexander J. Lazar, 7 Ryan C. Fields, 1,8,11 Michael C. Wendl, 2,9,10 Brian A. Van Tine, 1,11 Ravi Vij, 1,11 Feng Chen, 1,11 Matti Nykter, 12 Ilya Shmulevich, 4 and Li Ding 1,2,9,11,14, * 1 Department of Medicine, Washington University in St. Louis, St. Louis, MO 63110, USA 2 McDonnell Genome Institute, Washington University in St. Louis, St. Louis, MO 63108, USA 3 Institute of Signal Processing, Tampere University of Technology, 33101, Tampere, Finland 4 Institute for Systems Biology, Seattle, WA 98109, USA 5 H3 Biomedicine, Inc., Cambridge, MA 02139, USA 6 Department of Bioinformatics and Computational Biology, The University of Texas MD Anderson Cancer Center, Houston, TX 77230, USA 7 Departments of Pathology, Genomic Medicine, and Translational Molecular Pathology, The University of Texas MD Anderson Cancer Center, Houston, TX 77230, USA 8 Department of Surgery, Washington University in St. Louis, St. Louis, MO 63110, USA 9 Department of Genetics, Washington University in St. Louis, St. Louis, MO 63110, USA 10 Department of Mathematics, Washington University in St. Louis, St. Louis, MO 63130, USA 11 Siteman Cancer Center, Washington University in St. Louis, St. Louis, MO 63110, USA 12 Institute for Biosciences and Medical Technology, University of Tampere, 33520 Tampere, Finland 13 These authors contributed equally 14 Lead Contact *Correspondence: [email protected] https://doi.org/10.1016/j.celrep.2018.03.050 SUMMARY Gene fusions represent an important class of somatic alterations in cancer. We systematically investigated fusions in 9,624 tumors across 33 cancer types using multiple fusion calling tools. We identified a total of 25,664 fusions, with a 63% validation rate. Integration of gene expression, copy number, and fusion annotation data revealed that fusions involving oncogenes tend to exhibit increased expression, whereas fusions involving tumor suppressors have the opposite effect. For fusions involving kinases, we found 1,275 with an intact kinase domain, the proportion of which varied significantly across cancer types. Our study suggests that fusions drive the development of 16.5% of cancer cases and function as the sole driver in more than 1% of them. Finally, we identified druggable fusions involving genes such as TMPRSS2,RET,FGFR3, ALK, and ESR1 in 6.0% of cases, and we predicted immunogenic peptides, suggesting that fusions may provide leads for targeted drug and immune therapy. INTRODUCTION The ability to determine the full genomic portrait of a patient is a vital prerequisite for making personalized medicine a reality. To date, many studies have focused on determining the landscape of SNPs, insertions, deletions, and copy number alterations in cancer genomes (Kanchi et al., 2014; Kandoth et al., 2013; Kumar-Sinha et al., 2015; Lawrence et al., 2014; Vogelstein et al., 2013; Wang et al., 2014). Although such genomic alterations make up a large fraction of the typical tumor mutation burden, gene fusions also play a critical role in oncogenesis. Gene fusions or translocations have the potential to create chimeric proteins with altered function. These events may also rearrange gene promoters to amplify oncogenic function through protein overexpression or to decrease the expression of tumor suppressor genes. Gene fusions function as diagnostic markers for specific cancer types. For example, a frequent translocation between chromosomes 11 and 22 creates a fusion between EWSR1 and FLI1 in Ewing’s sarcoma. Also, the Philadelphia chromosome 9–22 translocation is characteristic of chronic myeloid leukemia, resulting in the fusion protein BCR–ABL1. This fusion leads to constitutive protein tyrosine kinase activity and downstream signaling of the PI3K and MAPK pathways, which enables cells to evade apoptosis and achieve increased cell proliferation (Cilloni and Saglio, 2012; Hantschel, 2012; Ren, 2005; Sinclair et al., 2013). Fibrolamellar carcinoma (FLC) in the liver is characterized by a DNAJB1–PRKACA fusion. A recent study of The Cancer Genome Atlas (TCGA) tumors revealed this fusion transcript is specific to FLC, differentiating it from other liver cancer samples (Dinh et al., 2017). In contrast, FGFR3–TACC3 is an inframe activating kinase fusion found in multiple cancer types, including glioblastoma multiforme (GBM) (Lasorella et al., 2017; Singh et al., 2012) and urothelial bladder carcinomas (BLCA) (Cancer Cell Reports 23, 227–238, April 3, 2018 ª2018 The Authors. 227 This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/). Genome Atlas Research Network, 2014). Other recurrent fusions have also been reported in multiple cancer types (Bass et al., 2011; Jones et al., 2008; Palanisamy et al., 2010), and functional characterization of a few selected fusion genes in cellular model systems has confirmed their oncogenic nature (Lu et al., 2017). Recently, large-scale genomic studies have used the TCGA RNA sequencing (RNA-seq) data corpus to systematically identify and compile fusion candidates across many cancer types. For example, as part of its goal to develop a comprehensive, genome-wide database of fusion genes, ChimerDB (Lee et al., 2017) has analyzed RNA-seq data of several thousand TCGA cases. Giacomini et al. (2013) performed breakpoint analysis on exon microarrays across 974 cancer samples and identified 198 candidate fusions in annotated cancer genes. A searchable portal of TCGA data includes 20,731 fusions called from 9,966 cancer and 648 normal samples (Hu et al., 2018). Some studies focus on important classes of genes, such as kinase fusions (Stransky et al., 2014), which may have particular structural properties that are selected for during oncogenesis and cancer progression. However, most efforts have used only a single fusion calling algorithm. Because disagreements among different callers are common, there is a need to develop a comprehensive approach that combines the strengths of various callers to achieve higher fusion calling accuracy. Furthermore, large-scale analyses are likely to expand the targetable landscape of fusions in cancer, revealing potential treatment options for patients. Here, we leverage multiple newly developed bioinformatic tools to methodically identify fusion transcripts across the TCGA RNA-seq data corpus using the Institute for Systems Biology (ISB) Cancer Genomics Cloud. These tools include STAR-Fusion, Breakfast, and EricScript (STAR Methods). Fusion calling across 9,624 TCGA tumor samples from 33 cancer types identified a total of 25,664 fusion transcripts, with a 63.3% validation rate for the samples having available whole-genome sequencing data. Furthermore, we investigated the relationship between fusion status and gene expression, the spectrum of kinase fusions, mutations, and fusions found in driver genes, and fusions as potential drug and immunotherapy targets. RESULTS Fusion Detection Pipeline and WGS-Based Validation of a Subset of Fusion Predictions We analyzed RNA-seq data from 9,624 tumor samples and 713 normal samples from TCGA using STAR-Fusion (STAR Methods), EricScript (Benelli et al., 2012), and Breakfast (STAR Methods;Table S1). A total of 25,664 fusions were identified after extensive filtering using several panel-of-normals databases, including fusions reported in TCGA normal samples, GTEx tissues (Consortium, 2013) and non-cancer cells (Babiceanu et al., 2016)(STAR Methods;Figure 1A; Table S1). Our pipeline detected 405 of 424 events curated from individual TCGA marker papers (Table S1) (95.5% sensitivity). We further cross-confirmed our transcriptome sequencingbased fusion detection pipeline by incorporating whole-genome sequencing (WGS) data, where available. WGS paired-end reads aligned to the partner genes of each fusion were used to validate fusions detected using RNA-seq. Using all available WGS, including both low-pass and high-pass data, from 1,725 of the 9,624 cancer samples across 25 cancer types, we were able to evaluate 18.2% (4,675 fusions) of our entire fusion call set. Of that subset, WGS validated 63.3% of RNA-seq-based fusions by requiring at least three supporting discordant read pairs from the WGS data (Figure S1). Fusion Landscape across 33 Cancer Types Categorizing the 25,664 fusions on the basis of their breakpoints, we found that the majority of breakpoints are in coding regions (CDS) of both partner genes (Figure 1B). Surprisingly, there are many more fusions in 50UTRs compared with 30UTRs for both partner genes, given that 30UTRs are generally longer (MannWhitney U test, p < 2.2e-16). This could be explained by having more open chromatin in the 50UTR region (Boyle et al., 2008), the larger number of exons in 50UTRs than 30UTRs (Mann-Whitney U test, p < 2.2e-16) (Mignone et al., 2002), but could also indicate some regulatory mechanisms, such as alternative use of the promoter region of a partner gene. For different cancer types, the total number of fusions per sample varies from 0 to 60, with a median value of 1 (Figure S1). Cancer types with the fewest number of fusions per sample are kidney chromophobe (KICH), kidney renal clear cell carcinoma (KIRC), kidney renal papillary cell carcinoma (KIRP), low-grade glioma (LGG), pheochromocytoma and paraganglioma (PCPG), testicular germ cell tumors (TGCT), thyroid carcinoma (THCA), thymoma (THYM), and uveal melanoma (UVM), each with a median of 0. Other cancer types show a range of medians between 0.5 and 5 fusions per sample, although most samples demonstrate zero or only one inframe, disruptive fusion relevant to oncogenesis. Frequencies of recurrent fusions found in each cancer are illustrated in Figure 1C(Table S1). The most recurrent example within any cancer type was TMPRSS2–ERG in prostate adenocarcinoma (PRAD; 38.2%). We found FGFR3–TACC3 to be the most recurrent fusion in BLCA (2.0%), cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC, 1.7%), and lung squamous cell carcinoma (LUSC, 1.2%). Other top recurrent fusions include EML4–ALK in lung adenocarcinoma (LUAD; 1.0%), CCDC6–RET in THCA (4.2%), and FGFR2– BICC1 in cholangiocarcinoma (CHOL; 5.6%). Fusion Gene Expression in Oncogenes and Tumor Suppressors Fusion events may be associated with altered expression of one or both of the fusion gene partners, a well-known example being multiple myeloma tumors in which translocation t(4;14) fuses the highly expressed IGH locus with the tyrosine protein kinase FGFR3 (Manier et al., 2017). We integrated gene expression, copy number, and fusion annotations to systematically test for associations between gene expression and fusion status. For each fusion having an oncogene, kinase, or tumor suppressor gene (TSG) (Table S2), we determined whether that sample was an expression outlier for that gene and subsequently examined resulting percentages of both underand overexpressed genes in each cancer type (Table S3). Figure 2A shows 228 Cell Reports 23, 227–238, April 3, 2018 that between 6% (mesothelioma [MESO]) and 28% (KIRP) of kinase fusions displayed outlier overexpression of the kinase partner. Oncogenes tended to show higher likelihoods of overexpression, whereas TSGs displayed lower likelihoods. Between 3% (breast invasive carcinoma [BRCA]) and 38% (PCPG) of TSG fusions showed outlier underexpression, generally higher than both oncogenes and kinases. Figure 2B illustrates the median percentile expression levels of the most highly recurrent oncogenes and TSGs involved in fusions (Table S3). Samples with fusions involving oncogenes, such as EGFR,ERBB2, and RET, showed increased expression of those genes relative to samples without fusions across cancer types. Most TSGs showed inconsistent patterns of expression across cancer types. However, the global trend for TSGs is decreased expression compared with non-fusion samples. We also examined the relationship between TSG mutations and fusions to determine whether frequently fused TSGs were also disrupted by other mutation types. A variety of patterns AB C Figure 1. Fusion Detection and Landscape in Cancer (A) Fusion calling and filtering pipeline. (B) Cartoon overview of fusion gene partner breakpoints. Purple indicates the 50gene partner, and green indicates the 30gene partner. For both the 50and 30gene partners, fusion gene breakpoints can occur in the following genomic regions: 50UTR (triangle), coding sequence (CDS; rectangle), 30UTR (circle), and non-coding region (rounded rectangle). For each fusion event, a dotted line connects the breakpoints in the 50and 30gene partners to create the predicted fusion and the circle size, while number represents the total fusion events classified into the associated fusion category. (C) The dot plot shows the frequency of recurrent fusions found in each cancer type. The most recurrent fusion in each cancer type is labeled. Cancer types without recurrent fusions are not shown. Cell Reports 23, 227–238, April 3, 2018 229 were noted. For example, TP53 is affected by mutations rather than fusions in most cancer types. However, in sarcoma (SARC), both fusions and mutations affecting TP53 were detected. In acute myeloid leukemia (LAML), several CBFB fusions but no mutations were observed, yet other cancer types also exhibited CBFB mutations (Table S3;Figure S2). Our results suggest that alternative mechanisms are used by tumor cells in a cancer type-specific manner. AC DB Figure 2. Fusion Expression Outliers (A) The dot plot indicates the percentage of fusions called in which one of the partner genes is an expression outlier (overexpression or underexpression). The size of the dot corresponds to the number of fusions called in each cancer type. Color corresponds to genes of interest coming from lists of oncogenes, protein kinases, and tumor suppressor genes. (B) The dot plot shows the relative expression level of samples with fusions compared with those without fusions. Each sample has a particular expression percentile at a given gene, and color indicates the median percentile of samples with a fusion in that gene. Genes are the 15 most recurrent oncogenes and tumor suppressor genes. Size corresponds to the number of samples in each cancer type with a fusion at that gene. (C and D) Expression of samples at RET and CBFB in thyroid carcinoma (THCA) (C) and acute myeloid leukemia (LAML) (D), respectively. Color indicates a categorical copy number ranging from deep deletion to high amplification. 230 Cell Reports 23, 227–238, April 3, 2018 We also observed associations between fusion status and expression level in well-known fusions (Table S3), such as RET–NTRK1 in thyroid cancer, EML4–ALK in lung cancer (Stransky et al., 2014), and DNAJB1–PRKACA in the FLC subtype of liver cancer (Dinh et al., 2017). RET fusions in thyroid THCA and LUAD are inframe protein kinase fusions with overexpression of the 30RET oncogene (Figure 2C). Recurrent CBFB– MYH11 fusions in LAML are significantly associated with decreased expression of the tumor suppressor CBFB, which functions as a transcriptional regulator (Haferlach et al., 2010) (Figure 2D). In breast cancer, copy number amplification is a well-known mechanism of ERBB2 overexpression, and treatment of these AB C Figure 3. Protein Kinase Fusions (A) The bar chart indicates the number of protein kinase fusions with the kinase at the 50or 30end, inframe or frameshift, and kinase domain intact or disrupted. (B) The left bar plot shows the percentage of samples with kinase fusions across different cancer types. The number of samples with a kinase fusion is also indicated at the end of each bar. Light green and blue denote 50kinase and 30 kinase fusions, respectively. The right bar plot shows the normalized percentage of kinase fusions broken down by kinase groups. (C) The dot plot shows the numbers of samples for recurrent fusions across different cancer types. Light green and blue denote 50kinase and 30kinase fusions, respectively. HER2 + patients with trastuzumab is an established and effective targeted therapy (Smith et al., 2007). Interestingly, three of four samples with ERBB2 fusions and two samples without a called fusion showed HPV integration within 1 Mb of ERBB2 (Cao et al., 2016). ERBB2 fusion gene partners PPP1R1B and IKZF3 are genomic neighbors of ERBB2, suggesting that these fusions could be a by-product of local instability, potentially induced by the viral integration and subsequent breakage fusion events. By careful analysis of the association between fusions and expression, we have identified strategies for improving both sensitivity and specificity of fusion calls. Structure and Spectrum of Kinase Fusions Some oncogenic kinase fusions are susceptible to kinase inhibitors (Stransky et al., 2014), suggesting that additional therapeutic candidates might be discovered by examining fusion transcripts involving protein kinase genes. In total, we detected 2,892 such events, comprising 1,172 with kinase at the 30end (30-kinase), 1,603 with kinase at the 50end (50-kinase), and 117 with both partners being kinases (both-kinase) (Figure 3A; Table S4). Analysis of the catalytic kinase domains using the UniProt/PFAM domain database (STAR Methods) showed that 1,275 kinase fusions (44.1%) retained an intact kinase domain (Figure 3A). We further predicted open reading frames for these fusions and separated them into three categories with respect to the frame of the 30gene: inframe, frameshift, and no frame information (e.g., breakpoint at UTR, intron, or non-coding RNA). In general, there were more inframe fusions than frameshift fusions, especially for 30-kinase fusions, because preserving the reading frame is required to keep the kinase domain intact. For subsequent kinase analyses, we focused Cell Reports 23, 227–238, April 3, 2018 231 only on those 1,275 fusions with intact domains, further classifying the both-kinase group into 30-kinase or 50-kinase on the basis of the position of the intact domain. Comparison of kinase fusions across different cancer types indicated that kinase fusions are significantly enriched in THCA (35.6%, Fisher’s exact test, p < 2.2e-16) (Figure 3B). Moreover, the majority were 30-kinase fusions (94.0%), a significantly higher percentage than what we observed in other cancer types (Fisher’s exact test, p < 2.2e-16). We further divided these fusions into eight categories on the basis of different kinase groups, including AGC, CAMK, CK1, CMGC, STE, TK, and TKL. In general, we found that the percentages of different categories vary across cancer types (Figure 3B). For example, there are more TK fusions in THCA and GBM, more CK1 fusions in uterine corpus endometrial carcinoma (UCEC), colon adenocarcinoma (COAD), and esophageal carcinoma (ESCA) and more AGC fusions in liver hepatocellular carcinoma (LIHC). Across different cancer types, we found an enrichment of TK and TKL kinase fusions for 30-kinases but no strong preference for 50-kinases (Figure S3). Recurrent kinase fusions are of great interest as potential drug targets. Overall, we detected 744 50-kinase and 531 30-kinase fusions. Of these, 147 and 99 were recurrent, respectively, mostly across cancer types rather than within cancer types (Figure S3). As expected, fusions in the FGFR kinase family (FGFR2 and FGFR3) are the most frequent 50-kinase fusions, given their high recurrence in individual cancer types (Figure 3C). WNK kinase family fusions (WNK1 and WNK2) were also detected in multiple cancer types. The WNK family is phylogenetically distinct from the major kinase families, and there is emerging evidence of its role in cancer development (Moniz and Jordan, 2010). Here, we found a total of 23 WNK-family fusions, most of which resulted in higher expression of WNK mRNA (Figure S4). The increased expression was not generally accompanied by copy number amplification; for example, neither WNK1 nor WNK2 was amplified in ESCA or LIHC. Incidentally, ERC1–WNK1 was also detected recently in an independent Chinese esophageal cancer cohort (Chang et al., 2017). For 30-kinase fusions, all the top ten kinase genes are tyrosine kinases, most of which are enriched in THCA, including RET,BRAF,NTRK1,NTRK3,ALK, and REF1 (Figure 3C). FGR fusions were found in seven samples the same partner gene WASF2, five of which showed higher expression of FGR gene. In these five samples, the breakpoints for the two genes are the same (50UTR of both genes) resulting in usage of the stronger WASF2 promoter for the FGR gene. Interestingly, recurrent MERTK fusions are singletons in each individual cancer type with TMEM87B, and PRKACA fusions are observed only in liver cancer with DNAJB1 (Figure S3). To further understand the regulation of kinase fusions, we compared the gene expression patterns between the kinase gene and partner gene. There are in total 1,035 kinase fusions with both gene expression and copy number data available. To control for the effect of copy number amplification on gene expression, we focused on the fusions with copy numbers between 1 and 3, including 439 50-kinase and 339 30-kinase fusions (Figures 4A and 4B). For 50-kinase fusions, the kinase gene expression quantiles are uniformly distributed, indicating that the kinase gene expressions in the samples with fusion are not significantly different from the samples without fusion (Figure 4A). However, 30-kinase genes tend to show higher expression in samples with a fusion compared with the ones without. To explain this, we classified the fusion events into three categories on the basis of the relative expression pattern between the kinase gene and its partner in samples from the same cancer type. Most (66.7% [293 of 439]) 50-kinase fusions showed lower expression in the partner gene compared with the kinase. In contrast, 70.5% of 30-kinase fusions (239 of 339) showed higher partner expression (Figures 4A and 4B). Moreover, those 30-kinase fusions involving a more highly expressed 50partner also show higher kinase expression (Figure 4C). For example, we found a TRABD–DDR2 fusion in one head and neck squamous cell carcinoma (HNSC) sample, which fused the stronger TRABD promoter with DDR2, resulting in its overexpression (Figure 4D). This patient could potentially be treated using dasatinib, which targets overexpressed DDR2 in HNSC (von Massenhausen et al., 2016). DDR2 fusions were also detected in another nine samples from five different cancer types, which could be treated similarly given sufficient DDR2 overexpression (Table S1). Mutual Exclusivity between Fusions and Mutations Although mutations in oncogenes or TSGs may lead to tumorigenesis, fusions involving those genes are also an important class of cancer driver events. We systematically profiled mutations and fusions in 299 cancer driver genes (Table S2;Bailey et al., 2018)to assess the contributions of fusion genes in carcinogenesis in the 8,955 TCGA patients who overlap between the mutation call set (Key Resources Table, Public MC3 MAF; Ellrott et al., 2018)and our fusion call set. We characterized patients as having a driver mutation, a mutation in a driver gene, and/or a driver fusion (fusion involving a driver gene). Although the majority of cancer cases have known driver mutations (48.6%, mean 6.8 mutations) or mutations in driver genes (28.1%, mean 4.2 mutations), we found that 8.3% have both driver mutations and driver fusion events (mean 5.5 mutations and 1.2 fusions), 6.4% have both mutations and fusions in driver genes (mean 4.2 mutations and 1.3 fusions), and 1.8% have driver fusions only (mean 1.1 fusions) (Figure 5A). This distribution is consistent with the notion that only a few driver events are required for tumor development (Kandoth et al., 2013). We further examined the total number of mutations for samples and observed a low mutational burden in the group with driver fusion only, which is comparable with the group with no driver alterations (Figure 5B). The significant decrease in the numbers of mutations (Mann-Whitney U test, p < 2.2e-16) reflects the functionality of fusions across multiple cancer types. Moreover, within cancer types, we observed a range of 0.2% (HNSC) to 14.0% (LAML) of tumors with fusions but no driver gene mutations. Among those LAML tumors that have fusions and no driver gene mutations, we identified several well-recognized fusions relevant to leukemia, such as CBFB–MYH11 (number of samples = 3), BCR–ABL1 (n = 2), and PML–RAR (n = 2). We also identified the leukemia-initiating fusion NUP98–NSD1 in two LAML tumors (Cancer Genome Atlas Research Network et al., 2013b). We then examined the relationship of fusions and mutations in the same driver gene (Figure 5C). The result shows that when fusion events are present in a gene, mutations in the same gene are rarely found, supporting a pattern of mutual exclusivity 232 Cell Reports 23, 227–238, April 3, 2018 of the two types of genomic alteration. This trend was observed across many patients and many cancer types. Our results suggest that a considerable number of tumors are driven primarily or solely by fusion events. Contributions of Fusions to Cancer Treatment We investigated potentially druggable fusion events in our call set using our curated Database of Evidence for Precision Oncology (DEPO; Sun et al., unpublished data) (Table S5). We defined a fusion as druggable if there is literature supporting the use of a drug against that fusion, regardless of cancer type (allowing for ‘‘off-label’’ drug treatment). We found potentially druggable fusions across 29 cancer types, with major recurrent druggable targets in PRAD (TMPRSS2, 205 samples), THCA (RET, 33 samples), and LAML (PML–RARA, 16 samples) (Figure 6A). FGFR3 was a potential target (both on-label and AC D B Figure 4. Kinase Gene Expression Regulated by Fusion (A) The scatterplot shows the gene expression quantile (y axis) for the 50-kinase without copy number variation (between one and three copies; x axis). All genes are classified among three categories: kinase expression higher, equal, and lower, compared with partner expression, marked in blue, gray, and red, respectively. The density plot for expression quantile is also shown on the right. (B) The scatterplot shows the gene expression quantile (y axis) for the 30-kinase without copy number variation (between one and three copies; x axis). The colors represent the same three categories as (A). The density plot for expression quantile is also shown. (C) Boxplot comparing the distribution of kinase gene expression quantile between the three groups defined in (A) for 50-kinase and 30-kinase, respectively. (D) Schematic of TBABD–DDR2 fusion gene structure in an HNSC sample and scatterplot of DDR2 copy number versus mRNA expression in HNSC. The samples with and without this fusion are marked in red and blue, respectively. Cell Reports 23, 227–238, April 3, 2018 233 off-label) in 15 cancer types. Overall, we found 6.0% of samples (574 of 9,624 samples) to be potentially druggable by one or more fusion targeted treatments. Further study of fusions in human cancer will facilitate the development of precision cancer treatments. We analyzed patterns of fusion druggability in LUAD, stratifying by smoking status. In this dataset, 15% of LUAD samples (75 of 500 samples with known smoking status) were from never smokers, while a significantly higher percentage of never smokers (15 of 75 samples) versus smokers (9 of 425 samples) were found to have druggable fusion (chi-square test, p < 1e-6) (Figure 6B). Several Food and Drug Administration (FDA)- approved drugs exist to target ALK fusions in lung and other cancer types. We observed ALK fusions in 20 samples from eight cancer types (5 samples in LUAD). In most cases, fusion status corresponded to copy number neutral overexpression of ALK (Figure 6D). In 17 of 20 cases, ALK was the 30partner of the fusion pair, with EML4 being the most frequent 50partner (7 of 17). ESR1 encodes an estrogen receptor with important and druggable relevance to breast cancer (Li et al., 2013). We detected ESR1 fusions in 16 samples from five different cancer types (9 samples from BRCA). Of the 9 BRCA samples, 8 are known be from the luminal A or B subtype. We observed strict mutual exclusivity between ESR1 mutations and fusions (Figure 5C). Of the 16 fusions, 11 have ESR1 at the 50end and 5 at the 30 end. When ESR1 is the 50gene in the fusion, the transactivation (AF1) domain is always included (Figure 6D). When ESR1 is the 30gene, the transactivation (AF2) domain is always included. Those samples with ESR1 fusion tend of have higher ESR1 expression, especially in the 9 BRCA samples (Figure S5). Similarly, ESR1 expression is higher when ESR1 is mutated in BRCA, CESC, and UCEC, which are all hormone receptor-related cancer types (Cancer Genome Atlas, 2012; Cancer Genome Atlas Research Network et al., 2013a, 2017). Further functional study to determine the mechanism of ESR1 fusions could suggest drug development directions. AB C Figure 5. Mutual Exclusivity between Driver Mutations and Driver Fusions (A) The bar plot shows the percentages of samples with driver mutations only (green), mutations only (orange), driver mutation and fusion (blue), mutation and fusion (pink), or fusion only (light green) events in 299 cancer driver genes. (B) Distribution of mutation burden across each alteration group designated in all figures. (C) All samples with fusions or mutations in any of the genes indicated on the left are displayed on the x axis. For each gene, samples are clustered by the alteration group. Bottom bar indicates cancer type. 234 Cell Reports 23, 227–238, April 3, 2018 DEPO DEPO is a curated list of druggable variants filtered such that each variant corresponds to one of several categories: single nucleotide polymorphisms or SNPs (missense, frameshift, and nonsense mutations), inframe insertions and deletions (indels), copy number variations (CNVs) or expression changes. Each variant/drug entry in DEPO was paired with several annotations of potential interest to oncologists. DEPO is available as a web portal (http://dinglab.wustl.edu/depo). Cell Reports 23, 227–238.e1–e3, April 3, 2018 e3