* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
Download Supplemental Information - Molecular Cancer Research
Copy-number variation wikipedia , lookup
Gene therapy of the human retina wikipedia , lookup
Epigenetics in learning and memory wikipedia , lookup
Long non-coding RNA wikipedia , lookup
No-SCAR (Scarless Cas9 Assisted Recombineering) Genome Editing wikipedia , lookup
Non-coding DNA wikipedia , lookup
Essential gene wikipedia , lookup
Human genome wikipedia , lookup
Epigenetics of neurodegenerative diseases wikipedia , lookup
Genetic engineering wikipedia , lookup
Epigenetics of diabetes Type 2 wikipedia , lookup
Cancer epigenetics wikipedia , lookup
Metagenomics wikipedia , lookup
Gene nomenclature wikipedia , lookup
Transposable element wikipedia , lookup
Gene therapy wikipedia , lookup
Polycomb Group Proteins and Cancer wikipedia , lookup
Public health genomics wikipedia , lookup
Vectors in gene therapy wikipedia , lookup
Point mutation wikipedia , lookup
Gene desert wikipedia , lookup
Pathogenomics wikipedia , lookup
Ridge (biology) wikipedia , lookup
Biology and consumer behaviour wikipedia , lookup
Genomic imprinting wikipedia , lookup
Gene expression programming wikipedia , lookup
History of genetic engineering wikipedia , lookup
Nutriepigenomics wikipedia , lookup
Genome editing wikipedia , lookup
Therapeutic gene modulation wikipedia , lookup
Epigenetics of human development wikipedia , lookup
Minimal genome wikipedia , lookup
Helitron (biology) wikipedia , lookup
Oncogenomics wikipedia , lookup
Genome evolution wikipedia , lookup
Genome (book) wikipedia , lookup
Microevolution wikipedia , lookup
Site-specific recombinase technology wikipedia , lookup
Designer baby wikipedia , lookup
A Sleeping Beauty forward genetic screen in mice identifies Cul3 as a tumor suppressor along with 76 additional potential lung cancer driver genes. Authors Casey Dorr1,2, Callie Janik1, Madison Weg1, Raha A. Been2,8, Justin Bader5, Ryan Kang1, Brandon Ng1, Lindsey Foran1, Sean R. Landman4, M. Gerard O'Sullivan6,9, Michael Steinbach4, Aaron L. Sarver1, Kevin A. T. Silverstein7, David A. Largaespada2,3, Timothy K. Starr1,2,3 Affiliations 1 Department of Obstetrics, Gynecology & Women's Health, University of Minnesota, Minneapolis, MN 2 Masonic Cancer Center, University of Minnesota, Minneapolis, MN 3 Department of Genetic, Cell Biology & Development, University of Minnesota, Minneapolis, MN 4 Department of Computer Science and Engineering, University of Minnesota, Minneapolis, MN 5 Massachusetts Institute of Technology, Cambridge, MA 6 Department of Veterinary Population Medicine, College of Veterinary Medicine, University of Minnesota, St. Paul, MN 7 Minnesota Supercomputing Institute, University of Minnesota, Minneapolis, MN 8 Department of Comparative and Molecular Biosciences, University of Minnesota, St. Paul, MN 9 Comparative Pathology Shared Resource, Masonic Cancer Center, University of Minnesota, Minneapolis, MN Supplemental Information Supplemental Methods Immunohistochemistry & necropsy images Specificity was confirmed using negative control goat serum/immunoglobulins (BioGenex) in place of the primary antibody for CC10 and Dako's Universal Negative Control rabbit reagent (solid phase absorbed immunoglobulin fraction of serum from nonimmunized rabbits) in place of the primary antibody for proSP-C. Mouse 1188 was used in Fig. 2 Panels A-D and G. Mouse 1100 was used in Fig. 2 panels E, F and H. Mouse 1004 was used in Fig. S5. Fig. S4 mouse numbers: A) 901-1, B) 901-2, C) 1004, D) 1100-1, E) 1100-2, F) 1188, G) 1274, H) 894, I) 2461, J) 2607 and K) 1188. Transposon insertion analysis Linker-Mediated PCR. Linkers [described previously (1)] were ligated to NlaIII- (right-side) or BfaI- (left side) digested genomic DNA using T4 DNA ligase. A secondary digest (BamHI) was performed to destroy concatamer-generated products. Primary and secondary PCR was performed using primers specific for linker and SB transposon sequences along with Illumina GAIIx fusion sequences and barcode sequences (sequences available upon request). PCR amplicons were sequenced using the Illumina GAIIx platform. Sequence Analysis. Sequences were mapped to the mouse genome using BOWTIE(2) using the TAPDANCE(3) bioinformatics pipeline. TAPDANCE identifies CISs based on analysis of varying genomic window sizes, tested for significance using the Poisson distribution (p < 0.05) utilizing a Bonferroni correction based on number of windows examined. Loss- and Gain-of-Function analysis: To predict the effect of the transposon insertions for each CIS, the pattern of transposon insertions from all tumors for each CIS was manually analyzed. If the majority of transposons were in the enhancer/promoter region of gene or in a single intron and over 75% of the insertions were oriented such that the MSCV-LTR promoter would drive transcription, the predicted effect was gain-offunction. If transposons were generally equally spaced throughout the genomic loci and roughly 50% were in one orientation and 50% were in the opposite orientation, the predicted effects was loss-of-function. If neither of these patterns were apparent, the predicted effect was listed as unknown. TCGA analysis The level 2 MAF file (broad.mit.edu__Illumina_Genome_Analyzer_DNA_Sequencing_level2.maf) deposited in TCGA was downloaded from https://tcga-data.nci.nih.gov/tcga/dataAccessMatrix.html on 7/22/13. Using a custom perl script we extracted the list of mutations predicted to cause a change in protein sequence based on the Variant Classification field. The final list of mutated genes, along with the frequency of mutations was compared with the CIS human ortholog list. Significance was determined using the Fisher’s Exact Test. COSMIC analysis The COSMIC Mutant Export Including Fusions file (CosmicMutantExportIncFus_v64_270313.tsv) was downloaded from http://cancer.sanger.ac.uk/cancergenome/projects/cosmic/download on 5/5/13. Using a custom perl script we extracted the list of mutations predicted to cause a change in protein sequence along with the frequency of mutations. This list was compared with the CIS human orthologs list and significance was determined using the Fisher’s Exact Test. Cancer Gene Census Cancer Gene Census file (cancer_gene_census.tsv) was downloaded from http://cancer.sanger.ac.uk/cancergenome/projects/cosmic/download on 6/11/14. Frequent Itemset Mining We used frequent itemset mining to determine groups of genes that co-occur in multiple tumors (4,5). Specifically, closed frequent itemsets (a condensed form of frequent itemset results) were extracted from the full list of insertion locations (mapped to their nearest gene) using an apriori-based algorithm (6-8). The result of this algorithm was a list of candidate gene sets that occur in at least three different tumors (i.e. the support count). Some support counts were then modified to reflect the number of unique mice that had the gene pattern, rather than the number of tumors. This was to correct for similar gene sets in tumors originating from the same mouse. A p-value was calculated for each candidate gene set by modeling the support of the pattern as the test statistic. The null distribution was modeled as a binomial with the number of trials equal to the number of tumors and the probability of success equal to the joint probability of the individual genes in the gene set occurring together (based on their individual frequencies in the dataset). In order to account for multiple hypotheses testing, the significance of each candidate gene set was determined by empirically estimating its q-value (9), which is the minimum False Discovery Rate (FDR) at which the test may be called significant (10). Specifically, a set of 10,000 simulated results were generated by randomizing the tumor that each insertion appears in while preserving the overall set of insertion locations and the number of insertions in each tumor. The q-value for each candidate gene set was calculated as the percent of simulated results that had a p-value better than or equal to the p-value of the candidate gene set divided by the percent of real patterns with a p-value better than or equal to the candidate gene set. The top 10 gene sets by q-value (q-value < 0.8) were also the same as the top 10 gene sets by p-value (-Log(p-value) > 5), and these were considered to be the final set of significant gene sets (Table S5). Statistical analysis of overlap between NRF2 target genes and genes dysregulated in Cul3/Pten DKD A549 cells The list of genes regulated by NRF2 was identified by Chorley, et al., 2012 (11) using ChIP-Seq on cells treated with sulforaphane to induce oxidative stress. The Supplementary Data Set (worksheet: ChIP-Seq SFN) in Chorley, et al., lists 849 NRF2 binding regions, of which 243 are considered "High Confidence". The 243 High Confidence regions are associated with 232 unique gene symbols. The overlap between this list of 232 unique genes and the list of 120 genes (Table S6, this publication) differentially regulated in Cul3/Pten double knockdown A549 cells compared to controls was 7 genes (Table S7, this publication). To assess the statistical significance we used Fisher's Exact Test (12). The input numbers for the Fisher's Exact Test were as follows: Genes altered in DKD Genes not altered in DKD Genes bound by NRF2 7 225 Genes not bound by NRF2 113 22,000 Fisher's exact test computes a 2-tailed p-value = 0.000261. The gene universe used in this calculation (~22,000) is not an exact known number. However, running the test with a smaller gene universe (20,000) or a larger gene universe (30,000) still generates a significant p-value (20,000 p-value = 0.000457, 30,000 p-value = 0.000039). In the Supplementary Data Set (worksheet: Expression) Chorley, et al., also published the fold change difference between gene expression in sulforaphane treated cells and vehicle treated cells for 60 unrelated CEU HapMap lymphoblastoid cell lines. Four genes were down > 2-fold and 14 genes were up > 2-fold. Of these 18 genes, 7 overlapped with the 120 genes altered in the Cul3/Pten double knockdown cells (Table S8). We performed a Fisher's Exact Test using the following input numbers Genes Genes not altered in altered in DKD DKD Genes with > 2 fold change 7 11 Genes with < 2-fold change 113 22,000 Fisher's exact test computes a 2-tailed p-value = 3.5 e-12. Similar to the previous analysis, raising or lowering the gene universe still results in a highly significant result. References 1. 2. 3. 4. 5. 6. 7. 8. 9. 10. 11. 12. Wu X, Li Y, Crise B, Burgess SM. Transcription start regions in the human genome are favored targets for MLV integration. Science 2003;300(5626):1749-51. Langmead B, Trapnell C, Pop M, Salzberg SL. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome biology 2009;10(3):R25. Sarver AL, Erdman J, Starr T, Largaespada DA, Silverstein KA. TAPDANCE: An Automated tool to identify and annotate Transposon insertion CISs and associations between CISs from next generation sequence data. BMC Bioinformatics 2012;13(1):154. Agrawal R, Imielinski T, Swami A. Database mining: A performance perspective. Knowledge and Data Engineering, IEEE Transactions on 1993;5(6):914-25. Agrawal R, Imieliński T, Swami A. Mining association rules between sets of items in large databases. 1993. ACM. p 207-16. Zaki MJ, Ogihara M. Theoretical foundations of association rules. 1998. Citeseer. p 71-78. Pasquier N, Bastide Y, Taouil R, Lakhal L. Discovering Frequent Closed Itemsets for Association Rules. In: Beeri C, Buneman P, editors. Database Theory — ICDT’99. Volume 1540, Lecture Notes in Computer Science: Springer Berlin Heidelberg; 1999. p 398-416. Agrawal R, Srikant R. Fast Algorithms for Mining Association Rules in Large Databases. Proceedings of the 20th International Conference on Very Large Data Bases: Morgan Kaufmann Publishers Inc.; 1994. p 487-99. Subramanian A, Tamayo P, Mootha VK, Mukherjee S, Ebert BL, Gillette MA, et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci U S A 2005;102(43):15545-50. Storey JD. A direct approach to false discovery rates. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 2002;64(3):479-98. Chorley BN, Campbell MR, Wang X, Karaca M, Sambandan D, Bangura F, et al. Identification of novel NRF2-regulated genes by ChIP-Seq: influence on retinoid X receptor alpha. Nucleic Acids Res 2012;40(15):7416-29. Agresti A. A Survey of Exact Inference for Contingency Tables. Staistical Science 1992;7(1):131-53.