Download Supplemental Information - Molecular Cancer Research

Survey
yes no Was this document useful for you?
   Thank you for your participation!

* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project

Document related concepts

Copy-number variation wikipedia , lookup

Gene therapy of the human retina wikipedia , lookup

Epigenetics in learning and memory wikipedia , lookup

Long non-coding RNA wikipedia , lookup

Genomics wikipedia , lookup

No-SCAR (Scarless Cas9 Assisted Recombineering) Genome Editing wikipedia , lookup

Non-coding DNA wikipedia , lookup

Essential gene wikipedia , lookup

Human genome wikipedia , lookup

Epigenetics of neurodegenerative diseases wikipedia , lookup

Genetic engineering wikipedia , lookup

Epigenetics of diabetes Type 2 wikipedia , lookup

Cancer epigenetics wikipedia , lookup

Metagenomics wikipedia , lookup

Gene nomenclature wikipedia , lookup

Transposable element wikipedia , lookup

Gene therapy wikipedia , lookup

Polycomb Group Proteins and Cancer wikipedia , lookup

Public health genomics wikipedia , lookup

Vectors in gene therapy wikipedia , lookup

Point mutation wikipedia , lookup

Gene desert wikipedia , lookup

NEDD9 wikipedia , lookup

Pathogenomics wikipedia , lookup

Ridge (biology) wikipedia , lookup

Biology and consumer behaviour wikipedia , lookup

Genomic imprinting wikipedia , lookup

Gene expression programming wikipedia , lookup

History of genetic engineering wikipedia , lookup

Nutriepigenomics wikipedia , lookup

Genome editing wikipedia , lookup

Therapeutic gene modulation wikipedia , lookup

Epigenetics of human development wikipedia , lookup

Gene wikipedia , lookup

Minimal genome wikipedia , lookup

Helitron (biology) wikipedia , lookup

Oncogenomics wikipedia , lookup

Genome evolution wikipedia , lookup

Genome (book) wikipedia , lookup

RNA-Seq wikipedia , lookup

Microevolution wikipedia , lookup

Site-specific recombinase technology wikipedia , lookup

Designer baby wikipedia , lookup

Artificial gene synthesis wikipedia , lookup

Gene expression profiling wikipedia , lookup

Transcript
A Sleeping Beauty forward genetic screen in mice identifies Cul3 as a
tumor suppressor along with 76 additional potential lung cancer driver
genes.
Authors
Casey Dorr1,2, Callie Janik1, Madison Weg1, Raha A. Been2,8, Justin Bader5, Ryan Kang1, Brandon Ng1,
Lindsey Foran1, Sean R. Landman4, M. Gerard O'Sullivan6,9, Michael Steinbach4, Aaron L. Sarver1,
Kevin A. T. Silverstein7, David A. Largaespada2,3, Timothy K. Starr1,2,3
Affiliations
1
Department of Obstetrics, Gynecology & Women's Health, University of Minnesota, Minneapolis, MN
2
Masonic Cancer Center, University of Minnesota, Minneapolis, MN
3
Department of Genetic, Cell Biology & Development, University of Minnesota, Minneapolis, MN
4
Department of Computer Science and Engineering, University of Minnesota, Minneapolis, MN
5
Massachusetts Institute of Technology, Cambridge, MA
6
Department of Veterinary Population Medicine, College of Veterinary Medicine, University of Minnesota, St.
Paul, MN
7
Minnesota Supercomputing Institute, University of Minnesota, Minneapolis, MN
8
Department of Comparative and Molecular Biosciences, University of Minnesota, St. Paul, MN
9
Comparative Pathology Shared Resource, Masonic Cancer Center, University of Minnesota, Minneapolis, MN
Supplemental Information
Supplemental Methods
Immunohistochemistry & necropsy images
Specificity was confirmed using negative control goat serum/immunoglobulins (BioGenex) in place of the
primary antibody for CC10 and Dako's Universal Negative Control rabbit reagent (solid phase absorbed
immunoglobulin fraction of serum from nonimmunized rabbits) in place of the primary antibody for proSP-C.
Mouse 1188 was used in Fig. 2 Panels A-D and G. Mouse 1100 was used in Fig. 2 panels E, F and H. Mouse
1004 was used in Fig. S5. Fig. S4 mouse numbers: A) 901-1, B) 901-2, C) 1004, D) 1100-1, E) 1100-2, F)
1188, G) 1274, H) 894, I) 2461, J) 2607 and K) 1188.
Transposon insertion analysis
Linker-Mediated PCR. Linkers [described previously (1)] were ligated to NlaIII- (right-side) or BfaI- (left
side) digested genomic DNA using T4 DNA ligase. A secondary digest (BamHI) was performed to destroy
concatamer-generated products. Primary and secondary PCR was performed using primers specific for linker
and SB transposon sequences along with Illumina GAIIx fusion sequences and barcode sequences (sequences
available upon request). PCR amplicons were sequenced using the Illumina GAIIx platform.
Sequence Analysis. Sequences were mapped to the mouse genome using BOWTIE(2) using the
TAPDANCE(3) bioinformatics pipeline. TAPDANCE identifies CISs based on analysis of varying genomic
window sizes, tested for significance using the Poisson distribution (p < 0.05) utilizing a Bonferroni correction
based on number of windows examined.
Loss- and Gain-of-Function analysis: To predict the effect of the transposon insertions for each CIS, the
pattern of transposon insertions from all tumors for each CIS was manually analyzed. If the majority of
transposons were in the enhancer/promoter region of gene or in a single intron and over 75% of the insertions
were oriented such that the MSCV-LTR promoter would drive transcription, the predicted effect was gain-offunction. If transposons were generally equally spaced throughout the genomic loci and roughly 50% were in
one orientation and 50% were in the opposite orientation, the predicted effects was loss-of-function. If neither
of these patterns were apparent, the predicted effect was listed as unknown.
TCGA analysis
The level 2 MAF file (broad.mit.edu__Illumina_Genome_Analyzer_DNA_Sequencing_level2.maf) deposited
in TCGA was downloaded from https://tcga-data.nci.nih.gov/tcga/dataAccessMatrix.html on 7/22/13. Using a
custom perl script we extracted the list of mutations predicted to cause a change in protein sequence based on
the Variant Classification field. The final list of mutated genes, along with the frequency of mutations was
compared with the CIS human ortholog list. Significance was determined using the Fisher’s Exact Test.
COSMIC analysis
The COSMIC Mutant Export Including Fusions file (CosmicMutantExportIncFus_v64_270313.tsv) was
downloaded from http://cancer.sanger.ac.uk/cancergenome/projects/cosmic/download on 5/5/13. Using a
custom perl script we extracted the list of mutations predicted to cause a change in protein sequence along with
the frequency of mutations. This list was compared with the CIS human orthologs list and significance was
determined using the Fisher’s Exact Test.
Cancer Gene Census
Cancer Gene Census file (cancer_gene_census.tsv) was downloaded from
http://cancer.sanger.ac.uk/cancergenome/projects/cosmic/download on 6/11/14.
Frequent Itemset Mining
We used frequent itemset mining to determine groups of genes that co-occur in multiple tumors (4,5).
Specifically, closed frequent itemsets (a condensed form of frequent itemset results) were extracted from the
full list of insertion locations (mapped to their nearest gene) using an apriori-based algorithm (6-8). The result
of this algorithm was a list of candidate gene sets that occur in at least three different tumors (i.e. the support
count). Some support counts were then modified to reflect the number of unique mice that had the gene pattern,
rather than the number of tumors. This was to correct for similar gene sets in tumors originating from the same
mouse.
A p-value was calculated for each candidate gene set by modeling the support of the pattern as the test statistic.
The null distribution was modeled as a binomial with the number of trials equal to the number of tumors and the
probability of success equal to the joint probability of the individual genes in the gene set occurring together
(based on their individual frequencies in the dataset). In order to account for multiple hypotheses testing, the
significance of each candidate gene set was determined by empirically estimating its q-value (9), which is the
minimum False Discovery Rate (FDR) at which the test may be called significant (10). Specifically, a set of
10,000 simulated results were generated by randomizing the tumor that each insertion appears in while
preserving the overall set of insertion locations and the number of insertions in each tumor. The q-value for
each candidate gene set was calculated as the percent of simulated results that had a p-value better than or equal
to the p-value of the candidate gene set divided by the percent of real patterns with a p-value better than or
equal to the candidate gene set. The top 10 gene sets by q-value (q-value < 0.8) were also the same as the top 10
gene sets by p-value (-Log(p-value) > 5), and these were considered to be the final set of significant gene sets
(Table S5).
Statistical analysis of overlap between NRF2 target genes and genes
dysregulated in Cul3/Pten DKD A549 cells
The list of genes regulated by NRF2 was identified by Chorley, et al., 2012 (11) using ChIP-Seq on cells treated
with sulforaphane to induce oxidative stress. The Supplementary Data Set (worksheet: ChIP-Seq SFN) in
Chorley, et al., lists 849 NRF2 binding regions, of which 243 are considered "High Confidence". The 243 High
Confidence regions are associated with 232 unique gene symbols. The overlap between this list of 232 unique
genes and the list of 120 genes (Table S6, this publication) differentially regulated in Cul3/Pten double
knockdown A549 cells compared to controls was 7 genes (Table S7, this publication). To assess the statistical
significance we used Fisher's Exact Test (12). The input numbers for the Fisher's Exact Test were as follows:
Genes
altered in
DKD
Genes not
altered in DKD
Genes bound by
NRF2
7
225
Genes not bound by
NRF2
113
22,000
Fisher's exact test computes a 2-tailed p-value = 0.000261. The gene universe used in this calculation (~22,000)
is not an exact known number. However, running the test with a smaller gene universe (20,000) or a larger gene
universe (30,000) still generates a significant p-value (20,000 p-value = 0.000457, 30,000 p-value = 0.000039).
In the Supplementary Data Set (worksheet: Expression) Chorley, et al., also published the fold change
difference between gene expression in sulforaphane treated cells and vehicle treated cells for 60 unrelated CEU
HapMap lymphoblastoid cell lines. Four genes were down > 2-fold and 14 genes were up > 2-fold. Of these 18
genes, 7 overlapped with the 120 genes altered in the Cul3/Pten double knockdown cells (Table S8). We
performed a Fisher's Exact Test using the following input numbers
Genes
Genes not
altered in
altered in DKD
DKD
Genes with > 2 fold
change
7
11
Genes with < 2-fold
change
113
22,000
Fisher's exact test computes a 2-tailed p-value = 3.5 e-12. Similar to the previous analysis, raising or lowering
the gene universe still results in a highly significant result.
References
1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
11.
12.
Wu X, Li Y, Crise B, Burgess SM. Transcription start regions in the human genome are favored targets
for MLV integration. Science 2003;300(5626):1749-51.
Langmead B, Trapnell C, Pop M, Salzberg SL. Ultrafast and memory-efficient alignment of short DNA
sequences to the human genome. Genome biology 2009;10(3):R25.
Sarver AL, Erdman J, Starr T, Largaespada DA, Silverstein KA. TAPDANCE: An Automated tool to
identify and annotate Transposon insertion CISs and associations between CISs from next generation
sequence data. BMC Bioinformatics 2012;13(1):154.
Agrawal R, Imielinski T, Swami A. Database mining: A performance perspective. Knowledge and Data
Engineering, IEEE Transactions on 1993;5(6):914-25.
Agrawal R, Imieliński T, Swami A. Mining association rules between sets of items in large databases.
1993. ACM. p 207-16.
Zaki MJ, Ogihara M. Theoretical foundations of association rules. 1998. Citeseer. p 71-78.
Pasquier N, Bastide Y, Taouil R, Lakhal L. Discovering Frequent Closed Itemsets for Association
Rules. In: Beeri C, Buneman P, editors. Database Theory — ICDT’99. Volume 1540, Lecture Notes in
Computer Science: Springer Berlin Heidelberg; 1999. p 398-416.
Agrawal R, Srikant R. Fast Algorithms for Mining Association Rules in Large Databases. Proceedings
of the 20th International Conference on Very Large Data Bases: Morgan Kaufmann Publishers Inc.;
1994. p 487-99.
Subramanian A, Tamayo P, Mootha VK, Mukherjee S, Ebert BL, Gillette MA, et al. Gene set
enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles.
Proc Natl Acad Sci U S A 2005;102(43):15545-50.
Storey JD. A direct approach to false discovery rates. Journal of the Royal Statistical Society: Series B
(Statistical Methodology) 2002;64(3):479-98.
Chorley BN, Campbell MR, Wang X, Karaca M, Sambandan D, Bangura F, et al. Identification of novel
NRF2-regulated genes by ChIP-Seq: influence on retinoid X receptor alpha. Nucleic Acids Res
2012;40(15):7416-29.
Agresti A. A Survey of Exact Inference for Contingency Tables. Staistical Science 1992;7(1):131-53.