Survey
* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
THE HUMAN GENOME VARIATION SOCIETY: POSSIBLE FUTURE DIRECTIONS? Bruce Gottlieb Lady Davis Institute for Medical Research, Sir Mortimer B. Davis-Jewish General Hospital, 3755 Chemin de la Cote-Ste-Catherine, Department of Human Genetics, McGill University, Montreal, Quebec, H3T 1E2, Canada, and Department of Biology, John Abbott College, P.O. Box 2000, Ste Anne de Bellevue, Quebec, H9X 3L9, Canada The Human Genome Variation Society and its precursor the Mutation Database Initiative of HUGO have been in existence for over 10 years. In that time the organization has evolved into its present form, with a somewhat limited membership, whose activities have principally involved organizing two major meetings a year. The organization has in addition become the sponsor of a major scientific publication - Human Mutation, as well as on occasion supporting the efforts of some of it’s members to obtain funding for a group grant from the NIH. Initially, its major interest was to serve as a forum for individuals that had set up mutation databases. However, with the ‘sequencing’ of the human genome, interest in mutation databases has markedly waned and with it in HGVS as well. In fact, it might have been expected that interest in human genome variation would have increased following the sequencing of the human genome, but this seems not to have happened. This has meant that the society has experienced little growth over the past few years. The challenge now for the society is how to reinvigorate itself by both expanding its membership and perhaps more importantly its influence in the scientific community. This paper will present a number of suggestions and ideas, which it is hoped will not only stimulate discussion, but perhaps also lead to both a reinvigoration of its membership, as well as a possible new purpose and direction for HGVS itself. LSDB-IN-A-BOX: STATUS REPORT Alastair F. Brown, Keith Brown, Ewan McDowall MRC Human Genetics Unit, Crewe Road, Edinburgh EH4 2XU, U.K. We are developing a system that integrates a Locus Specific Database (LSDB) with a Patient Database, Sample Database, and Phenotype Database. For security reasons, these databases are “stand-alone” in their own right, but can be linked to form an integrated system. This also has the advantage that, given certain constraints, alternative database designs could be linked into the system. We are currently using the Patient Database locally, and are carrying out final tests and modifications before releasing the software for general use in the near future. Development of the Sample Database is ongoing, with the intention that details of sample processing and tracking will be included, as well as basic data on sample content and storage location. We plan to integrate the existing LSDB software developed by Johan den Dunnen into the system, and this stage of the process will be started next. The final stage will be integration of a phenotype database. This has proved the most difficult part of the system to design in a generic way, and no final decision on a design has yet been taken. As always, comments and suggestions are welcome. Details of the current status can be found on the project web site at http://lsdb.hgu.mrc.ac.uk/ LOVD AN LSDB-IN-A-BOX FACILITATING SIMPLE CREATION OF A SEQUENCE VARIATION DATABASE USING OPEN SOURCE SOFTWARE Ivo F.A.C. Fokkema, Johan T. den Dunnen, and Peter E.M. Taschner Center of Human and Clinical Genetics, Department of Human Genetics, Leiden University Medical Center, Nederland The completion of the human genome project has provided the basis for the collection and study of sequence variations in any human gene. Efficient access to sequence variation information is currently provided most conveniently through web-based databases, so called LSDB's (Locus-Specific DataBases). Ideally, to guarantee the most efficient access to the data for any gene, these databases would have to be set up in such a way that similar data for a group of genes might be retrieved using standardized queries. We have developed an LSDB-in-a-Box, the Leiden Open Source Variation Database (LOVD) software, using freely available open source software only (PHP and MySQL). LOVD is fully web-based, platform-independent and creates LSDB's using a format in accordance with the recommendations of the Human Genome Variation Society (HGVS). Upon installation, LOVD generates gene-centered LSDB's, including several levels of access (e.g. database manager, curator, submitter) and web pages for data submission, curation and visualization of the data collected. The basic design is directed at cataloguing DNA variation but facilitates extension with phenotype and patient data with minimal effort. The LOVD software has been used to establish the databases displayed at the Leiden Muscular Dystrophy pages (http://www.DMD.nl) and is freely available (http://www.humgen.nl/LOVD/). ANALYSIS OF GENETIC VARIATION IN CONSERVED NON-CODING REGIONS OF THE HUMAN GENOME Gregory Kryukov, Steffen Schmidt and Shamil Sunyaev Genetics Division, Brigham & Women’s Hospital and Harvard Medical School Overwhelming majority of known mutations with large phenotypic effect, including mutations causing Mendelian human diseases correspond to changes in protein coding genes. Effect of these mutations can be experimentally tested at the molecular level, and computational methods for predicting the effect of amino acid changes have been developed. However, recent data on several mammalian genomes showed that many non-coding regions exhibit sequence conservation comparable to that of protein coding genes. The high sequence conservation suggests that many mutations in non-coding regions are under pressure of negative selection, and, therefore, should have phenotypic effect. Functional importance of DNA variation in conserved non-coding regions and its relation to phenotypic variation remains to be an open question. It was hypothesized that polymorphic variants in non-coding DNA contribute to human complex diseases. Although it is not currently possible to evaluate the effect of non-coding mutations on function experimentally, it is possible to estimate this effect via analysis of statistical signatures of natural selection. This will help evaluate importance of non-coding genetic variation and assess the potential of comparative genomics to highlight DNA variants of large phenotypic effect. We analyzed nucleotide diversity and allele frequency spectrum of human SNPs in genomic regions highly conserved between primate and rodent genomes. We also compared substitution rate in the human lineage after divergence from chimpanzee to the substitution rate in the mouse lineage after divergence from rat. The well established theoretical relationship between the strength of purifying natural selection and substitution rate predicts that the relative substitution rate should be lower in large (mouse) than in small (human) population. Although, the relative substitution rate was indeed higher in the human lineage than in the mouse lineage, this effect was much more profound in non-coding regions compared to protein-coding genes. Non-coding regions have higher polymorphism density and smaller fraction of rare alleles compared to equally conserved coding regions. Based on these data we estimated fraction of sites with large effect on fitness and, consequently, on phenotypes, in both coding and non-coding regions of the genome. This analysis suggests that most mutations in conserved non-coding regions are only slightly deleterious, in sharp contrast with protein coding regions, even though the conservation level of these regions is the same. We propose that conserved non-coding DNA have relatively few sites which, if mutated, result in large phenotypic effects. However, conserved non-coding regions harbor a large number of slightly deleterious SNPs, and their cumulative effect on fitness and phenotypes may be substantial. FROM LOCUS SPECIFIC DATABASE TO PHYLOGENY Béroud C, Hamroun D, Desgeorge M, Guittard C and Claustres M. Laboratoire de Montpellier, France. Génétique Moléculaire, Almost 50 years after the discovery of the DNA’s double helix by James Watson and Francis Crick, the International Human Genome Sequencing Consortium announced the successful completion of the Human Genome Project. Concomitantly to this sequencing effort, many biotechnology companies have produced new tools to rapidly scan large sets of samples for mutations. So every year thousands of variations are thus identified in diagnostic and research laboratories. The knowledge of these variations associated with clinical and biological data is essential for clinicians, geneticists and researchers. If in most cases the knowledge of the disease causing mutations is sufficient, additional information are usually available such as polymorphisms or unclassified variations. These data are today wasted for the community as they are not available through Locus Specific DataBases (LSDBs). It is now time to collect these data and to move from mutations to haplotypes. For this purpose, we have implemented new features in the UMD® software. The user can collect a virtually unlimited number of polymorphisms and mutations for a specific patient. He can specify if these variations are localized on the same chromosome (cis) or on different ones (trans). The UMD® software can thus build haplotypes for all samples. Specific routines have been designed to produce association tables between variations (allele-allele associations), variations and haplotypes (allele-haplotype associations) and haplotypes (haplotype-haplotype associations). These tools are useful to identify association disequilibrium, to analyze complex alleles and to identify modifier effects. Here, we report the creation of the UMD-CFTRCBAVD database, which includes mutations of the CFTR gene specifically associated with congenital bilateral absence of the vas deferens (CBAVD). Today, this database contains 242 mutations and 1543 polymorphisms from 159 individuals and six mutations account for 70.7% of mutations. Surprisingly, if 3 haplotypes are frequently found (39%), 62 among the 90 different haplotypes have been reported only once. This observation led us to develop a unique module for LSDBs: Phylogeny. This tool creates a PHYLIP formatted file of the various haplotypes. It then runs the dnamlk program from the PHYLIP package. This program implements the maximum likelihood method for DNA sequences under the constraint that the trees estimated must be consistent with a molecular clock. At this step, the UMD® software collects the output file and draws the estimated tree. To easily interpret this tree, additional color codes are drawn for each variation. Thus, we have shown that, in the context of CBAVD, the c.1327G>T (p.Asp443Tyr) mutation is specifically found in haplotypes harboring the TG10 and 7T alleles from intron 8 in addition to c.1727G>C (p.Gly576Ala) and c.2002C>T (p.Arg668Cys) polymorphisms. Similarly, the c.1522_1524delTTT (p.Phe508del), accounting for 32% of mutations, is specifically found on chromosomes harboring the TG10 and 9T alleles from intron 8. These results show that the collection of polymorphism information resulting from largescale sequencing should be reported in LSDBs. This will open a new area and contribute to the identification of highly conserved ancestral chromosomal segments. INTERPRETING MISSENSE VARIANTS IN OCULOCUTANEOUS ALBINISM GENES TYROSINASE AND P GENE: AN EVOLUTIONARY APPROACH M.S. Greenblatt, S. Duraisamy, C. McBride, J.P. Bond, W.S. Oetting Univ of Vermont, Univ of Minnesota BACKGROUND: Oculocutaneous Albinism (OCA) is characterized by lack of skin and eye pigment, UV sensitivity, predisposition to skin cancer, and developmental eye defects. OCA is caused by mutations in several genes, most commonly the tyrosinase (OCA1) and P genes (OCA2). Tyrosinase is the key enzyme catalyzing melanin pigment synthesis from tyrosine; the P gene encodes a transport protein. Understanding their intragenic conservation patterns can help to predict functionally critical amino acids (AA), where variants are likely to cause OCA. OBJECTIVES: To assess quantitatively the predictive value of AA conservation we have: 1) made sequence alignments and phylogenetic trees of the tyrosinase and P genes, 2) computationally studied their intragenic AA conservation patterns, and 3) tested how well the patterns predict OCA-associated variants and non-OCA-associated polymorphisms, using the Albinism database (http://albinismdb.med.umn.edu). RESULTS: Evolutionary variation in the existing database of sequences is sufficient to determine statistically significant conservation of codons for tyrosinase but not for P gene. The SIFT program correctly predicted 82 % of the OCA1 associated tyrosinase variants and both of the non-OCAassociated polymorphisms, and 94% of the OCA2-associated P gene variants. However, 4 of 6 P gene polymorphisms were incorrectly predicted to be deleterious based on conservation in the three known P gene sequences (human, pig, mouse). To achieve 82% sensitivity for predicting OCA1-associated tyrosinase variants, cutoff scores for AA conservation would classify 50% of codons as conserved, with resulting 50% specificity, 100% Positive Predictive Value (PPV), and only 9% Negative Predictive Value (NPV). For predicting OCA2-associated P gene variants, cutoff scores to achieve 85% sensitivity would classify 73% of codons as conserved, with resulting 50% specificity, 91% PPV, and 38% NPV. CONCLUSIONS: 1) >80% prediction of deleterious mutations is possible. 2) Adequate databases of sequences, mutations, and polymorphisms are needed to confirm the validity of predictions based on evolutionary AA conservation. Since few polymorphisms have been reported in these genes, the power of this technique is reduced. 3) More P Gene sequences are needed to validate evolutionary predictions. COSMIC (CATALOGUE OF SOMATIC MUTATIONS IN CANCER) - A RESOURCE FOR CANCER MUTATIONS Richard Wooster, P Andrew Futreal, Michael R. STratton Cancer Genome Project, Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridgeshire, CB101SA, UK COSMIC is a resource for cancer genetics; a database of published mutations involved in human cancer (http://www.sanger.ac.uk/cosmic). Focusing on genes of key interest in carcinogenesis, the project is analysing an expanding gene set, starting with a group of 17 including BRAF, FGFR3, HRAS, RET, PTEN. As of October 2004, data have been collated from more than 2,800 journal publications, comprising over 114,000 tumours analysed for tumourigenic mutations. 17,929 of these samples had mutations, comprising 1,728 unique sequence alterations. COSMIC collates all key genomic data describing a mutation event in a tumour sample, including the experimental procedures used, and the tissue type / histological classification of the tumour sample itself. Web pages overlay this data, providing highly selectable graphical views and tabulated summaries of the information together with links to external sources such as PubMed, Swissprot and Pfam. COSMIC is also being used as a vehicle to display data from the systematic mutation screens being performed by the Cancer Genome Project. AUTOMATED SPLICE SITE MUTATION ANALYSIS BY INFORMATION THEORY Vijay K. Nalla & Peter K. Rogan. Children’s Mercy Hospital, Schools of Medicine and Computer Science & Engineering, University of Missouri-Kansas City, USA Significance: Accurate interpretation of mutations that alter non-coding, conserved sequence elements in human genes is important for diagnosis and prognosis of inherited or acquired genetic disorders. The effects of such mutations can be predicted in silico by information theory (Hum. Mut. 6:74-76; Hum Mut. 12: 153-171). This is because changes in the affinity of a protein or protein complex for its cognate binding site can be estimated from the individual information content of the sequence of a site, which is based on an information weight matrix describing a set of sites recognized by the same protein(s) (J. Theor. Biol. 189: 427441). Implementation: Perl and C++ software was developed to assist in interpretation of noncoding sequence variation in functional elements within human genes. Initially, algorithms were implemented to analyze mutations expressed in standard HUGO nomenclature from a wide variety of peer-reviewed sources, including published literature and conforming locus specific databases. The software is also able to process some non-standard mutation designations used in popular locus specific databases. Genomic coordinates homologous to the mRNA accession number or gene are parsed from a MySQL database, then the reference and corresponding variant or haplotype sequences are verified and retrieved. The software introduces mutation(s) into the reference sequence and computes the resulting information contents (and changes, if any) at splice sites and/or regulatory sequences (human and murine donor and acceptor sites, SC35, ASF/SF2, and SRp40 regulatory binding sites). These tools were incorporated into a secure web server that detects changes in information content at binding sites due to mutations in any catalogued human gene, genome-mapped mRNA or user-defined sequence [https://splice.cmh.edu]. Individual information analyses with splice junction binding site matrices detects activated cryptic splice sites, associated splicing regulatory sites and distinguishes null alleles from those that are partially functional. Standard gene and mutation nomenclature (HUGO-approved format) are entered using a CGI-based front end supported by backend server. Changes in information content are tabulated and visualized as walker figures showing the binding sites on the sequence. Performance: Depending upon the set of options that are selected, the Web server requires 30 ~ 60” to analyze one mutation under optimal CPU loads. Batch submission of 50 mutations has been clocked at 2’ 25” on a single 2.4 MHz I686 processor running Redhat Linux. Results: This application has been validated by analysis of ~600 mutations (including haplotypes) parsed directly or interactively from Human Mutation and locusspecific databases. We confirmed that all of the previous recognized splicing mutations affected splice site strength or activated cryptic splice sites. Information analysis identified 8 examples of missense mutations that concomitantly affected adjacent splice donor and acceptor sites, 4 partially functional splice sites previously thought to be null alleles, and 6 unrecognized cryptic splice site mutations. Alterations in presumed SR protein binding sites that may impact known splicing regulatory elements were also detected in a number of cases. Conclusions: The system has been designed to facilitate addition of other information weight matrices and genomes and software for scanning multipartite binding sites in future versions. It should be feasible to develop more comprehensive models of splicing phenotypes and eventually, to analyze binding sites in other regulatory sequences. Acknowledgements: Grant support [PHS ES10855] from the NIEHS is acknowledged. MUTATIONVIEW : DEVELOPMENT OF AN ENHANCED SEARCH SYSTEM FOR THE CLINICAL WORDS IN OMIM WITH A NEW ALGORITHM 1,2) 2) Shinsei Minoshima , Masafumi Ohtsubo , 3) 3) Katsue Daicho , Kouichi Kawaguchi , Susumu 2) 2) Mitsuyama , Takashi Kawamura , Tomoyoshi 3) 2) Horisawa , Nobuyoshi Shimizu 1) Photon Med. Res. Ctr., Hamamatsu Univ. Sch. 2) Med., Dept. Mol. Biol., Keio Univ. Sch. Med., 3) Chi Co., Ltd. In order to investigate the relevance between disease and genetic diversity, we have developed an integrated knowledge-base system, MutationView (http://mutview.dmb.med. keio.ac.jp/). Current MutationView has principally focused on monogenic disease mutations, and its characteristic features are as follows: (1) Various data display: genomic/cDNA structure, functional domain of protein, histogram for the case number of mutations, changes in the nucleotide/amino acid sequence and restriction sites with graphical environment, (2) Analysis functions: classification based on the various information included in each case record (e.g. ethnic origin, onset age and symptom). To date, we have collected 10,166 entries of mutations from 1736 literatures dealing with 259 genes involved in 407 distinct diseases according to particular categories such as eye, brain, muscle, ear, heart, autoimmunity and familial tumor. Recently, the systemic bone disease was added as a new category. Moreover, we have developed a genome browser, which has genome-wide data presentation function. V arious information such as chromosome band, contig, gene, transcription, and DNA marker covering all the chromosomes was automatically imported from Ensembl (http://www.ensembl.org/). Recently, we have developed a new search system for OMIM. Scientifically significant words were picked up using an existing dictionary, and then sets of meaningful interrelation between them were computationally extracted from the OMIM with a statistical analysis based on coincidental appearance of the words in the same section of each OMIM document. These interrelation are quite useful for the search based on word association. Thus, we are developing a new algorithm of an ab initio method for extraction of important words without dictionaries, in which “Variance” value of each word in OMIM is utilized. The system will be demonstrated at the meeting. PHENOTYPES VS GENOTYPES IN THE WORLD OF BLOOD GROUP ANTIGENS 1 2 Olga O. Blumenfeld and Santosh K. Patnaik , 1 2 Deparments of Biochemistry and Cell Biology , Albert Einstein College of Medicine, New York, NY 10461 Blood group antigens are proteins, glycans or glycolipids, of a variety of functions, whose common feature is that all are expressed on the surface of red cells and are polymorphic in the population. The hallmark of each antigen is an epitope, a linear or spatially arranged amino acid or carbohydrate sequence which due to its variant nature can be recognized as non-self by the immune system. The science (art) of serology is based on this recognition, and its goal is to decipher and assign blood group phenotypes using antibodies to the polymorphic epitopes as tools. A blood group system is a set of variant antigens encoded by alleles of a single locus, each defining a related form of a common blood group phenotype. The Blood Group Antigen Gene Mutation Database* documents 36 genes encoding 29 blood group systems and comprising close to 700 alleles that result in surface expression of at least 400 different blood group phenotypes. Here we examine the correlation between the genotype, the structure of the allele and the blood group phenotype. This is a rare example in which a direct correlation between a DNA alteration and a single physiologic function (antibody response) can be established, with modifier genes or environmental factors playing a minimal role. In the database, DNA alterations are documented in donors who were selected for study on the basis of a variant blood group phenotype. Thus an alteration of the epitopic and/or other segments of the coding regions is expected. As generally observed for many documented human sequence variations, among total DNA alterations, single nucleotide mutations predominate (sense ~8%, missense ~50%, nonsense ~6%). Nonsense mutations and small deletions, insertions, or splice site alterations, often accompanied by additional upstream or downstream mutations, give rise to a number of related alleles whose products are truncated or defective and result, directly or indirectly (e.g., in case of inactive glycosyltransferases), in the absence of epitopes from the cell surface. Thus, for several blood group systems such different alleles result in the null phenotype (e.g., the O phenotype of the ABO system). Gene rearrangements based on different mechanisms (gene conversions, unequal recombinations) can give rise to the same epitopic sequence and the same phenotype specified by different hybrid alleles (e.g., the MN system). In contrast, in most instances, missense mutations linked to the epitope give rise to variant phenotypes, each characteristic of a specific mutation and the amino acid change. This is observed whether the protein exhibits a single or multiple epitopes; the effect of the mutation can be direct or indirect (affecting the level of cell surface expression). The majority of DNA alterations have no apparent effect on the function of the erythrocyte. The subtle relationships among a single amino acid replacement affecting the epitopic region, the immune response, and the ability to detect it, become apparent from a survey of the different blood group systems documented in the database. *http://www.bioc.aecom.yu.edu/bgmut/index.htm DEVELOPMENT OF A PUTATIVE FUNCTIONAL CODING SNP SET FOR WHOLE-GENOME DIRECT ASSOCIATION STUDIES Francisco M. De La Vega, Charles R. Scafe, Anish Kejariwal, Eugene G. Spier, Paul D. Thomas, and Dennis A. Gilbert. Applied Biosystems, Foster City, CA, USA Genetic association studies of complex disease can be carried out by: indirect association, in which surrogate markers assumed to be in linkage disequilibrium (LD) with the disease allele are tested for trait association; or direct association, in which a list of putatively functional SNPs are tested for their disease relevance directly. Current estimates of the number of markers needed for a genome scan via LD in humans range from 120,000 to over a million SNPs, translating into enormous genotyping costs and a challenging problem of statistical inference. On the other hand, a genome scan with putative functional SNPs may require typing tens of thousands of common causative variants implicated in complex disease. The feasibility of whole-genome direct association studies requires that most of the variants influencing the susceptibility to disease are typed in the study. Currently, about 40,000 non-synonymous coding SNPs (nsSNPs) are deposited on the public databases. An additional 30,000 novel nsSNPs were discovered through the resequencing of exonic regions of 23,363 genes by the Applera Genomics Initiative. Combined, these datasets provide a comprehensive resource of nsSNPs making the development of a whole-genome direct association marker set feasible. We are currently pursuing implementation of such a set on the SNPlex™ Genotyping System, a multiplexed high throughput genotyping platform based on the oligonucleotide ligation/PCR assay. SNPs can be prioritized as to their assay conversion potential, heterozygosity, location in the gene, gene’s product classification, and their putative functional impact. In our preliminary design, we compiled 28,709 nsSNPs including over 9,000 proprietary SNPs based on their measured or expected heterozygosity in populations of European and African descent. SNPs were grouped when possible by protein families based on the PANTHER protein classification before submission to the SNPlex assay design pipeline. We designed SNPlex assays for about 70% of the SNPs. Currently we are engaged in improving the assay design conversion rate and in collaborative studies for the validation of the set in complex disease cohorts. RESEQUENCING ON A CHIP: IS IT READY FOR PRIME TIME? Arupa Ganguly1, Courtney MacMullen2, Charles A. Stanley2 1Department of Genetics, University of Pennsylvania School of Medicine; 2Division of Endocrinology, The Children’s Hospital of Philadelphia Resequencing of multiple large genes in parallel can be achieved by using high density oligonucleotide microarrays. We summarize our experience of resequencing using Custom-Seq chips synthesized by Affymetrix, CA in the context of Congenital Hyperinsulinism (HI). HI, the most common cause of persistent hypoglycemia in infants, is caused by mutations in at least 4 different genes: the KATP potassium channel comprised of the sulfonylurea receptor (SUR1) and inward rectifying agent (Kir6.2), glutamate dehydrogenase (GDH), and glucokinase (GK). We designed the HI-Chip for detecting variations in the 65 coding exons including flanking intronic regions of the 4 HI genes. We evaluated the resequencing method using genomic DNA from 13 individuals: 1 negative control, 2 positive controls, and 10 unknowns, and compared the results to conventional automated fluorescent single pass direct sequencing on an ABI 3100 machine. Initial base call rates averaged 92% (9393bp/10210bp) with a range of 61 to 96%. This translates to an average of 12 ‘n’ calls per tiled exon. However, the settings of the GDAS analysis software can be adjusted to improve the call rate at a cost of accuracy. Under modified condition, the average base call rate improved to 96.49%(9851bp/10210bp) with a range of 84.18- 98.45%. This resulted in an average of 6 ‘n’ calls per exon. Resequencing on a chip has the potential to return HI genotypes in a time efficient manner. The method has to be optimized to perform as well as direct sequencing for any custom set of genes. About 60% of the amplicons have to be directly sequenced after scanning due to presence of undefined ‘n’ calls. The sequence context of these n’s is predominantly a stretch of Cs (>3 in a row). Visual inspection of the probe hybridization data indicate the possibility of better base calls if the algorithm could be trained to look at the data from one strand only. In time, the HI-CHIP may prove to be a clinically useful pre-operative diagnostic tool and guide the extent of pancreatectomy in non-medically responsive HI cases. ARRAY-BASED MLPA ANALYSIS 1,2 1 1 2 M.E. Kalf , S.J.White , M.Kriek , L. Vahlkamp , 2 1 R. van Beuningen , M.H. Breuning , J.T. den 1 1 Dunnen . Leiden University Medical Center, 2 Leiden, Nederland, PamGene International B.V., 's Hertogenbosch, Nederland Due to its simplicity and broad applicability, Multiplex Ligation-dependent Probe Amplification (MLPA) is becoming the method of choice for detecting deletions and duplications in genomic DNA. A drawback of the current methodology is that probes are separated according to length, limiting the number of loci that can be amplified and analysed simultaneously. The use of probes of uniform length can be expected to facilitate the simultaneous amplification of hundreds of probes in one multiplex PCR. Furthermore, analysis of these products using micro-arrays should allow the read-out of potentially thousands of loci at a resolution currently unattainable with array CGH. To test this approach we have quantitatively scored MLPA-products using a porous microTM array substrate (PamChip arrays). Compared with planar arrays, these arrays have several advantages, including a larger surface area, the possibility to vary hybridization stringency during analysis and, most interestingly, a decreased hybridization time of about 10 minutes. We combined four commercially available probe sets (product range 130 bp to 490 bp) with our own synthetically produced probes (product range 80 bp to 125 bp), allowing 180 loci to be analysed simultaneously. Analysis of samples with known mutations gave results that were comparable with those derived using capillary electrophoresis. It was noticeable that in this complex amplification longer probes were amplified with a reduced and variable efficiency, resulting in lower signals and less reliable scores. This suggests that using probes of identical lengths should improve consistency of the results. A reduction of the hybridization time prior to probe ligation from 16 hours to 2.5 hours did not affect MLPA-accuracy. This indicates that a complete MLPA-analysis, including DNA isolation, can be obtained within 8 hours. HIGHLY SENSITIVE, EFFICIENT AND RAPID MUTATION IDENTIFICATION FOR GENES WITH A HIGH FREQUENCY OF NOVEL MUTATIONS ENHANCES QUALITY OF CARE AND REDUCES COSTS Nadia Prigoda, Katherine Zhang, Kirk Vandezande, Diane Rushlow, Beata Piovesan, Ning Chen and Brenda L. Gallie Solutions by Sequence, Retinoblastoma Solutions and HHT Solutions, University Health Network, University of Toronto, Toronto, Canada We describe a strategy for mutation identification in genes harboring a wide spectrum of disease-causing mutations. This multi-assay strategy is optimized for each disease gene studied, detects a wide variety of mutations, and achieves high test sensitivity and cost-efficiency in a relatively short turnaround time. We have optimized the methods for RB1 gene mutations that lead to retinoblastoma and for ALK-1 and ENG gene mutations that lead to hereditary hemorrhagic telangiectasia (HHT). Quantitative Multiplex PCR (QM-PCR) efficiently detects changes in the size or copy number of exons, approximately 35% of mutated RB1 alleles. Half (18%) of the RB1 mutations detected by QM-PCR are small insertions and deletions; half are whole exon deletions or duplications not detected by sequencing. A single allele-specific PCR (ASPCR) reaction detects eight commonly recurring RB1 mutations, 21% of the mutations identified. We optimized sequencing reactions by combining two exons in each reaction with different labels and ordering duplex reactions to maximize the mutation discovery rate. Sequencing alone detects approximately 68% of mutated RB1 alleles. Methylation-specific PCR identifi es RB1 promoter hypermethylation in 13% of unilateral retinoblastoma tumors. If all of these DNA -based methods fail to identify the mutated allele, we use RNA-based methods to search for deep intronic splice mutations. Overall, these methods enabled us to identify the germline mutation in 92.4% of 314 blood samples from individuals with bilateral or familial unilateral retinoblastoma, and both somatic mutations in 89% of the 210 tumor samples from individuals with unilateral sporadic retinoblastoma. Median turnaround time in 2003 was 4.6 weeks. This strategy for RB1 mutation identification simultaneously achieved significant saving in health costs, improved the standard of care, and improved the clinical outcomes for most of the retinoblastoma families studied. Using a similar technical strategy combining QM-PCR and duplex sequencing, we have detected ALK -1 and endoglin (ENG) mutations in approximately 85% of patients clinically diagnosed with HHT. This suggests that our highly sensitive, rapid and efficient approach to mutation detection can be successfully extended to other similar genes. We advocate the use of our hierarchical mutation identification strategy to screen other large complex genes with a wide spectrum of disease-causing mutations, such as BRCA 1/2. THE EFFECT OF CODING SYNONYMOUS AND 3’ UTR SNPS ON MRNA SECONDARY STRUCTURE Huiqi Qu and Constantin Polychronakos, Endocrine Genetics Laboratory, Department of Pediatrics, McGill University Health Center, Montreal, Québec, Canada Background In the search for functional features of the human genome altered by singlenucleotide polymorphisms (SNPs), mRNA secondary structure has not received much attention. However, mRNA secondary structure differences caused by SNPs have a significant potential to influence biologically important functions by changing mRNA stability, translational efficency or other, less well understood RNA functions. Since consensus sequences identifying functionally important RNA regions are not nearly as well understood as the corresponding DNA features, a nonspecific but potentially powerful screening approach would use effects on secondary structure to identify priority SNPs for functional evaluation from among a list of candidates—for example in a search for the functional variant within a tight linkage disequilibrium block associated with a complex disease. In this study, we explored the feasibility of predicting SNP effects on mRNA secondary structure computationally, by minimum free energy change, compared effects by different nucleotide substitutions and made a first attempt at correlating this to experimentally verified function. Methods Using the NCBI dbSNP database (http://www.ncbi.nlm.nih.gov/SNP/), synonymous SNPs (sSNPs) and 3’ UTR SNPs were selected for this study. The human dbSNP database build 122 includes 7893 sSNPs and 52388 UTR SNPs from 22 autosomal chromosomes with heterozygosity>0.10. From these SNPs, 100 sSNPs and 100 3’ UTR SNPs were selected randomly. Because of the uncertain delineation of the 5’ UTRs in many genes, SNP’s in this region were not included. This decision was made before any results were known. Minimum free energy (MFE) of whole length mRNA was computed on the basis of an energy minimization algorithm (Zuker and Stiegler 1981), by the Vienna RNA Package (Hofacker 2003, http://rna.tbi.univie.ac.at/cgibin/RNAfold.cgi). Because of a limitation of this algorithm, only SNPs located in mRNA <4000 nucleotides in length were included. The change of minimium free energy of mRNA moleculars was compared among different nucleotide substitutions and between sSNPs and 3’ UTR SNPs. Six synonymous or 3' UTR SNPs with experimentally demonstrated functional effects were found in an exhaustive search of Pub Med. MFE changes by these SNPs were compared to thos e of the randomly selected SNPs. Results Between sSNPs and 3’ UTR SNPs, there was no difference as to the distribution of each type of nucleotide substitution or MFE changes, either by type of substitution or as a whole. Therefore for subsequent analysis the two types were pooled. Among those 200 SNPs, the most common nucleodite substitution was C/T, twice as frequent as its complementary A/G (54.5% vs. 26.5%, p=0.000017). One way ANOVA suggested statistically significant differences in MEF change among different substitutions (F=2.53, p=0.03). The difference was further confirmed in an additional group of 240 SNPs, randomly selected so that each type of substitution was equally represented (F=4.3, p=0.001). G/T substitutions made up 4% of the total and had the highest average MFE change at 2.02 ± 0.24 kcal/mol (mean+SEM). The most common C/T was found to have the lowest average MFE changes at 1.03± 0.084. The mean of MFE changes by the six sSNPs or 3’ UTR SNPs with experimental evidence of functional effect were more than twice that of random SNPs (2.60+0.82 vs. 1.23+ 1.22 kcal/mol). Conclusion This study suggests that computational prediction of mRNA secondary structure change by MFE assessment can predict experimentally verified functional effects of a SNP. This justifies further exploration of its potential to identify SNPs for further in-depth experimental analysis in search of functional variants. Hofacker, I. L. (2003). "Vienna RNA secondary structure server." Nucleic Acids Res 31(13): 3429-31. Zuker, M. and P. Stiegler (1981). "Optimal computer folding of large RNA sequences using thermodynamics and auxiliary Nucleic Acids Res 9(1): 133-48. information." PMSG, A NOVEL TECHNOLOGY TO ANALYZE DNA AND RNA SEQUENCES FOR VARIATION 1 1 1 Dan Graziano , Chaof u Shi , Angela Alexander , 2 1 Jon Jarvik and Cheryl Telmer 1 SpectraGenetics LLC, 4415 Fifth Ave., Suite 160, Pittsburgh PA, 15213 2 Department of Biological Sciences, Carnegie Mellon University, 4400 Fifth Ave., Pittsburgh PA, 15213 Peptide mass signature genotyping, PMSG, is a scanning genotyping method that detects and characterizes known and novel mutations and polymorphisms by (1) amplifying the sequences of interest by PCR, (2) translating the amplicons in more than one reading frame, (3) affinity purifying the resulting peptides, and (4) analyzing the peptides by mass spectrometry. The set of peptide masses encoded in a given sequence comprises a “peptide mass signature” characteristic of that sequence. Individual sequence variants typically yield different and distinct mass signatures because they change the amino acid composition, and hence the mass, of one or more of the peptides of which the signature is comprised. Once a given signature has been detected and verified by dideoxy sequencing, the signature serves to identify that sequence variant in subsequent analyses. We have applied PMSG technology to several genes including the tumor suppressor geneTP53. The TP53 test analyzes exons 2 to 11 of the gene and all splice sites in a multiplexed configuration at a cost comparable to dideoxy sequencing but with greater sensitivity. A PMSG test that detects mutations and splice variants in mRNA is also under development. The status of the DNA and RNA tests will be discussed, and the advantages and limitations of PMSG will be addressed with respect to detecting and characterizing germline and somatic sequence variation. Telmer, C.A., Retchless, A.R., Kinsey, A.D., Conley, Y., Rigatti, B., Gorin, M.B., Jarvik, J.W. 2003. Detection and assignment of mutations and minihaplotypes in human DNA using Peptide Mass Signature Genotyping (PMSG): application to the human RDS/Peripherin gene. Genome Research 13: 1944-1951. Telmer, C.A., An, J., Malehorn, D.E., Zeng, X, Gollin, S.D., Ishwad, C., Jarvik, J.W. 2003. Detection and assignment of TP53 mutations in tumor DNA using Peptide Mass Signature Genotyping. Human Mutation 22: 158-165. DEVELOPMENT AND VALIDATION OF ECONOMIC, HIGH THROUGHPUT MELTMADGE ASSAYS FOR THE BRCA1 CODING REGION Aldahmesh M.A, Spanakis E, Alharbi K.K, Sillibourne J, *Day I.N.M, Eccles D.M Human Genetics Division, School of Medicine, Duthie Building (MP 808), Southampton University Hospitals NHS Trust, Tremona Road, Southampton SO16 6YD, UK * Author for correspondence. Email:[email protected] Telephone +44 1703 795063 (Sec); fax +44 1703 794264 MeltMADGE offers 30-100 fold economy and throughput advantages for mutation scanning over standard prevalent techniques such as DHPLC, SSCP and CSGE, but has not previously been evaluated in any diagnostically relevant genes. In this study we have developed 54 PCR-meltMADGE assays representing the BRCA1 coding region, and part of the 3'noncoding region. Ten known SNPs were detected and also different pathogenic mutations previously characterised by other methods were detected in blind trials on a panel of 94 unrelated subjects. In addition, a new SNP in the 3'noncoding region and two new sequence variants were also identified. However, the same system can run an assay on > 1,000 subjects simultaneously with the main cost increment being that of the PCR reactions. This approach has reasonable sensitivity to base changes and the potential to be applied to complete description of population missense mutation diversity (“reference range” studies) and to initial economic scans of diagnostic lab backlogs and of larger lower risk groups. AN EVALUATION OF THE ACCURACY AND PRECISION OF TWO DNA POOLING STUDY DESIGNS BY COMPARISON OF POOLING RESULTS WITH INDIVIDUAL GENOTYPE DATA FOR 18,000 SNPS INTERACTION AND ASSOCIATION EFFECTS OF MYOC, OPTN AND APOE IN PATIENTS WITH PRIMARY OPEN ANGLE GLAUCOMA 1 1 1 1 CP Pang, BJ Fan, DY Wang, POS Tam, YF 1,2 1 Leung, DSC Lam 1 1 2 2 Ansar Jawaid , Kelly Frazer , David Hinds and 1 Neil J Gibson 1. Research and Development Genetics, AstraZeneca Pharmaceuticals, Alderley Park, Macclesfield, SK10 4TG, United Kingdom. 2. Perlegen Sciences, Inc. 2021 Stierlin Court, Mountain View, CA 94043-4655, USA Pooling of individual DNA samples prior to genotyping has been proposed as a practical method of performing whole genome association studies with very large numbers of SNP markers 1 . DNA pooling can significantly reduce the consumable and labour costs of a study as well as reducing DNA usage. However, experimental errors inherent in a pooling design lead to a loss in power compared to individual genotyping. Furthermore, such errors increase the false positive rate, if not appropriately quantified and considered in the analyses, leading to significant 2 difficulty in exploiting the results . Experimental errors arise from not obtaining exactly equal amounts of DNA from each individual included in the pool, variation in the quantitative accuracy of the method to measure allele frequencies, and differential amplification of alleles. Here we examine the sources of experimental error using empirical data from large pool and small pool designs with corresponding individual genotype counts for 18,000 SNPs typed in 685 samples using a high density array -based genotyping 3 platform. References 1. Norton N, Williams NM, O'Donovan MC, Owen MJ. Ann Med. 2004; 36(2): 146-52 2. Zou G, Zhao H. Genet Epidemiol. 2004 Jan;26(1):1-10. 3. Patil N, Berno AJ, Hinds DA et al. Science. 2001 Nov 23;294(5547):1719-23. Department of Ophthalmology & Visual Sciences, the Chinese University of Hong Kong, 2 Hong Kong. Present affiliation: Bauer Center for Genomics Research, Harvard University, Cambridge, USA. Primary open-angle glaucoma (POAG) is a leading cause of visual impairment and blindness worldwide. Myocilin (MYOC) and optineurin (OPTN) are the two known diseasecausing genes for POAG. Apolipoprotein E (APOE) is a potential candidate gene. To explore the interactions among these genes in POAG, we performed a multi-gene association study in 200 sporadic POAG patients and 201 unrelated control subjects. We identified disease-causing mutations (DCMs) in the MYOC gene, R91X, E300K and Y471C, accounting for 1.5% of POAG patients. DCMs identified in the OPTN gene were E103D and H486R, found in 1% of POAG patients. Two non-coding OPTN variants, IVS6-5T>C and IVS6-10G>A, and the APOE e4 allele, decreased POAG risk. The OPTN IVS6-5T>C reduced POAG risk by 5- fold based on multivariable analysis (P = 0.014) and haplotype analysis (P = 0.0003). Two APOE promoter polymorphisms, -491A>T and 219T>G, increased POAG susceptibility. Three pairs of MYOC and OPTN polymorphisms, 1000C>G (MYOC) and M98K (O PTN), A260A (MYOC) and R545Q (OPTN), I288I (MYOC) and IVS7+24G>A (OPTN), showed linkage disequilibrium in POAG, indicating interactions between the MYOC and OPTN genes. Meanwhile, R545Q in OPTN decreased cup-disc ratio (P = 0.004), and IVS8+20G>A and IVS1548C>A lowered IOP at diagnosis (P = 0.006 and 0.030 respectively). Our findings showed that besides the DCMs, non-coding sequence changes in MYOC and OPTN might also affect the susceptibility of POAG. While our data did not reveal modifying effects of APOE on the glaucoma phenotype, we have shown potential interactions between MYOC and OPTN to be involved in the pathogenesis of POAG, indicating a digenic etiology for sporadic POAG.