Download THE HUMAN GENOME VARIATION SOCIETY: POSSIBLE FUTURE

Survey
yes no Was this document useful for you?
   Thank you for your participation!

* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project

Document related concepts
no text concepts found
Transcript
THE HUMAN GENOME VARIATION
SOCIETY: POSSIBLE FUTURE
DIRECTIONS?
Bruce Gottlieb
Lady Davis Institute for Medical Research, Sir
Mortimer B. Davis-Jewish General Hospital,
3755 Chemin de la Cote-Ste-Catherine,
Department of Human Genetics, McGill
University, Montreal, Quebec, H3T 1E2,
Canada, and Department of Biology, John
Abbott College, P.O. Box 2000, Ste Anne de
Bellevue, Quebec, H9X 3L9, Canada
The Human Genome Variation Society and its
precursor the Mutation Database Initiative of
HUGO have been in existence for over 10 years.
In that time the organization has evolved into its
present form, with a somewhat limited
membership, whose activities have principally
involved organizing two major meetings a year.
The organization has in addition become the
sponsor of a major scientific publication - Human
Mutation, as well as on occasion supporting the
efforts of some of it’s members to obtain funding
for a group grant from the NIH.
Initially, its major interest was to serve as a
forum for individuals that had set up mutation
databases. However, with the ‘sequencing’ of
the human genome, interest in mutation
databases has markedly waned and with it in
HGVS as well. In fact, it might have been
expected that interest in human genome
variation would have increased following the
sequencing of the human genome, but this
seems not to have happened. This has meant
that the society has experienced little growth
over the past few years.
The challenge now for the society is how to
reinvigorate itself by both expanding its
membership and perhaps more importantly its
influence in the scientific community.
This paper will present a number of suggestions
and ideas, which it is hoped will not only
stimulate discussion, but perhaps also lead to
both a reinvigoration of its membership, as well
as a possible new purpose and direction for
HGVS itself.
LSDB-IN-A-BOX: STATUS REPORT
Alastair F. Brown, Keith Brown, Ewan McDowall
MRC Human Genetics Unit, Crewe Road,
Edinburgh EH4 2XU, U.K.
We are developing a system that integrates a
Locus Specific Database (LSDB) with a Patient
Database, Sample Database, and Phenotype
Database. For security reasons, these
databases are “stand-alone” in their own right,
but can be linked to form an integrated system.
This also has the advantage that, given certain
constraints, alternative database designs could
be linked into the system.
We are currently using the Patient Database
locally, and are carrying out final tests and
modifications before releasing the software for
general use in the near future. Development of
the Sample Database is ongoing, with the
intention that details of sample processing and
tracking will be included, as well as basic data
on sample content and storage location.
We plan to integrate the existing LSDB software
developed by Johan den Dunnen into the
system, and this stage of the process will be
started next.
The final stage will be integration of a phenotype
database. This has proved the most difficult part
of the system to design in a generic way, and no
final decision on a design has yet been taken.
As always, comments and suggestions are
welcome. Details of the current status can be
found
on
the
project
web
site
at
http://lsdb.hgu.mrc.ac.uk/
LOVD AN LSDB-IN-A-BOX
FACILITATING SIMPLE CREATION
OF A SEQUENCE VARIATION
DATABASE USING OPEN SOURCE
SOFTWARE
Ivo F.A.C. Fokkema, Johan T. den Dunnen, and
Peter E.M. Taschner
Center of Human and Clinical Genetics,
Department of Human Genetics, Leiden
University Medical Center, Nederland
The completion of the human genome project
has provided the basis for the collection and
study of sequence variations in any human
gene. Efficient access to sequence variation
information
is
currently
provided
most
conveniently through web-based databases, so
called LSDB's (Locus-Specific DataBases).
Ideally, to guarantee the most efficient access to
the data for any gene, these databases would
have to be set up in such a way that similar data
for a group of genes might be retrieved using
standardized queries. We have developed an
LSDB-in-a-Box, the Leiden Open Source
Variation Database (LOVD) software, using
freely available open source software only (PHP
and MySQL). LOVD is fully web-based,
platform-independent and creates LSDB's using
a
format
in
accordance
with
the
recommendations of the Human Genome
Variation Society (HGVS). Upon installation,
LOVD
generates
gene-centered
LSDB's,
including several levels of access (e.g. database
manager, curator, submitter) and web pages for
data submission, curation and visualization of
the data collected. The basic design is directed
at cataloguing DNA variation but facilitates
extension with phenotype and patient data with
minimal effort. The LOVD software has been
used to establish the databases displayed at the
Leiden
Muscular
Dystrophy
pages
(http://www.DMD.nl) and is freely available
(http://www.humgen.nl/LOVD/).
ANALYSIS OF GENETIC VARIATION
IN CONSERVED NON-CODING
REGIONS OF THE HUMAN GENOME
Gregory Kryukov, Steffen Schmidt and Shamil
Sunyaev
Genetics Division, Brigham & Women’s Hospital
and Harvard Medical School
Overwhelming majority of known mutations with
large phenotypic effect, including mutations
causing Mendelian human diseases correspond
to changes in protein coding genes. Effect of
these mutations can be experimentally tested at
the molecular level, and computational methods
for predicting the effect of amino acid changes
have been developed. However, recent data on
several mammalian genomes showed that many
non-coding
regions
exhibit
sequence
conservation comparable to that of protein
coding genes. The high sequence conservation
suggests that many mutations in non-coding
regions are under pressure of negative
selection,
and,
therefore,
should
have
phenotypic effect.
Functional importance of
DNA variation in conserved non-coding regions
and its relation to phenotypic variation remains
to be an open question. It was hypothesized
that polymorphic variants in non-coding DNA
contribute to human complex diseases.
Although it is not currently possible to evaluate
the effect of non-coding mutations on function
experimentally, it is possible to estimate this
effect via analysis of statistical signatures of
natural selection.
This will help evaluate
importance of non-coding genetic variation and
assess the potential of comparative genomics to
highlight DNA variants of large phenotypic
effect.
We analyzed nucleotide diversity and allele
frequency spectrum of human SNPs in genomic
regions highly conserved between primate and
rodent genomes.
We also compared
substitution rate in the human lineage after
divergence from chimpanzee to the substitution
rate in the mouse lineage after divergence from
rat. The well established theoretical relationship
between the strength of purifying natural
selection and substitution rate predicts that the
relative substitution rate should be lower in large
(mouse) than in small (human) population.
Although, the relative substitution rate was
indeed higher in the human lineage than in the
mouse lineage, this effect was much more
profound in non-coding regions compared to
protein-coding genes. Non-coding regions have
higher polymorphism density and smaller
fraction of rare alleles compared to equally
conserved coding regions. Based on these data
we estimated fraction of sites with large effect on
fitness and, consequently, on phenotypes, in
both coding and non-coding regions of the
genome.
This analysis suggests that most
mutations in conserved non-coding regions are
only slightly deleterious, in sharp contrast with
protein coding regions, even though the
conservation level of these regions is the same.
We propose that conserved non-coding DNA
have relatively few sites which, if mutated, result
in large phenotypic effects. However, conserved
non-coding regions harbor a large number of
slightly deleterious SNPs, and their cumulative
effect on fitness and phenotypes may be
substantial.
FROM LOCUS SPECIFIC DATABASE
TO PHYLOGENY
Béroud C, Hamroun D, Desgeorge M, Guittard C
and Claustres M.
Laboratoire
de
Montpellier, France.
Génétique
Moléculaire,
Almost 50 years after the discovery of the DNA’s
double helix by James Watson and Francis
Crick, the International Human Genome
Sequencing
Consortium
announced
the
successful completion of the Human Genome
Project. Concomitantly to this sequencing effort,
many biotechnology companies have produced
new tools to rapidly scan large sets of samples
for mutations. So every year thousands of
variations are thus identified in diagnostic and
research laboratories. The knowledge of these
variations associated with clinical and biological
data is essential for clinicians, geneticists and
researchers. If in most cases the knowledge of
the disease causing mutations is sufficient,
additional information are usually available such
as polymorphisms or unclassified variations.
These data are today wasted for the community
as they are not available through Locus Specific
DataBases (LSDBs). It is now time to collect
these data and to move from mutations to
haplotypes. For this purpose, we have
implemented new features in the UMD®
software. The user can collect a virtually
unlimited number of polymorphisms and
mutations for a specific patient. He can specify if
these variations are localized on the same
chromosome (cis) or on different ones (trans).
The UMD® software can thus build haplotypes
for all samples. Specific routines have been
designed to produce association tables between
variations (allele-allele associations), variations
and haplotypes (allele-haplotype associations)
and
haplotypes
(haplotype-haplotype
associations). These tools are useful to identify
association disequilibrium, to analyze complex
alleles and to identify modifier effects.
Here, we report the creation of the UMD-CFTRCBAVD database, which includes mutations of
the CFTR gene specifically associated with
congenital bilateral absence of the vas deferens
(CBAVD). Today, this database contains 242
mutations and 1543 polymorphisms from 159
individuals and six mutations account for 70.7%
of mutations. Surprisingly, if 3 haplotypes are
frequently found (39%), 62 among the 90
different haplotypes have been reported only
once.
This observation led us to develop a unique
module for LSDBs: Phylogeny. This tool creates
a PHYLIP formatted file of the various
haplotypes. It then runs the dnamlk program
from the PHYLIP package. This program
implements the maximum likelihood method for
DNA sequences under the constraint that the
trees estimated must be consistent with a
molecular clock. At this step, the UMD®
software collects the output file and draws the
estimated tree. To easily interpret this tree,
additional color codes are drawn for each
variation. Thus, we have shown that, in the
context
of
CBAVD,
the
c.1327G>T
(p.Asp443Tyr) mutation is specifically found in
haplotypes harboring the TG10 and 7T alleles
from intron 8 in addition to c.1727G>C
(p.Gly576Ala) and c.2002C>T (p.Arg668Cys)
polymorphisms.
Similarly,
the
c.1522_1524delTTT (p.Phe508del), accounting
for 32% of mutations, is specifically found on
chromosomes harboring the TG10 and 9T
alleles from intron 8.
These results show that the collection of
polymorphism information resulting from largescale sequencing should be reported in LSDBs.
This will open a new area and contribute to the
identification of highly conserved ancestral
chromosomal segments.
INTERPRETING MISSENSE VARIANTS IN
OCULOCUTANEOUS ALBINISM GENES
TYROSINASE AND P GENE: AN
EVOLUTIONARY APPROACH
M.S. Greenblatt, S. Duraisamy, C. McBride, J.P.
Bond, W.S. Oetting
Univ of Vermont, Univ of Minnesota
BACKGROUND:
Oculocutaneous Albinism
(OCA) is characterized by lack of skin and eye
pigment, UV sensitivity, predisposition to skin
cancer, and developmental eye defects. OCA is
caused by mutations in several genes, most
commonly the tyrosinase (OCA1) and P genes
(OCA2). Tyrosinase is the key enzyme
catalyzing melanin pigment synthesis from
tyrosine; the P gene encodes a transport
protein.
Understanding
their
intragenic
conservation patterns can help to predict
functionally critical amino acids (AA), where
variants are likely to cause OCA.
OBJECTIVES: To assess quantitatively the
predictive value of AA conservation we have: 1)
made sequence alignments and phylogenetic
trees of the tyrosinase and P genes, 2)
computationally studied their intragenic AA
conservation patterns, and 3) tested how well
the patterns predict OCA-associated variants
and non-OCA-associated polymorphisms, using
the
Albinism
database
(http://albinismdb.med.umn.edu).
RESULTS: Evolutionary variation in the existing
database of sequences is sufficient to determine
statistically significant conservation of codons for
tyrosinase but not for P gene. The SIFT program
correctly predicted 82 % of the OCA1 associated
tyrosinase variants and both of the non-OCAassociated polymorphisms, and 94% of the
OCA2-associated P gene variants. However, 4
of 6 P gene polymorphisms were incorrectly
predicted to be deleterious based on
conservation in the three known P gene
sequences (human, pig, mouse). To achieve
82% sensitivity for predicting OCA1-associated
tyrosinase variants, cutoff scores for AA
conservation would classify 50% of codons as
conserved, with resulting 50% specificity, 100%
Positive Predictive Value (PPV), and only 9%
Negative Predictive Value (NPV). For predicting
OCA2-associated P gene variants, cutoff scores
to achieve 85% sensitivity would classify 73% of
codons as conserved, with resulting 50%
specificity, 91% PPV, and 38% NPV.
CONCLUSIONS: 1) >80% prediction of
deleterious mutations is possible. 2) Adequate
databases of sequences, mutations, and
polymorphisms are needed to confirm the
validity of predictions based on evolutionary AA
conservation. Since few polymorphisms have
been reported in these genes, the power of this
technique is reduced. 3) More P Gene
sequences are needed to validate evolutionary
predictions.
COSMIC (CATALOGUE OF SOMATIC
MUTATIONS IN CANCER)
- A RESOURCE FOR CANCER
MUTATIONS
Richard Wooster, P Andrew Futreal, Michael R.
STratton
Cancer Genome Project, Sanger Institute,
Wellcome Trust Genome Campus, Hinxton,
Cambridgeshire, CB101SA, UK
COSMIC is a resource for cancer genetics; a
database of published mutations involved in
human cancer
(http://www.sanger.ac.uk/cosmic). Focusing on
genes of key interest in carcinogenesis, the
project is analysing an expanding gene set,
starting with a group of 17 including BRAF,
FGFR3, HRAS, RET, PTEN. As of October
2004, data have been collated from more than
2,800 journal publications, comprising over
114,000 tumours analysed for tumourigenic
mutations. 17,929 of these samples had
mutations, comprising 1,728 unique sequence
alterations.
COSMIC collates all key genomic data
describing a mutation event in a tumour sample,
including the experimental procedures used, and
the tissue type / histological classification of the
tumour sample itself. Web pages overlay this
data, providing highly selectable graphical views
and tabulated summaries of the information
together with links to external sources such as
PubMed, Swissprot and Pfam. COSMIC is also
being used as a vehicle to display data from the
systematic mutation screens being performed by
the Cancer Genome Project.
AUTOMATED SPLICE SITE
MUTATION ANALYSIS BY
INFORMATION THEORY
Vijay K. Nalla & Peter K. Rogan. Children’s
Mercy Hospital, Schools of Medicine
and Computer Science & Engineering,
University of Missouri-Kansas City, USA
Significance: Accurate interpretation of
mutations that alter non-coding, conserved
sequence elements in human genes is important
for diagnosis and prognosis of inherited or
acquired genetic disorders. The effects of such
mutations can be predicted in silico by
information theory (Hum. Mut. 6:74-76; Hum Mut.
12: 153-171). This is because changes in the
affinity of a protein or protein complex for its
cognate binding site can be estimated from the
individual information content of the sequence of
a site, which is based on an information weight
matrix describing a set of sites recognized by
the same protein(s) (J. Theor. Biol. 189: 427441).
Implementation: Perl and C++ software
was developed to assist in interpretation of noncoding sequence variation in functional elements
within human genes. Initially, algorithms were
implemented to analyze mutations expressed in
standard HUGO nomenclature from a wide
variety of peer-reviewed sources, including
published literature and conforming locus
specific databases. The software is also able to
process
some
non-standard
mutation
designations used in popular locus specific
databases. Genomic coordinates homologous
to the mRNA accession number or gene are
parsed from a MySQL database, then the
reference and corresponding variant or
haplotype sequences are verified and retrieved.
The software introduces mutation(s) into the
reference sequence and computes the resulting
information contents (and changes, if any) at
splice sites and/or regulatory sequences (human
and murine donor and acceptor sites, SC35,
ASF/SF2, and SRp40 regulatory binding sites).
These tools were incorporated into a
secure web server that detects changes in
information content at binding sites due to
mutations in any catalogued human gene,
genome-mapped
mRNA
or
user-defined
sequence [https://splice.cmh.edu]. Individual
information analyses with splice junction binding
site matrices detects activated cryptic splice
sites, associated splicing regulatory sites and
distinguishes null alleles from those that are
partially functional. Standard gene and mutation
nomenclature (HUGO-approved format) are
entered using a CGI-based front end supported
by backend server. Changes in information
content are tabulated and visualized as walker
figures showing the binding sites on the
sequence.
Performance: Depending upon the set
of options that are selected, the Web server
requires 30 ~ 60” to analyze one mutation under
optimal CPU loads. Batch submission of 50
mutations has been clocked at 2’ 25” on a single
2.4 MHz I686 processor running Redhat Linux.
Results: This application has been
validated by analysis of ~600 mutations
(including haplotypes) parsed directly or
interactively from Human Mutation and locusspecific databases. We confirmed that all of the
previous recognized splicing mutations affected
splice site strength or activated cryptic splice
sites. Information analysis identified 8 examples
of missense mutations that concomitantly
affected adjacent splice donor and acceptor
sites, 4 partially functional splice sites previously
thought to be null alleles, and 6 unrecognized
cryptic splice site mutations.
Alterations in
presumed SR protein binding sites that may
impact known splicing regulatory elements were
also detected in a number of cases.
Conclusions: The system has been
designed to facilitate addition of other
information weight matrices and genomes and
software for scanning multipartite binding sites in
future versions. It should be feasible to develop
more comprehensive models of splicing
phenotypes and eventually, to analyze binding
sites in other regulatory sequences.
Acknowledgements: Grant support [PHS
ES10855] from the NIEHS is acknowledged.
MUTATIONVIEW : DEVELOPMENT OF
AN ENHANCED SEARCH SYSTEM
FOR THE CLINICAL WORDS IN
OMIM WITH A NEW ALGORITHM
1,2)
2)
Shinsei Minoshima , Masafumi Ohtsubo ,
3)
3)
Katsue Daicho , Kouichi Kawaguchi , Susumu
2)
2)
Mitsuyama , Takashi Kawamura , Tomoyoshi
3)
2)
Horisawa , Nobuyoshi Shimizu
1)
Photon Med. Res. Ctr., Hamamatsu Univ. Sch.
2)
Med., Dept. Mol. Biol., Keio Univ. Sch. Med.,
3)
Chi Co., Ltd.
In order to investigate the relevance between
disease and genetic diversity, we have
developed an integrated knowledge-base
system, MutationView (http://mutview.dmb.med.
keio.ac.jp/).
Current
MutationView
has
principally focused on monogenic disease
mutations, and its characteristic features are as
follows: (1) Various data display: genomic/cDNA
structure, functional domain of protein,
histogram for the case number of mutations,
changes in the nucleotide/amino acid sequence
and restriction sites with graphical environment,
(2) Analysis functions: classification based on
the various information included in each case
record (e.g. ethnic origin, onset age and
symptom).
To date, we have collected 10,166 entries of
mutations from 1736 literatures dealing with 259
genes involved in 407 distinct diseases
according to particular categories such as eye,
brain, muscle, ear, heart, autoimmunity and
familial tumor. Recently, the systemic bone
disease was added as a new category.
Moreover, we have developed a genome
browser,
which
has genome-wide
data
presentation function. V arious information such
as
chromosome
band,
contig,
gene,
transcription, and DNA marker covering all the
chromosomes was automatically imported from
Ensembl (http://www.ensembl.org/). Recently,
we have developed a new search system for
OMIM.
Scientifically significant words were
picked up using an existing dictionary, and then
sets of meaningful interrelation between them
were computationally extracted from the OMIM
with a statistical analysis based on coincidental
appearance of the words in the same section of
each OMIM document. These interrelation are
quite useful for the search based on word
association. Thus, we are developing a new
algorithm of an ab initio method for extraction of
important words without dictionaries, in which
“Variance” value of each word in OMIM is
utilized. The system will be demonstrated at the
meeting.
PHENOTYPES VS GENOTYPES
IN THE WORLD OF
BLOOD GROUP ANTIGENS
1
2
Olga O. Blumenfeld and Santosh K. Patnaik ,
1
2
Deparments of Biochemistry and Cell Biology ,
Albert Einstein College of Medicine, New York,
NY 10461
Blood group antigens are proteins, glycans or
glycolipids, of a variety of functions, whose
common feature is that all are expressed on the
surface of red cells and are polymorphic in the
population. The hallmark of each antigen is an
epitope, a linear or spatially arranged amino acid
or carbohydrate sequence which due to its
variant nature can be recognized as non-self by
the immune system. The science (art) of
serology is based on this recognition, and its
goal is to decipher and assign blood group
phenotypes using antibodies to the polymorphic
epitopes as tools. A blood group system is a set
of variant antigens encoded by alleles of a single
locus, each defining a related form of a common
blood group phenotype.
The Blood Group Antigen Gene
Mutation Database* documents 36 genes
encoding 29 blood group systems and
comprising close to 700 alleles that result in
surface expression of at least 400 different blood
group phenotypes. Here we examine the
correlation between the genotype, the structure
of the allele and the blood group phenotype.
This is a rare example in which a direct
correlation between a DNA alteration and a
single physiologic function (antibody response)
can be established, with modifier genes or
environmental factors playing a minimal role.
In the database, DNA alterations are
documented in donors who were selected for
study on the basis of a variant blood group
phenotype. Thus an alteration of the epitopic
and/or other segments of the coding regions is
expected. As generally observed for many
documented human sequence variations,
among total DNA alterations, single nucleotide
mutations predominate (sense ~8%, missense
~50%, nonsense ~6%). Nonsense mutations
and small deletions, insertions, or splice site
alterations, often accompanied by additional
upstream or downstream mutations, give rise to
a number of related alleles whose products are
truncated or defective and result, directly or
indirectly
(e.g.,
in
case
of
inactive
glycosyltransferases), in the absence of
epitopes from the cell surface. Thus, for several
blood group systems such different alleles result
in the null phenotype (e.g., the O phenotype of
the ABO system). Gene rearrangements based
on different mechanisms (gene conversions,
unequal recombinations) can give rise to the
same epitopic sequence and the same
phenotype specified by different hybrid alleles
(e.g., the MN system). In contrast, in most
instances, missense mutations linked to the
epitope give rise to variant phenotypes, each
characteristic of a specific mutation and the
amino acid change. This is observed whether
the protein exhibits a single or multiple epitopes;
the effect of the mutation can be direct or
indirect (affecting the level of cell surface
expression). The majority of DNA alterations
have no apparent effect on the function of the
erythrocyte. The subtle relationships among a
single amino acid replacement affecting the
epitopic region, the immune response, and the
ability to detect it, become apparent from a
survey of the different blood group systems
documented in the database.
*http://www.bioc.aecom.yu.edu/bgmut/index.htm
DEVELOPMENT OF A PUTATIVE
FUNCTIONAL CODING SNP SET FOR
WHOLE-GENOME DIRECT
ASSOCIATION STUDIES
Francisco M. De La Vega, Charles R. Scafe,
Anish Kejariwal, Eugene G. Spier, Paul D.
Thomas, and Dennis A. Gilbert. Applied
Biosystems, Foster City, CA, USA
Genetic association studies of complex disease
can be carried out by: indirect association, in
which surrogate markers assumed to be in
linkage disequilibrium (LD) with the disease
allele are tested for trait association; or direct
association, in which a list of putatively
functional SNPs are tested for their disease
relevance directly. Current estimates of the
number of markers needed for a genome scan
via LD in humans range from 120,000 to over a
million SNPs, translating into enormous
genotyping costs and a challenging problem of
statistical inference.
On the other hand, a
genome scan with putative functional SNPs may
require typing tens of thousands of common
causative variants implicated in complex
disease. The feasibility of whole-genome direct
association studies requires that most of the
variants influencing the susceptibility to disease
are typed in the study. Currently, about 40,000
non-synonymous coding SNPs (nsSNPs) are
deposited on the public databases. An additional
30,000 novel nsSNPs were discovered through
the resequencing of exonic regions of 23,363
genes by the Applera Genomics Initiative.
Combined,
these
datasets
provide
a
comprehensive resource of nsSNPs making the
development of a whole-genome direct
association marker set feasible. We are
currently pursuing implementation of such a set
on the SNPlex™ Genotyping System, a
multiplexed high throughput genotyping platform
based on the oligonucleotide ligation/PCR
assay. SNPs can be prioritized as to their assay
conversion potential, heterozygosity, location in
the gene, gene’s product classification, and their
putative functional impact. In our preliminary
design, we compiled 28,709 nsSNPs including
over 9,000 proprietary SNPs based on their
measured or expected heterozygosity in
populations of European and African descent.
SNPs were grouped when possible by protein
families based on the PANTHER protein
classification before submission to the SNPlex
assay design pipeline. We designed SNPlex
assays for about 70% of the SNPs. Currently we
are engaged in improving the assay design
conversion rate and in collaborative studies for
the validation of the set in complex disease
cohorts.
RESEQUENCING ON A CHIP: IS IT
READY FOR PRIME TIME?
Arupa Ganguly1, Courtney MacMullen2, Charles
A. Stanley2
1Department of Genetics, University of
Pennsylvania School of Medicine; 2Division of
Endocrinology, The Children’s Hospital of
Philadelphia
Resequencing of multiple large genes in parallel
can be achieved by using high density oligonucleotide microarrays. We summarize our
experience of resequencing using Custom-Seq
chips synthesized by Affymetrix, CA in the
context of Congenital Hyperinsulinism (HI). HI,
the most common cause of persistent
hypoglycemia in infants, is caused by mutations
in at least 4 different genes: the KATP
potassium
channel
comprised
of
the
sulfonylurea receptor (SUR1) and inward
rectifying
agent
(Kir6.2),
glutamate
dehydrogenase (GDH), and glucokinase (GK).
We designed the HI-Chip for detecting variations
in the 65 coding exons including flanking intronic
regions of the 4 HI genes. We evaluated the
resequencing method using genomic DNA from
13 individuals: 1 negative control, 2 positive
controls, and 10 unknowns, and compared the
results to conventional automated fluorescent
single pass direct sequencing on an ABI 3100
machine.
Initial
base
call
rates
averaged
92%
(9393bp/10210bp) with a range of 61 to 96%.
This translates to an average of 12 ‘n’ calls per
tiled exon. However, the settings of the GDAS
analysis software can be adjusted to improve
the call rate at a cost of accuracy. Under
modified condition, the average base call rate
improved to 96.49%(9851bp/10210bp) with a
range of 84.18- 98.45%. This resulted in an
average of 6 ‘n’ calls per exon.
Resequencing on a chip has the potential to
return HI genotypes in a time efficient manner.
The method has to be optimized to perform as
well as direct sequencing for any custom set of
genes. About 60% of the amplicons have to be
directly sequenced after scanning due to
presence of undefined ‘n’ calls. The sequence
context of these n’s is predominantly a stretch of
Cs (>3 in a row). Visual inspection of the probe
hybridization data indicate the possibility of
better base calls if the algorithm could be trained
to look at the data from one strand only. In time,
the HI-CHIP may prove to be a clinically useful
pre-operative diagnostic tool and guide the
extent of pancreatectomy in non-medically
responsive HI cases.
ARRAY-BASED MLPA ANALYSIS
1,2
1
1
2
M.E. Kalf , S.J.White , M.Kriek , L. Vahlkamp ,
2
1
R. van Beuningen , M.H. Breuning , J.T. den
1 1
Dunnen . Leiden University Medical Center,
2
Leiden, Nederland, PamGene International
B.V., 's Hertogenbosch, Nederland
Due to its simplicity and broad applicability,
Multiplex
Ligation-dependent
Probe
Amplification (MLPA) is becoming the method of
choice for detecting deletions and duplications in
genomic DNA. A drawback of the current
methodology is that probes are separated
according to length, limiting the number of loci
that
can
be
amplified
and
analysed
simultaneously. The use of probes of uniform
length can be expected to facilitate the
simultaneous amplification of hundreds of
probes in one multiplex PCR. Furthermore,
analysis of these products using micro-arrays
should allow the read-out of potentially
thousands of loci at a resolution currently
unattainable with array CGH.
To test this approach we have quantitatively
scored MLPA-products using a porous microTM
array substrate (PamChip arrays). Compared
with planar arrays, these arrays have several
advantages, including a larger surface area, the
possibility to vary hybridization stringency during
analysis and, most interestingly, a decreased
hybridization time of about 10 minutes.
We combined four commercially available probe
sets (product range 130 bp to 490 bp) with our
own synthetically produced probes (product
range 80 bp to 125 bp), allowing 180 loci to be
analysed simultaneously. Analysis of samples
with known mutations gave results that were
comparable with those derived using capillary
electrophoresis. It was noticeable that in this
complex amplification longer probes were
amplified with a reduced and variable efficiency,
resulting in lower signals and less reliable
scores. This suggests that using probes of
identical lengths should improve consistency of
the results. A reduction of the hybridization time
prior to probe ligation from 16 hours to 2.5 hours
did not affect MLPA-accuracy. This indicates
that a complete MLPA-analysis, including DNA
isolation, can be obtained within 8 hours.
HIGHLY SENSITIVE, EFFICIENT AND
RAPID MUTATION IDENTIFICATION
FOR GENES WITH A HIGH
FREQUENCY OF NOVEL MUTATIONS
ENHANCES QUALITY OF CARE AND
REDUCES COSTS
Nadia Prigoda, Katherine Zhang, Kirk
Vandezande, Diane Rushlow, Beata Piovesan,
Ning Chen and Brenda L. Gallie
Solutions by Sequence, Retinoblastoma
Solutions and HHT Solutions, University Health
Network, University of Toronto, Toronto, Canada
We describe a strategy for mutation
identification in genes harboring a wide
spectrum of disease-causing mutations. This
multi-assay strategy is optimized for each
disease gene studied, detects a wide variety of
mutations, and achieves high test sensitivity and
cost-efficiency in a relatively short turnaround
time. We have optimized the methods for RB1
gene mutations that lead to retinoblastoma and
for ALK-1 and ENG gene mutations that lead to
hereditary hemorrhagic telangiectasia (HHT).
Quantitative Multiplex PCR (QM-PCR) efficiently
detects changes in the size or copy number of
exons, approximately 35% of mutated RB1
alleles. Half (18%) of the RB1 mutations
detected by QM-PCR are small insertions and
deletions; half are whole exon deletions or
duplications not detected by sequencing. A
single allele-specific PCR (ASPCR) reaction
detects eight commonly recurring
RB1
mutations, 21% of the mutations identified. We
optimized sequencing reactions by combining
two exons in each reaction with different labels
and ordering duplex reactions to maximize the
mutation discovery rate. Sequencing alone
detects approximately 68% of mutated RB1
alleles. Methylation-specific PCR identifi es RB1
promoter hypermethylation in 13% of unilateral
retinoblastoma tumors. If all of these DNA -based
methods fail to identify the mutated allele, we
use RNA-based methods to search for deep
intronic splice mutations. Overall, these methods
enabled us to identify the germline mutation in
92.4% of 314 blood samples from individuals
with bilateral or familial unilateral retinoblastoma,
and both somatic mutations in 89% of the 210
tumor samples from individuals with unilateral
sporadic retinoblastoma. Median turnaround
time in 2003 was 4.6 weeks. This strategy for
RB1 mutation identification simultaneously
achieved significant saving in health costs,
improved the standard of care, and improved the
clinical outcomes for most of the retinoblastoma
families studied. Using a similar technical
strategy combining QM-PCR and duplex
sequencing, we have detected ALK -1 and
endoglin (ENG) mutations in approximately 85%
of patients clinically diagnosed with HHT. This
suggests that our highly sensitive, rapid and
efficient approach to mutation detection can be
successfully extended to other similar genes.
We advocate the use of our hierarchical
mutation identification strategy to screen other
large complex genes with a wide spectrum of
disease-causing mutations, such as BRCA 1/2.
THE EFFECT OF CODING
SYNONYMOUS AND 3’ UTR SNPS ON
MRNA SECONDARY STRUCTURE
Huiqi Qu and Constantin Polychronakos,
Endocrine Genetics Laboratory, Department of
Pediatrics, McGill University Health Center,
Montreal, Québec, Canada
Background In the search for functional
features of the human genome altered by singlenucleotide polymorphisms (SNPs), mRNA
secondary structure has not received much
attention. However, mRNA secondary structure
differences caused by SNPs have a significant
potential to influence biologically important
functions
by
changing
mRNA
stability,
translational efficency or other, less well
understood RNA functions. Since consensus
sequences identifying functionally important
RNA regions are not nearly as well understood
as the corresponding DNA features, a nonspecific but potentially powerful screening
approach would use effects on secondary
structure to identify priority SNPs for functional
evaluation from among a list of candidates—for
example in a search for the functional variant
within a tight linkage disequilibrium block
associated with a complex disease. In this study,
we explored the feasibility of predicting SNP
effects
on
mRNA
secondary
structure
computationally, by minimum free energy
change, compared effects by different nucleotide
substitutions and made a first attempt at
correlating this to experimentally verified
function.
Methods
Using the NCBI dbSNP database
(http://www.ncbi.nlm.nih.gov/SNP/),
synonymous SNPs (sSNPs) and 3’ UTR SNPs
were selected for this study. The human dbSNP
database build 122 includes 7893 sSNPs and
52388 UTR SNPs from 22 autosomal
chromosomes with heterozygosity>0.10. From
these SNPs, 100 sSNPs and 100 3’ UTR SNPs
were selected randomly. Because of the
uncertain delineation of the 5’ UTRs in many
genes, SNP’s in this region were not included.
This decision was made before any results were
known. Minimum free energy (MFE) of whole
length mRNA was computed on the basis of an
energy minimization algorithm (Zuker and
Stiegler 1981), by the Vienna RNA Package
(Hofacker 2003, http://rna.tbi.univie.ac.at/cgibin/RNAfold.cgi). Because of a limitation of this
algorithm, only SNPs located in mRNA <4000
nucleotides in length were included.
The
change of minimium free energy of mRNA
moleculars was compared among different
nucleotide substitutions and between sSNPs
and 3’ UTR SNPs. Six synonymous or 3' UTR
SNPs
with
experimentally
demonstrated
functional effects were found in an exhaustive
search of Pub Med. MFE changes by these
SNPs were compared to thos e of the randomly
selected SNPs.
Results Between sSNPs and 3’ UTR SNPs,
there was no difference as to the distribution of
each type of nucleotide substitution or MFE
changes, either by type of substitution or as a
whole. Therefore for subsequent analysis the
two types were pooled. Among those 200 SNPs,
the most common nucleodite substitution was
C/T, twice as frequent as its complementary A/G
(54.5% vs. 26.5%, p=0.000017). One way
ANOVA suggested statistically significant
differences in MEF change among different
substitutions (F=2.53, p=0.03). The difference
was further confirmed in an additional group of
240 SNPs, randomly selected so that each type
of substitution was equally represented (F=4.3,
p=0.001). G/T substitutions made up 4% of the
total and had the highest average MFE change
at 2.02 ± 0.24 kcal/mol (mean+SEM). The most
common C/T was found to have the lowest
average MFE changes at 1.03± 0.084.
The mean of MFE changes by the six sSNPs
or 3’ UTR SNPs with experimental evidence of
functional effect were more than twice that of
random SNPs (2.60+0.82 vs. 1.23+ 1.22
kcal/mol).
Conclusion
This
study
suggests
that
computational prediction of mRNA secondary
structure change by MFE assessment can
predict experimentally verified functional effects
of a SNP. This justifies further exploration of its
potential to identify SNPs for further in-depth
experimental analysis in search of functional
variants.
Hofacker, I. L. (2003). "Vienna RNA secondary
structure server." Nucleic Acids Res 31(13):
3429-31.
Zuker, M. and P. Stiegler (1981). "Optimal
computer folding of large RNA sequences using
thermodynamics and auxiliary
Nucleic Acids Res 9(1): 133-48.
information."
PMSG, A NOVEL TECHNOLOGY TO
ANALYZE DNA AND RNA
SEQUENCES FOR VARIATION
1
1
1
Dan Graziano , Chaof u Shi , Angela Alexander ,
2
1
Jon Jarvik and Cheryl Telmer
1
SpectraGenetics LLC, 4415 Fifth Ave., Suite
160, Pittsburgh PA, 15213
2
Department of Biological Sciences, Carnegie
Mellon University, 4400 Fifth Ave., Pittsburgh
PA, 15213
Peptide mass signature genotyping, PMSG, is a
scanning genotyping method that detects and
characterizes known and novel mutations and
polymorphisms by (1) amplifying the sequences
of interest by PCR, (2) translating the amplicons
in more than one reading frame, (3) affinity
purifying the resulting peptides, and (4)
analyzing the peptides by mass spectrometry.
The set of peptide masses encoded in a given
sequence comprises a “peptide mass signature”
characteristic of that sequence. Individual
sequence variants typically yield different and
distinct mass signatures because they change
the amino acid composition, and hence the
mass, of one or more of the peptides of which
the signature is comprised. Once a given
signature has been detected and verified by
dideoxy sequencing, the signature serves to
identify that sequence variant in subsequent
analyses.
We have applied PMSG technology to several
genes
including
the
tumor
suppressor
geneTP53. The TP53 test analyzes exons 2 to
11 of the gene and all splice sites in a
multiplexed configuration at a cost comparable
to dideoxy sequencing but with greater
sensitivity. A PMSG test that detects mutations
and splice variants in mRNA is also under
development. The status of the DNA and RNA
tests will be discussed, and the advantages and
limitations of PMSG will be addressed with
respect to detecting and characterizing germline
and somatic sequence variation.
Telmer, C.A., Retchless, A.R., Kinsey, A.D.,
Conley, Y., Rigatti, B., Gorin, M.B., Jarvik, J.W.
2003. Detection and assignment of mutations
and minihaplotypes in human DNA using
Peptide Mass Signature Genotyping (PMSG):
application to the human RDS/Peripherin gene.
Genome Research 13: 1944-1951.
Telmer, C.A., An, J., Malehorn, D.E., Zeng, X,
Gollin, S.D., Ishwad, C., Jarvik, J.W. 2003.
Detection and assignment of TP53 mutations in
tumor DNA using Peptide Mass Signature
Genotyping. Human Mutation 22: 158-165.
DEVELOPMENT AND VALIDATION
OF ECONOMIC, HIGH THROUGHPUT
MELTMADGE ASSAYS
FOR THE BRCA1 CODING REGION
Aldahmesh M.A, Spanakis E, Alharbi K.K,
Sillibourne J, *Day I.N.M, Eccles D.M
Human Genetics Division, School of Medicine,
Duthie Building (MP 808), Southampton University
Hospitals
NHS
Trust,
Tremona
Road,
Southampton SO16 6YD, UK
* Author for correspondence.
Email:[email protected]
Telephone +44 1703 795063 (Sec); fax +44
1703 794264
MeltMADGE offers 30-100 fold economy and
throughput advantages for mutation scanning
over standard prevalent techniques such as
DHPLC, SSCP and CSGE, but has not
previously been evaluated in any diagnostically
relevant genes. In this study we have developed
54 PCR-meltMADGE assays representing the
BRCA1 coding region, and part of the 3'noncoding region. Ten known SNPs were
detected and also different pathogenic mutations
previously characterised by other methods were
detected in blind trials on a panel of 94 unrelated
subjects. In addition, a new SNP in the 3'noncoding region and two new sequence
variants were also identified. However, the same
system can run an assay on > 1,000 subjects
simultaneously with the main cost increment
being that of the PCR reactions. This approach
has reasonable sensitivity to base changes and
the potential to be applied to complete
description of population missense mutation
diversity (“reference range” studies) and to initial
economic scans of diagnostic lab backlogs and
of larger lower risk groups.
AN EVALUATION OF THE
ACCURACY AND PRECISION OF
TWO DNA POOLING STUDY
DESIGNS BY COMPARISON OF
POOLING RESULTS WITH
INDIVIDUAL GENOTYPE
DATA FOR 18,000 SNPS
INTERACTION AND ASSOCIATION
EFFECTS OF MYOC, OPTN AND
APOE IN PATIENTS WITH PRIMARY
OPEN ANGLE GLAUCOMA
1
1
1
1
CP Pang, BJ Fan, DY Wang, POS Tam, YF
1,2
1
Leung, DSC Lam
1
1
2
2
Ansar Jawaid , Kelly Frazer , David Hinds and
1
Neil J Gibson
1. Research and Development Genetics,
AstraZeneca Pharmaceuticals, Alderley Park,
Macclesfield, SK10 4TG, United Kingdom.
2. Perlegen Sciences, Inc. 2021 Stierlin Court,
Mountain View, CA 94043-4655, USA
Pooling of individual DNA samples prior to
genotyping has been proposed as a practical
method of performing whole genome association
studies with very large numbers of SNP markers
1
. DNA pooling can significantly reduce the
consumable and labour costs of a study as well
as reducing DNA usage. However, experimental
errors inherent in a pooling design lead to a loss
in power compared to individual genotyping.
Furthermore, such errors increase the false
positive rate, if not appropriately quantified and
considered in the analyses, leading to significant
2
difficulty in exploiting the results . Experimental
errors arise from not obtaining exactly equal
amounts of DNA from each individual included in
the pool, variation in the quantitative accuracy of
the method to measure allele frequencies, and
differential amplification of alleles. Here we
examine the sources of experimental error using
empirical data from large pool and small pool
designs with corresponding individual genotype
counts for 18,000 SNPs typed in 685 samples
using a high density array -based genotyping
3
platform.
References
1. Norton N, Williams NM, O'Donovan MC,
Owen MJ. Ann Med. 2004; 36(2): 146-52
2. Zou G, Zhao H. Genet Epidemiol. 2004
Jan;26(1):1-10.
3. Patil N, Berno AJ, Hinds DA et al. Science.
2001 Nov 23;294(5547):1719-23.
Department of Ophthalmology & Visual
Sciences, the Chinese University of Hong Kong,
2
Hong Kong. Present affiliation: Bauer Center
for Genomics Research, Harvard University,
Cambridge, USA.
Primary open-angle glaucoma (POAG) is a
leading cause of visual impairment and
blindness worldwide. Myocilin (MYOC) and
optineurin (OPTN) are the two known diseasecausing genes for POAG. Apolipoprotein E
(APOE) is a potential candidate gene. To
explore the interactions among these genes in
POAG, we performed a multi-gene association
study in 200 sporadic POAG patients and 201
unrelated control subjects. We identified
disease-causing mutations (DCMs) in the MYOC
gene, R91X, E300K and Y471C, accounting for
1.5% of POAG patients. DCMs identified in the
OPTN gene were E103D and H486R, found in
1% of POAG patients. Two non-coding OPTN
variants, IVS6-5T>C and IVS6-10G>A, and the
APOE e4 allele, decreased POAG risk. The
OPTN IVS6-5T>C reduced POAG risk by 5- fold
based on multivariable analysis (P = 0.014) and
haplotype analysis (P = 0.0003). Two APOE
promoter polymorphisms, -491A>T and 219T>G, increased POAG susceptibility. Three
pairs of MYOC and OPTN polymorphisms, 1000C>G (MYOC) and M98K (O PTN), A260A
(MYOC) and R545Q (OPTN), I288I (MYOC) and
IVS7+24G>A
(OPTN),
showed
linkage
disequilibrium in POAG, indicating interactions
between the MYOC and OPTN genes.
Meanwhile, R545Q in OPTN decreased cup-disc
ratio (P = 0.004), and IVS8+20G>A and IVS1548C>A lowered IOP at diagnosis (P = 0.006 and
0.030 respectively). Our findings showed that
besides the DCMs, non-coding sequence
changes in MYOC and OPTN might also affect
the susceptibility of POAG. While our data did
not reveal modifying effects of APOE on the
glaucoma phenotype, we have shown potential
interactions between MYOC and OPTN to be
involved in the pathogenesis of POAG,
indicating a digenic etiology for sporadic POAG.