Download MIABE-parent - HUPO Proteomics Standards Initiative

Survey
yes no Was this document useful for you?
   Thank you for your participation!

* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project

Document related concepts

List of types of proteins wikipedia , lookup

Biochemistry wikipedia , lookup

Drug design wikipedia , lookup

Nuclear magnetic resonance spectroscopy of proteins wikipedia , lookup

Transcript
Minimum Information About a Bioactive Entity (MIABE)
Version 0.4
Sandra Orchard
Henning Hermjakob
Input/support to date
Steve Bryant
NCBI
Dominic Clark
EMBL-EBI
Ian Dix
AstraZeneca
Ola Engkvist
AstraZeneca
Mark Forster
Syngenta
Michael Gilson
BindingDB
Martin Grigorov
Nestlé
Kim Hammond-Kosack Rothampsted
Lee Harland
Pfizer
Andrew Hopkins
U. Dundee
Christopher Larminie GSK
Elena Lo Piparo
Nestlé
John Overington
EMBL-EBI
Chris Southern
EMBL-EBI
Christoph Steinbeck EMBL-EBI
Janet Thornton
EMBL-EBI
David Wishart
DrugBank
Feedback to [email protected]
Introduction The process of the identification and development of molecules
with useful bioactive properties, such as pharmaceuticals and pesticides, is
fraught with difficulty and many compounds will fall by the wayside on the road
from New Chemical Entity to licensed product. In the pharmaceutical industry
only a very small percentage of Investigational New Drugs will make it through
to clinical usage. The causes for this high attrition rate are many, with lack of
efficacy, unexpected drug side effects, and undesirable drug-drug interactions
being just some of the more common pitfalls in the drug discovery process.
Similarly, in the world of pesticides, compounds which prove to be active show
undesirable side-effects against organisms other than their original target will
not make it through to the market.
However, published reports of the activities of these ‘failed’ compounds, in
addition to detailed information on those which go on to become fully licensed,
commercially-available bioactive entities, are crucial for an understanding of
how improved molecules may be developed. Details of their molecular
structure and mechanism of action may give clues as to how related
analogues may be developed that hit the same target, but with improved
effectiveness or increased selectivity for a specific target over closely related
molecules. A full disclosure of observed toxicity or an understanding of the
pharmacokinetic properties of an agent may assist in improving these
properties in subsequent generations of molecules. Even those molecules
which fell by the wayside at an early stage in the development process may
have a role to play as tools, to enable the verification of potential new targets
which may have been identified by micro-array or proteomic studies in
diseased tissues.
In 2002, Hopkins and Groom introduced the concept of the ‘Druggable
Genome’ [1] suggesting that some proteins/protein families in the (human)
genome are more amenable to modulation by small exogenous molecules
than others. Since that time, work has been directed at finding ways to
develop a computational approach to calculate druggability and to predict
‘druggable’ proteins. In 2006 Overington et al estimated the number of
molecular targets for approved drugs to be as low as 266 protein targets [2],
out of the 20,400 protein coding genes [3] in the human genome. A further 58
targets are either from pathogenic organisms or are non-protein molecules
and similar models may be developed to look at the susceptibility of plant,
parasitic, bacterial or viral genomes to xenobiotics. This apparent very low
percentage of success may be significantly increased by adding to the
available data on licensed drugs information which lies in the literature and in
databases held by commercial companies, of molecular targets which have
proven to be successfully modulated but the agents have subsequently failed
to reached the end of the discovery pipeline due to unfavourable
pharmacokinetics, or toxicity linked to molecule class rather than target.
As a consequence of the current productivity crisis, the pharmaceutical and
biotechnology industries are increasingly disposed towards pre-competitive
release of compound-related bioactivity data into public domain repositories. It
has becoming widely acknowledged that the production of such an
aggregated resource allows a combinatorial increase in the value of the
information, in that the total amount of information that can be mined from it
will be of much more valuable than from a single isolated collection.
Increasingly, companies are recognising that regarding such an exercise as a
pre-competitive activity allows not only a general benefit to the fields of human
health and welfare but also a commercial advantage, in that the data need
only be collected once. The manual harvesting and curation of data is an
expensive process and involves resources often not available in even the
largest pharmaceutical company. Additionally, any one company has only a
limited range of in-house products targeting an equally restricted number of
targets. By adding to these the knowledge contributed by other groups in both
the commercial and academic sectors, a broader understanding of the
druggability of a potential target protein may be reached, well before
expensive resources are committed to an active research program. This precompetitive activity has wider benefits, for example in the fields of academic
or non-profit development of drugs for orphan and tropical diseases.
However, in order to fully understand the context, methods, data and
conclusions that pertain to the published description of an experiment,
detailed background information needs to be included. The diversity of
experimental designs and analytical techniques is becoming more of a
problem, not only as new methods supplant old ones but also as the scale of
data production increases due to the widespread application of automation.
While the problem is unlikely to be completely soluble it can be significantly
improved by the specification of this metadata (‘data about the data’). By
being associated with the results this metadata makes both the biological and
methodological contexts of the experiment explicit. The archetype of such a
specification is the Minimum Information about a Microarray Experiment
(MIAME) [4]. Many journals and funding agencies now require that authors
reporting microarray-based transcriptomics experiments comply with the
MIAME checklist as a prerequisite for publication. The adoption and
development of such specifications has had a much broader impact beyond
merely increasing the comprehension and comparability of journal articles, the
most important of which is facilitating the transfer of data from journal articles
into databases (i.e. converting unstructured to structured data) in a form that
will allow mining across combined data sets.
We therefore propose a new such document, the Minimum Information About
a Bioactive Entity (MIABE) that is predominantly concerned with, but is not
restricted to, bioactive chemical compounds. We believe the timing of this is
apposite for a number of reasons. The first of these is the revolution in
bioactive chemical information catalysed by the appearance of ChEBI [5] and
PubChem [6] towards the end of 2004. When the ”missing entity” of chemical
structure was embedded within the global Web of bioinformatic relationships it
became possible to search across biological effects, protein names, sequence
data, and chemical information. The second is that we now have the
deposition, not just of HTS results but also other types of bioactivity screening
data, directly linked to chemical structure information in public repositories
such as PubChem Bioassay [6], ChemBank [7], and, in the near future,
ChEMBL (www.ebi.ac.uk/chembl). Finally, increasing legislation is requiring
that information on these compounds be more readily available. For example,
on 1st June 2007, the REACH (Registration, Evaluation, Authorisation and
Restriction of Chemical substances) legislation came into force across the
European Union. This requires that additional information on chemicals be
made available, dependent on the quantity imported or manufactured, and it
has been estimated that 30,000 chemicals may need to be re-evaluated in
accordance with these requirements [8].
MIABE: Principles and process.
In order that the maximum benefit be derived from the publication of data on
one or a series of bioactive entities, it is important that certain crucial
information is included in every such paper such that the properties of these
molecules are fully represented, their effects on biological systems (both
positive and negative) are accurately detailed and any factors which may
contribute to the activity of the molecule are stated. To this end, a group of
both industry and academic groups have come together to layout a checklist
of the information that is felt to be important to include with any published
dataset, the minimum information required for the activity of a molecule to be
both fully understood and compared with other molecules, either sharing a
common chemical structure or similar mechanism of action. It is intended that
these guidelines be adhered to by anyone planning on publishing such a
paper and also by the implementers of resources such as a databases to hold
this information and any body or institution funding such discovery work and
requiring the results to be published at the end of the granting period. As with
other such reporting guidelines [4], MIABE adheres to the criteria of
Sufficiency (a reader should be able to understand and critically evaluate the
interpretation and conclusions, and to support their experimental
corroboration) and Practicability (the guidelines should not be so burdensome
as to prohibit its widespread use).
It should be noted that the scope of this series of documents is limited to data
regarded as pre-clinical in drug studies, a discussion of the publication of
clinical data is a subject more appropriate for a separate effort. It is also
recognised that the list of requirements described in the documentation
represents the entire path taken by a molecule from synthesis to pre-clinical
development. Many compounds, which fail to fulfil one or more of the criteria
necessary for the development of a lead compound, may only travel a part of
this route but still be a valuable research tool so all data generated on this
compound should appear in any publication on its activity. Similarly, data on a
successful agent may be published in multiple papers – in this case the
minimum reporting guidelines should be followed over the series of articles if
this is more appropriate. Finally, it should also be remembered that, although
this document targets bioactive agents, data on inactive, closely related
analogues and orthologues is often of equal value, providing negative controls
and information for those trying to build activity into a molecular scaffold. Such
data should also be fully reported using these guidelines.
The production and development of MIABE documents
This parent document exists to make explicit the scope, purpose and manner
of use of the MIABE guidelines that accompany it and, as such, should be
stable, as the principles described should therefore remain valid for the
forseeable future. Each domain within the drug discovery process will then be
discussed in a modular guideline document, the content of which may be
more subject to change over time as both experimental techniques and the
formats for recording data develop and increase. Initial versions of all modules
will be produced through the Pharmaceutical Industry Forum hosted at the
European Bioinformatics Institute and made available for extensive
community input prior to publication through pre-publication on Human
Proteomics Organisation, Proteomics Standards Initiative (HUPO-PSI)
website and documentation process [9] and also accessible via the MIBBI
portal (www.mibbi.org/index.php/Projects/MIABE, see below). New
documents and proposed updates to existing guidelines will be advertised on
appropriate websites, discussion groups and input requested from domain
experts. Each guideline document will then enter the formal peer-review
process of the journal in which it is to be published. Updates to the documents
will be discussed, agreed and published when necessary, but it intended to
make these as infrequent as advances in technology will allow, thus providing
long periods of stability to encourage adoption and implementation. Potential
authors may move between the guidelines (and other related community
efforts), ensuring that their data conform to those aspects of the checklists
which are relevant to their publication.
Data Formats
It is still common practise in many publication for a compound to be described
by no more than the structural representation of a pharmacophore, with the
modified regions then exemplified and the resulting molecules given an
identifier, often either specific to the originating source or to that particular
publication. Specific rules and conventions for compound nomenclature have
been developed and allow the specific reconstruction of a compound’s
structure, if strictly adhered to. However, as the needs of computational
chemistry become more central to the daily workings of molecular science, the
need for more systematic, and computationally translatable means of
describing compounds were required, and lead to the development of SMILES
strings [10] and the International Chemical Identifier (InChI) [11]. Whilst it is
recognised that it would be impractical to publish several hundred of these
identifiers in a large paper describing the structure-activity relationships of
several hundred compounds, it is becoming increasingly important that such
data be made available for published small molecules, such that subsequent
data capture no longer relies on individuals piecing together molecules
descriptions from a disparate figure and table(s) within a publication to redraw
the molecule in electronic format.
It has not previously been the practice in the field of bioactive molecules to
consider data exchange formats other than the published paper however as
public domain data repositories become established, the requirement to
exchange data between them, or for users to download non-redundant
datasets in a common format will become of increasing importance. Such a
common format does already exist and has been publicly available and in
wide usage for some years in the molecular interaction field. The HUPO PSIMI XML2.5 interchange format is capable of capturing extensive details about
many interactions types, including bioactive entities and their target
molecules, including the biological role of each molecule within that
interaction, detailed description of interacting domains, and the kinetic
parameters of the interaction [12]. The ability to describe the structure of
individual molecules and to carry meta-data on that molecule is also inherent
to the format. The format is supported by data management and analysis tools
and has been adopted by major interaction data providers, toll developers and
used in visualisation and analytical software. Additionally, a simpler, tabdelimited format MITAB2.5 has been developed for the benefit of users who
require only minimal information in an easy to access configuration. It is
suggested that supporting the continued development of this format, rather
than an attempt to “reinvent the wheel” will be the most practical step forward.
This will allow produces of drug-target information to merge their data with the
information existing the molecular interaction databases and use PSI-MI
XML2.5 compliant resources, such as Cytoscape (www.cytoscape.org), to
visualise small molecule data in conjunction with cellular interactomes or
pathways.
Controlled Vocabularies
Where possible, the user is asked to make use of existing controlled
vocabularies (CV), or ontologies, to describe entities, processes and
conditions described within a paper. This benefits both the human reader, in
that it reduces the possibility of ambiguity where a term is used which may
have multiple meanings, and also assists in the process of curation of the
information into data resources and subsequent search and analysis
procedures. A number of such controlled vocabularies are available on the
OBO website (www.obofoundry.org/) and may be used to describe, for
example, tissues, disease and molecular interactions (including
enzyme/substrate). The existing PSI-MI CV [12] used to annotate the format
described above has already been extended to allow for a full description of
the properties of a bioactive molecule and further input into the development
of this resource is welcomed.
MIABE in the context of related efforts
As already stated, the move to standardise the reporting and subsequent
collection and exchange of data is now being encouraged across the entire
biomedical field as the community joins in a united front to improve both data
quality and availability. In order to manage this process, the MIBBI project
(http://www.mibbi.org/) has been established, which maintains a web-based,
freely accessible resource for checklist projects, providing straightforward
access to extant checklists (and to complementary data formats, controlled
vocabularies, tools and databases) and ensures that new efforts are not
overlapping or redundant to an existing resource [13]. MIBBI is managed by
representatives of its various participant communities and its goal is to
facilitate the development of an integrated checklist resource site for the wider
bioscience community. Where MIABE can be seen to have areas of overlap
with existing resources, and obvious example of this are the MIMIx molecular
interaction standard [14], the MIAPAR protein affinity reagent standard
(www.mibbi.org/index.php/Projects/MIAPAR) and MIACA
(www.mibbi.org/index.php/Projects/MIACA), the standard for the description of
cellular assays, the work of these resources has been referred to rather than
repeated.
Reporting requirements for a bioactive entity
The minimum reporting requirements are divided into the following sections
MIABE document
MIABE-compound
Reporting Requirement
Physico-chemical properties of the molecule. Only
experimental parameters are included in the
minimum requirements, however if calculated
parameters are given, the method by which they
are calculated should be stated.
MIABE-compound
Molecule Source. The synthetic route by which a
molecule has been produced and /or extracted from
natural source.
In vitro assays – mechanism of action studies, to
ascertain off-target activities or drug metabolism
studies.
MIABE-assay
MIABE-assay
Cellular assays – mechanism of action studies, to
ascertain off-target activities or drug metabolism
studies.
MIABE-assay
In vivo assays – mechanism of action studies, to
ascertain off-target activities or drug metabolism
studies.
MIABE-assay
Pharmacokinetic studies – to include a full range of
observations made in vivo and in vitro.
MIABE-assay
Toxicological data to include post-mortem and
subsequent analyses.
Conclusion
The main aim of this document is to support the long-term bioactive molecule
discovery process, by ensuring that the data generated from expensive and
time-consuming previous projects is not lost but can be harvested and built on
in the future. By following such guidelines, authors will produce a richly
annotated dataset about the molecule, or series of molecules which they are
describing, which will allow effective data capture by anyone wishing to
reproduce or analyse the results. It is to be hoped that more groups will
recognise the value of putting such data into the public domain to become part
of an ever increasing knowledge bank on biomodulatory agents, allowing the
development of improved molecules targeting an expanded druggable
genome in an increasing number of species.
References
1. Hopkins, A.L., Groom, C.R. (2002) The druggable genome. Nature
reviews. Drug discovery 1:727-30
2. Overington, J.P., Al-Lazikani, B., Hopkins, A.L. (2006) How many drug
targets are there? Nature reviews. Drug discovery [2006 (5) ] page
info:993-6
3. The UniProt Consortium (2009) The Universal Protein Resource
(UniProt) 2009 Nucleic Acids Res. 37, 169-174
4. Brazma, A, Hingamp P, Quackenbush, J., Sherlock, G., Spellman ,P.,
Stoeckert, C, Aach, J., Ansorge, W., Ball, C.A., Causton, H.C.,
Gaasterland, T., Glenisson, P., Holstege, F.C, Kim, I.F., Markowitz, V.,
Matese, J.C., Parkinson, H., Robinson, A., Sarkans, U., SchulzeKremer, S., Stewart, J., Taylor, R., Vilo, J., Vingron, M.. Minimum
information about a microarray experiment (MIAME)-toward standards
for microarray data. Nat. Genet. 2001, 29, 365-371
5. Degtyarenko, K., de Matos, P., Ennis, M., Hastings, J., Zbinden, M.,
McNaught, A., Alcantara, R., Darsow, M., Guedj, M., Ashburner. M
(2008) ChEBI: a database and ontology for chemical entities of
biological interest. Nucleic acids research 36, 344-350
6. Sayers EW , Barrett T , Benson DA , Bryant SH , Canese K ,
Chetvernin V , Church DM , Dicuccio M , Edgar R , Federhen S ,
Feolo M , Geer LY , Helmberg W , Kapustin Y , Landsman D.,
Lipman, D.J., Madden, T.L., Maglott, D.R., Miller, V., Ostell, J., Pruitt,
K.D., Schuler, G.D., Shumway, M., Sequeira, E., Sherry, S.T., Sirotkin,
K., Souvorov, A., Starchenko, G., Tatusov, R.L., Tatusova, T.A.,
Wagner, L., Yaschenko, E. Database resources of the National Center
for Biotechnology Information (2008) Nucl. Acids Res. 36: 13-21
7. Seiler K.P., George, G.A., Happ, M.P., Bodycombe, N.E., Carrinski,
H.A., Norton, S., Brudz, S., Sullivan, J.P., Muhlich, J., Serrano, M.,
Ferraiolo, P., Tolliday, N.J., Schreiber, S.L., Clemons, P.A. (2008)
ChemBank: a small-molecule screening and cheminformatics resource
database. Nucleic acids research (36) p351-359
8. Pederson, F., de Brujin, J., Munn, S., van Leeuwen, K. (2003)
Assessment of additional testing needds under REACH. Effects of
(Q)SAR, risk based testing and voluntary industry activity. (2003)
European Commission report EUR 20863.
9. Vizcaino, J.A., Martens, L., Hermjakob, H ., Julian, R.K., Paton, N.W.
(2007) The PSI formal document process and its implementation on the
PSI website. Proteomics 7, 2355-2357
10. Weininger, D. SMILES, a Chemical Language and Information System
1. Introduction and Encoding Rules. J. Chem. Inf. Comput. Sci. 1988,
28, 31−36
11. Stein SE, Tchekhovskoi D, Heller SR. The IUPAC and NIST Chemical
Identifier. 2004
12. Kerrien, S., Orchard, S., Montecchi-Palazzi, L., Aranda, B., Quinn,
A.F., Vinod, N., Bader, G.D., Xenarios, I., Wojcik, J., Sherman, D.,
Tyers, M., Salama, J.J., Moore, S., Ceol, A., Chatr-aryamontri, A.,
Oesterheld, M., Stümpflen, V., Salwinski, L., Nerothin, J., Cerami, E.,
Cusick, M.E., Vidal, M., Gilson, M. Armstrong, J., Woollard, P., Hogue,
C., Eisenberg, D., Cesareni, G., Apweiler, R., Hermjakob, H. (2007)
Broadening the horizon--level 2.5 of the HUPO-PSI format for
molecular interactions. BMC biology 5: 44
13. Taylor, C.F., Field, D., Sansone, S.A., Aerts, J., Apweiler, R.,
Ashburner, M., Ball, C.A., Binz, P.A., Bogue, M., Booth, T., Brazma,
A., Brinkman, R.R., Clark A. M., Deutsch, E.W., Fiehn, O., Fostel, J.,
Ghazal, P., Gibson, F., Gray, T., Grimes, G., Hancock, J.M., Hardy,
N.W., Hermjakob, H., Julian, R.K., Kane, M., Kettner, C., Kinsinger, C.,
Kolker, E., Kuiper, M., Le Novère, N., Leebens-Mack, J., Lewis, S.E.,
Lord, P., Mallon, A-M., Marthandan, N., Masuya, H., McNally, R.,
Mehrle, A., Morrison, N., Orchard, S., Quackenbush, J., Reecy, J.M.,
Robertson, D.G., Rocca-Serra, P., Rodriguez, H., Rosenfelder, H.,
Santoyo-Lopez, J., Scheuermann, R.H., Schober, D., Smith, B., Snape,
J., Stoeckert, C.J., Tipton, K., Sterk, P., Untergasser, A.,
Vandesompele, J., Wiemann, S. Promoting coherent minimum
reporting guidelines for biological and biomedical investigations: the
MIBBI project. (2008) Nature biotechnology 26, 889-896
14. Orchard, S., Salwinski, L., Kerrien, S., Montecchi-Palazzi, L.,
Oesterheld, M., Stümpflen, V., Ceol, A., Chatr-aryamontri, A.
Armstrong, J., Woollard, P., Salama, J.J., Moore, S., Wojcik, J., Bader,
G.D., Vidal, M., Cusick, M.E., Gerstein, M., Gavin, A-C., Superti-Furga,
G., Greenblatt, J., Bader, J., Uetz, P., Tyers, M., Legrain, P., Fields,
S., Mulder, N., Gilson, M., Niepmann, M., Burgoon, L., De Las Rivas,
J. , Priesto, C., Perreau, V.M., Hogue, C., Mewes, H-W., Apweiler, R.,
Xenarios, I., Eisenberg, D., Cesareni, C., Hermjakob, H. The Minimum
Information required for reporting a Molecular Interaction Experiment
(MIMIx). (2007) Nat. Biotechnol. 25:894-898