Survey
* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
Minimum Information About a Bioactive Entity (MIABE) Version 0.4 Sandra Orchard Henning Hermjakob Input/support to date Steve Bryant NCBI Dominic Clark EMBL-EBI Ian Dix AstraZeneca Ola Engkvist AstraZeneca Mark Forster Syngenta Michael Gilson BindingDB Martin Grigorov Nestlé Kim Hammond-Kosack Rothampsted Lee Harland Pfizer Andrew Hopkins U. Dundee Christopher Larminie GSK Elena Lo Piparo Nestlé John Overington EMBL-EBI Chris Southern EMBL-EBI Christoph Steinbeck EMBL-EBI Janet Thornton EMBL-EBI David Wishart DrugBank Feedback to [email protected] Introduction The process of the identification and development of molecules with useful bioactive properties, such as pharmaceuticals and pesticides, is fraught with difficulty and many compounds will fall by the wayside on the road from New Chemical Entity to licensed product. In the pharmaceutical industry only a very small percentage of Investigational New Drugs will make it through to clinical usage. The causes for this high attrition rate are many, with lack of efficacy, unexpected drug side effects, and undesirable drug-drug interactions being just some of the more common pitfalls in the drug discovery process. Similarly, in the world of pesticides, compounds which prove to be active show undesirable side-effects against organisms other than their original target will not make it through to the market. However, published reports of the activities of these ‘failed’ compounds, in addition to detailed information on those which go on to become fully licensed, commercially-available bioactive entities, are crucial for an understanding of how improved molecules may be developed. Details of their molecular structure and mechanism of action may give clues as to how related analogues may be developed that hit the same target, but with improved effectiveness or increased selectivity for a specific target over closely related molecules. A full disclosure of observed toxicity or an understanding of the pharmacokinetic properties of an agent may assist in improving these properties in subsequent generations of molecules. Even those molecules which fell by the wayside at an early stage in the development process may have a role to play as tools, to enable the verification of potential new targets which may have been identified by micro-array or proteomic studies in diseased tissues. In 2002, Hopkins and Groom introduced the concept of the ‘Druggable Genome’ [1] suggesting that some proteins/protein families in the (human) genome are more amenable to modulation by small exogenous molecules than others. Since that time, work has been directed at finding ways to develop a computational approach to calculate druggability and to predict ‘druggable’ proteins. In 2006 Overington et al estimated the number of molecular targets for approved drugs to be as low as 266 protein targets [2], out of the 20,400 protein coding genes [3] in the human genome. A further 58 targets are either from pathogenic organisms or are non-protein molecules and similar models may be developed to look at the susceptibility of plant, parasitic, bacterial or viral genomes to xenobiotics. This apparent very low percentage of success may be significantly increased by adding to the available data on licensed drugs information which lies in the literature and in databases held by commercial companies, of molecular targets which have proven to be successfully modulated but the agents have subsequently failed to reached the end of the discovery pipeline due to unfavourable pharmacokinetics, or toxicity linked to molecule class rather than target. As a consequence of the current productivity crisis, the pharmaceutical and biotechnology industries are increasingly disposed towards pre-competitive release of compound-related bioactivity data into public domain repositories. It has becoming widely acknowledged that the production of such an aggregated resource allows a combinatorial increase in the value of the information, in that the total amount of information that can be mined from it will be of much more valuable than from a single isolated collection. Increasingly, companies are recognising that regarding such an exercise as a pre-competitive activity allows not only a general benefit to the fields of human health and welfare but also a commercial advantage, in that the data need only be collected once. The manual harvesting and curation of data is an expensive process and involves resources often not available in even the largest pharmaceutical company. Additionally, any one company has only a limited range of in-house products targeting an equally restricted number of targets. By adding to these the knowledge contributed by other groups in both the commercial and academic sectors, a broader understanding of the druggability of a potential target protein may be reached, well before expensive resources are committed to an active research program. This precompetitive activity has wider benefits, for example in the fields of academic or non-profit development of drugs for orphan and tropical diseases. However, in order to fully understand the context, methods, data and conclusions that pertain to the published description of an experiment, detailed background information needs to be included. The diversity of experimental designs and analytical techniques is becoming more of a problem, not only as new methods supplant old ones but also as the scale of data production increases due to the widespread application of automation. While the problem is unlikely to be completely soluble it can be significantly improved by the specification of this metadata (‘data about the data’). By being associated with the results this metadata makes both the biological and methodological contexts of the experiment explicit. The archetype of such a specification is the Minimum Information about a Microarray Experiment (MIAME) [4]. Many journals and funding agencies now require that authors reporting microarray-based transcriptomics experiments comply with the MIAME checklist as a prerequisite for publication. The adoption and development of such specifications has had a much broader impact beyond merely increasing the comprehension and comparability of journal articles, the most important of which is facilitating the transfer of data from journal articles into databases (i.e. converting unstructured to structured data) in a form that will allow mining across combined data sets. We therefore propose a new such document, the Minimum Information About a Bioactive Entity (MIABE) that is predominantly concerned with, but is not restricted to, bioactive chemical compounds. We believe the timing of this is apposite for a number of reasons. The first of these is the revolution in bioactive chemical information catalysed by the appearance of ChEBI [5] and PubChem [6] towards the end of 2004. When the ”missing entity” of chemical structure was embedded within the global Web of bioinformatic relationships it became possible to search across biological effects, protein names, sequence data, and chemical information. The second is that we now have the deposition, not just of HTS results but also other types of bioactivity screening data, directly linked to chemical structure information in public repositories such as PubChem Bioassay [6], ChemBank [7], and, in the near future, ChEMBL (www.ebi.ac.uk/chembl). Finally, increasing legislation is requiring that information on these compounds be more readily available. For example, on 1st June 2007, the REACH (Registration, Evaluation, Authorisation and Restriction of Chemical substances) legislation came into force across the European Union. This requires that additional information on chemicals be made available, dependent on the quantity imported or manufactured, and it has been estimated that 30,000 chemicals may need to be re-evaluated in accordance with these requirements [8]. MIABE: Principles and process. In order that the maximum benefit be derived from the publication of data on one or a series of bioactive entities, it is important that certain crucial information is included in every such paper such that the properties of these molecules are fully represented, their effects on biological systems (both positive and negative) are accurately detailed and any factors which may contribute to the activity of the molecule are stated. To this end, a group of both industry and academic groups have come together to layout a checklist of the information that is felt to be important to include with any published dataset, the minimum information required for the activity of a molecule to be both fully understood and compared with other molecules, either sharing a common chemical structure or similar mechanism of action. It is intended that these guidelines be adhered to by anyone planning on publishing such a paper and also by the implementers of resources such as a databases to hold this information and any body or institution funding such discovery work and requiring the results to be published at the end of the granting period. As with other such reporting guidelines [4], MIABE adheres to the criteria of Sufficiency (a reader should be able to understand and critically evaluate the interpretation and conclusions, and to support their experimental corroboration) and Practicability (the guidelines should not be so burdensome as to prohibit its widespread use). It should be noted that the scope of this series of documents is limited to data regarded as pre-clinical in drug studies, a discussion of the publication of clinical data is a subject more appropriate for a separate effort. It is also recognised that the list of requirements described in the documentation represents the entire path taken by a molecule from synthesis to pre-clinical development. Many compounds, which fail to fulfil one or more of the criteria necessary for the development of a lead compound, may only travel a part of this route but still be a valuable research tool so all data generated on this compound should appear in any publication on its activity. Similarly, data on a successful agent may be published in multiple papers – in this case the minimum reporting guidelines should be followed over the series of articles if this is more appropriate. Finally, it should also be remembered that, although this document targets bioactive agents, data on inactive, closely related analogues and orthologues is often of equal value, providing negative controls and information for those trying to build activity into a molecular scaffold. Such data should also be fully reported using these guidelines. The production and development of MIABE documents This parent document exists to make explicit the scope, purpose and manner of use of the MIABE guidelines that accompany it and, as such, should be stable, as the principles described should therefore remain valid for the forseeable future. Each domain within the drug discovery process will then be discussed in a modular guideline document, the content of which may be more subject to change over time as both experimental techniques and the formats for recording data develop and increase. Initial versions of all modules will be produced through the Pharmaceutical Industry Forum hosted at the European Bioinformatics Institute and made available for extensive community input prior to publication through pre-publication on Human Proteomics Organisation, Proteomics Standards Initiative (HUPO-PSI) website and documentation process [9] and also accessible via the MIBBI portal (www.mibbi.org/index.php/Projects/MIABE, see below). New documents and proposed updates to existing guidelines will be advertised on appropriate websites, discussion groups and input requested from domain experts. Each guideline document will then enter the formal peer-review process of the journal in which it is to be published. Updates to the documents will be discussed, agreed and published when necessary, but it intended to make these as infrequent as advances in technology will allow, thus providing long periods of stability to encourage adoption and implementation. Potential authors may move between the guidelines (and other related community efforts), ensuring that their data conform to those aspects of the checklists which are relevant to their publication. Data Formats It is still common practise in many publication for a compound to be described by no more than the structural representation of a pharmacophore, with the modified regions then exemplified and the resulting molecules given an identifier, often either specific to the originating source or to that particular publication. Specific rules and conventions for compound nomenclature have been developed and allow the specific reconstruction of a compound’s structure, if strictly adhered to. However, as the needs of computational chemistry become more central to the daily workings of molecular science, the need for more systematic, and computationally translatable means of describing compounds were required, and lead to the development of SMILES strings [10] and the International Chemical Identifier (InChI) [11]. Whilst it is recognised that it would be impractical to publish several hundred of these identifiers in a large paper describing the structure-activity relationships of several hundred compounds, it is becoming increasingly important that such data be made available for published small molecules, such that subsequent data capture no longer relies on individuals piecing together molecules descriptions from a disparate figure and table(s) within a publication to redraw the molecule in electronic format. It has not previously been the practice in the field of bioactive molecules to consider data exchange formats other than the published paper however as public domain data repositories become established, the requirement to exchange data between them, or for users to download non-redundant datasets in a common format will become of increasing importance. Such a common format does already exist and has been publicly available and in wide usage for some years in the molecular interaction field. The HUPO PSIMI XML2.5 interchange format is capable of capturing extensive details about many interactions types, including bioactive entities and their target molecules, including the biological role of each molecule within that interaction, detailed description of interacting domains, and the kinetic parameters of the interaction [12]. The ability to describe the structure of individual molecules and to carry meta-data on that molecule is also inherent to the format. The format is supported by data management and analysis tools and has been adopted by major interaction data providers, toll developers and used in visualisation and analytical software. Additionally, a simpler, tabdelimited format MITAB2.5 has been developed for the benefit of users who require only minimal information in an easy to access configuration. It is suggested that supporting the continued development of this format, rather than an attempt to “reinvent the wheel” will be the most practical step forward. This will allow produces of drug-target information to merge their data with the information existing the molecular interaction databases and use PSI-MI XML2.5 compliant resources, such as Cytoscape (www.cytoscape.org), to visualise small molecule data in conjunction with cellular interactomes or pathways. Controlled Vocabularies Where possible, the user is asked to make use of existing controlled vocabularies (CV), or ontologies, to describe entities, processes and conditions described within a paper. This benefits both the human reader, in that it reduces the possibility of ambiguity where a term is used which may have multiple meanings, and also assists in the process of curation of the information into data resources and subsequent search and analysis procedures. A number of such controlled vocabularies are available on the OBO website (www.obofoundry.org/) and may be used to describe, for example, tissues, disease and molecular interactions (including enzyme/substrate). The existing PSI-MI CV [12] used to annotate the format described above has already been extended to allow for a full description of the properties of a bioactive molecule and further input into the development of this resource is welcomed. MIABE in the context of related efforts As already stated, the move to standardise the reporting and subsequent collection and exchange of data is now being encouraged across the entire biomedical field as the community joins in a united front to improve both data quality and availability. In order to manage this process, the MIBBI project (http://www.mibbi.org/) has been established, which maintains a web-based, freely accessible resource for checklist projects, providing straightforward access to extant checklists (and to complementary data formats, controlled vocabularies, tools and databases) and ensures that new efforts are not overlapping or redundant to an existing resource [13]. MIBBI is managed by representatives of its various participant communities and its goal is to facilitate the development of an integrated checklist resource site for the wider bioscience community. Where MIABE can be seen to have areas of overlap with existing resources, and obvious example of this are the MIMIx molecular interaction standard [14], the MIAPAR protein affinity reagent standard (www.mibbi.org/index.php/Projects/MIAPAR) and MIACA (www.mibbi.org/index.php/Projects/MIACA), the standard for the description of cellular assays, the work of these resources has been referred to rather than repeated. Reporting requirements for a bioactive entity The minimum reporting requirements are divided into the following sections MIABE document MIABE-compound Reporting Requirement Physico-chemical properties of the molecule. Only experimental parameters are included in the minimum requirements, however if calculated parameters are given, the method by which they are calculated should be stated. MIABE-compound Molecule Source. The synthetic route by which a molecule has been produced and /or extracted from natural source. In vitro assays – mechanism of action studies, to ascertain off-target activities or drug metabolism studies. MIABE-assay MIABE-assay Cellular assays – mechanism of action studies, to ascertain off-target activities or drug metabolism studies. MIABE-assay In vivo assays – mechanism of action studies, to ascertain off-target activities or drug metabolism studies. MIABE-assay Pharmacokinetic studies – to include a full range of observations made in vivo and in vitro. MIABE-assay Toxicological data to include post-mortem and subsequent analyses. Conclusion The main aim of this document is to support the long-term bioactive molecule discovery process, by ensuring that the data generated from expensive and time-consuming previous projects is not lost but can be harvested and built on in the future. By following such guidelines, authors will produce a richly annotated dataset about the molecule, or series of molecules which they are describing, which will allow effective data capture by anyone wishing to reproduce or analyse the results. It is to be hoped that more groups will recognise the value of putting such data into the public domain to become part of an ever increasing knowledge bank on biomodulatory agents, allowing the development of improved molecules targeting an expanded druggable genome in an increasing number of species. References 1. Hopkins, A.L., Groom, C.R. (2002) The druggable genome. Nature reviews. Drug discovery 1:727-30 2. Overington, J.P., Al-Lazikani, B., Hopkins, A.L. (2006) How many drug targets are there? Nature reviews. Drug discovery [2006 (5) ] page info:993-6 3. The UniProt Consortium (2009) The Universal Protein Resource (UniProt) 2009 Nucleic Acids Res. 37, 169-174 4. Brazma, A, Hingamp P, Quackenbush, J., Sherlock, G., Spellman ,P., Stoeckert, C, Aach, J., Ansorge, W., Ball, C.A., Causton, H.C., Gaasterland, T., Glenisson, P., Holstege, F.C, Kim, I.F., Markowitz, V., Matese, J.C., Parkinson, H., Robinson, A., Sarkans, U., SchulzeKremer, S., Stewart, J., Taylor, R., Vilo, J., Vingron, M.. Minimum information about a microarray experiment (MIAME)-toward standards for microarray data. Nat. Genet. 2001, 29, 365-371 5. Degtyarenko, K., de Matos, P., Ennis, M., Hastings, J., Zbinden, M., McNaught, A., Alcantara, R., Darsow, M., Guedj, M., Ashburner. M (2008) ChEBI: a database and ontology for chemical entities of biological interest. Nucleic acids research 36, 344-350 6. Sayers EW , Barrett T , Benson DA , Bryant SH , Canese K , Chetvernin V , Church DM , Dicuccio M , Edgar R , Federhen S , Feolo M , Geer LY , Helmberg W , Kapustin Y , Landsman D., Lipman, D.J., Madden, T.L., Maglott, D.R., Miller, V., Ostell, J., Pruitt, K.D., Schuler, G.D., Shumway, M., Sequeira, E., Sherry, S.T., Sirotkin, K., Souvorov, A., Starchenko, G., Tatusov, R.L., Tatusova, T.A., Wagner, L., Yaschenko, E. Database resources of the National Center for Biotechnology Information (2008) Nucl. Acids Res. 36: 13-21 7. Seiler K.P., George, G.A., Happ, M.P., Bodycombe, N.E., Carrinski, H.A., Norton, S., Brudz, S., Sullivan, J.P., Muhlich, J., Serrano, M., Ferraiolo, P., Tolliday, N.J., Schreiber, S.L., Clemons, P.A. (2008) ChemBank: a small-molecule screening and cheminformatics resource database. Nucleic acids research (36) p351-359 8. Pederson, F., de Brujin, J., Munn, S., van Leeuwen, K. (2003) Assessment of additional testing needds under REACH. Effects of (Q)SAR, risk based testing and voluntary industry activity. (2003) European Commission report EUR 20863. 9. Vizcaino, J.A., Martens, L., Hermjakob, H ., Julian, R.K., Paton, N.W. (2007) The PSI formal document process and its implementation on the PSI website. Proteomics 7, 2355-2357 10. Weininger, D. SMILES, a Chemical Language and Information System 1. Introduction and Encoding Rules. J. Chem. Inf. Comput. Sci. 1988, 28, 31−36 11. Stein SE, Tchekhovskoi D, Heller SR. The IUPAC and NIST Chemical Identifier. 2004 12. Kerrien, S., Orchard, S., Montecchi-Palazzi, L., Aranda, B., Quinn, A.F., Vinod, N., Bader, G.D., Xenarios, I., Wojcik, J., Sherman, D., Tyers, M., Salama, J.J., Moore, S., Ceol, A., Chatr-aryamontri, A., Oesterheld, M., Stümpflen, V., Salwinski, L., Nerothin, J., Cerami, E., Cusick, M.E., Vidal, M., Gilson, M. Armstrong, J., Woollard, P., Hogue, C., Eisenberg, D., Cesareni, G., Apweiler, R., Hermjakob, H. (2007) Broadening the horizon--level 2.5 of the HUPO-PSI format for molecular interactions. BMC biology 5: 44 13. Taylor, C.F., Field, D., Sansone, S.A., Aerts, J., Apweiler, R., Ashburner, M., Ball, C.A., Binz, P.A., Bogue, M., Booth, T., Brazma, A., Brinkman, R.R., Clark A. M., Deutsch, E.W., Fiehn, O., Fostel, J., Ghazal, P., Gibson, F., Gray, T., Grimes, G., Hancock, J.M., Hardy, N.W., Hermjakob, H., Julian, R.K., Kane, M., Kettner, C., Kinsinger, C., Kolker, E., Kuiper, M., Le Novère, N., Leebens-Mack, J., Lewis, S.E., Lord, P., Mallon, A-M., Marthandan, N., Masuya, H., McNally, R., Mehrle, A., Morrison, N., Orchard, S., Quackenbush, J., Reecy, J.M., Robertson, D.G., Rocca-Serra, P., Rodriguez, H., Rosenfelder, H., Santoyo-Lopez, J., Scheuermann, R.H., Schober, D., Smith, B., Snape, J., Stoeckert, C.J., Tipton, K., Sterk, P., Untergasser, A., Vandesompele, J., Wiemann, S. Promoting coherent minimum reporting guidelines for biological and biomedical investigations: the MIBBI project. (2008) Nature biotechnology 26, 889-896 14. Orchard, S., Salwinski, L., Kerrien, S., Montecchi-Palazzi, L., Oesterheld, M., Stümpflen, V., Ceol, A., Chatr-aryamontri, A. Armstrong, J., Woollard, P., Salama, J.J., Moore, S., Wojcik, J., Bader, G.D., Vidal, M., Cusick, M.E., Gerstein, M., Gavin, A-C., Superti-Furga, G., Greenblatt, J., Bader, J., Uetz, P., Tyers, M., Legrain, P., Fields, S., Mulder, N., Gilson, M., Niepmann, M., Burgoon, L., De Las Rivas, J. , Priesto, C., Perreau, V.M., Hogue, C., Mewes, H-W., Apweiler, R., Xenarios, I., Eisenberg, D., Cesareni, C., Hermjakob, H. The Minimum Information required for reporting a Molecular Interaction Experiment (MIMIx). (2007) Nat. Biotechnol. 25:894-898