Sunday, 29 September 2013
Congratulations to Ben Stauch PhD!
Ben Stauch in the group has just been examined on his these - 'Methods for the Investigation of Protein-Ligand Complexes'. This was a tour de force of many techniques - NMR, computational and X-ray crystallography. Ben will be around for a few more months, writing things up, and completing/starting some experimental work on Xe complex refinement and characterisation.
Congratulations to Ben from all the group!
In due course, the thesis will be downloadable from the EBI and EMBL websites, and I'll update this post when the files are there.
jpo
Friday, 27 September 2013
Team ChEMBL in Action
We usually blog about exciting scientific and technological updates, interesting concepts, ideas and publications within the realm of life sciences and drug discovery.
This post is slightly different, as it deals with something that might be (even) more important:
A number of us in the ChEMBL Group (Rita (not in the picture above), Patricia, Felix, Anna, Sam, Anne, Mark, Michal, George, Gerard and Ashwini) are doing a Fun Run at Victoria Park on 12th October to help raise money for Cancer Research UK. We are doing this to support a colleague who is currently receiving treatment for cancer.
We've set up a JustGiving page which makes donations fast, easy and secure.
Anything you can donate (in almost any currency :)) to this worthwhile cause would be really much appreciated.
The ChEMBL Group
Thursday, 26 September 2013
Document Similarity in ChEMBL - 2
Following up on yesterdays post by George and Mark, I put together a slide, hopefully illustrating the advantages of document comparison using objects other than words alone.
jpo
Wednesday, 25 September 2013
Document Similarity in ChEMBL - 1
Many of you will have noticed a new section on the ChEMBL interface, specifically at the Document Report Card page, called Related Documents. It consists of a table listing the links for up to 5 other ChEMBL documents (i.e. publications aka papers) that are scored to be the most similar to the one featured in the report card. Here's an example.
How does this work? There are examples of related documents sections online, e.g. in PubMed or in various journal publishers' websites. Document 'related-ness' or similarity can be assessed by comparing MeSH keywords or by clustering documents using TF-IDF weighted term vectors. Fortunately, ChEMBL puts a lot of effort in manually extracting and curating the compounds and biological targets from publications, so why not using these as descriptors to assess document similarity instead - as far as we know this is the first time this approach has been implemented?
So, here's how it works:
Firstly, for each document in ChEMBL, its list of references is retrieved using the excellent EuropePMC web services. By considering documents as nodes which are connected with an edge if one paper cites the other, a directed graph structure emerges. By doing this for all ~50K documents in ChEMBL, you get the massive graph illustrated above in Cytoscape. As a bonus, by measuring the in- and out- degree of the nodes, one could check which are the most cited papers in ChEMBL - but that's the topic of another blog post. This graph could be further annotated with protein target families, authors and institutions, as it has been elegantly done here.
Moving on, once a relationship between two documents is established, we need a way to quantify their similarity. As hinted above, we used the normalised overlap of compounds and targets reported in the two documents. This is done using the classic Tanimoto coefficient, so if doc A reports compounds (1,2,3) and doc B reports compounds (3,4,5), their compound Tanimoto similarity T is 1/5 or 0.2. Exactly the same applies for the target-based document similarity. The composite score we use to rank docs in the Related Documents section is simply the maximum of the two individual ones.
What does all that mean in practice? It means that 2 papers are listed as similar if they their reported compounds or biological targets overlap significantly (and one cites the other). For example, papers with follow-up experiments on the same candidate drug will be deemed similar, e.g. this one. The same will apply to two papers that involve kinase panel screening assays. A desirable side-effect is that by following the links, the tenacious user may traverse the whole graph displayed above!
George & Mark
Tuesday, 24 September 2013
Paper: Benchmarking of protein descriptor sets in proteochemometric modeling (part 2): modeling performance of 13 amino acid descriptor sets
A paper from Gerard in the group on some of his proteochemometric modelling work; a link to the paper is here. Z-scales rule! (the original Sandberg et al J Med Chem paper on the Z-scales was one of my 'lightbulb turning on' moments in my professional life - go hunt it down if you don't know it.)
%T Benchmarking of protein descriptor sets in proteochemometric modeling (part 2: modeling performance of 13 amino acid descriptor sets
%A G.J.P. van Westen
%A R.F. Swier
%A I. Cortes-Ciriano
%A J.K. Wegner
%A J.P Overington
%A A.P. IJzerman
%A H.W.T. van Vlijmen
%A A. Bender
%J J. Cheminformatics
%D 2013
%V 5
%O doi:10.1186/1758-2946-5-42
jpo
Sunday, 22 September 2013
Antibacterial Targets - Evidence for exclusion of targets for which host orthologues exist
One of the classic mantras for the genomics-based discovery of novel anti-bacterials is to ignore targets for which orthologues exist in the host (human) genome, but my hunch was that the majority of antibacterial mechanisms have clear host orthologues (they do, as you'll see below). I haven't come across any really simple papers supporting this dogma in the past, so decided to have a quick look this morning. I am the unwilling host for an oral bacterial infection myself at the moment, but I must stress that the throat above is not mine, but an anonymous one from Teh Interweb!
So, using a book I've just picked up at the ACS in Indy, I went through and did some quick analysis - the prose in the book is great, informal, and very very readable - buy it!
%T Antibacterial Agents: Chemistry, Mode of Action, Mechanisms of Resistance and Clinical Applications
%A R.J. Anderson
%A P.W. Groundwater
%A A. Todd
%A A.J. Worsley
%I Wiley
%D 2012
%O ISBN 978-0-470-97245-8
I went through, and at a drug class level, assigned the distinct mechanisms into three target classes.
- Orthologue of antibacterial target is present in humans.
- Orthologue of antibacterial target is absent in humans.
- Antibacterial acts through a non gene-derived target mechanism.
The counts for these three states are 8, 4, 3 or in a simple graphical form
So, no real great evidence to focus on bacteria specific genes - in fact it's 2:1 in favour of targets for which orthologues exist in the host. The key would seem to be more exploiting physicochemical differences between bacterial and human cells (e.g. acidity), or exploit differences in the binding sites. I guess the dogma arose for a couple of reasons - firstly everyone knowns about penicillins, and secondly, it is an easy filter to apply bioinformatically, and finally it just seems like a perfectly sensible thing to do with respect to elimination of mechanism-based toxicity - and so is often done.
On this latter point, that of mechanism-based toxicity, this can be really important, but remember, a drug dosed to a human is not magically attracted only to relatives of the bacterial target, it will sample and equilibrate across all accessible binding sites of all proteins, and drugs will have side-effects and toxicity related to 1) binding to orthologues, 2) binding to paralogues, and 3) binding to anything else. A nice example of this is the paper from Science earlier this year on the side effects of sulphonamide antibacterials via inhibition of host sepiapterin reductase.
To be clear what I did here. Firstly - I used the chapters in the Antibacterial Agents book to define a class, so the graph can be plotted in many other ways - however, I was interested in distinct mechanisms. Secondly, several antibacterials don't target proteins, but various parts of the ribosome - these are gene derived, so count above as gene products, but they are not proteins (well they are riboproteins). Even if you strip these RNA targets out though (there are 5) the numbers are still not compelling for the need to avoid host target orthologues - the ratio would be 3:4 instead of 8:4.
What are the implications of removing this bacterial specific filter? Well that is more than a quick job on a Sunday morning and two cups of Lapsong Suchong - but my feeling is it might be quite significant.
jpo
Seminar: Ruben Abgyan - The State of Docking, Modeling and Structure Based Molecular Discovery: GPCRs and Polypharmacology
We have Ruben Abagyan from Scripps visiting this coming Friday (27th of September 2013). Ruben will be well known to the computational chemistry and structural biology communities - he is giving a couple of talks, with the talk in the morning from 11am to noon "The State of Docking, Modeling and Structure Based Molecular Discovery: GPCRs and Polypharmacology" probably being of broad interest to many.
This will be an open seminar, but I will need to register any external attendees with security prior to Friday - so if you're in the Cambridge (UK) area, and wish to attend, please let me know by the end of Thursday. No web broadcasting or similar is possible for this. Sorry.
jpo
As an aside, Ruben is an EMBL Alumnus!
UPDATE - Please note change of time!! 11am to noon - Rosalind Franklin Room
Subscribe to:
Posts (Atom)




