TABLA 2 GRADO DE ANEMIA EN FUNCIÓN CONCENTRACIÓN HB (OMS)
1.5 ESTRATEGIAS TERAPÉUTICAS PARA LA CORRECCIÓN DE LA ANEMIA.
Having established the effectiveness of our approach on synthetic testbeds and in silico discovery of flavorful cocktails, we proceed to evaluate its effectiveness on a focused drug design problem where we search for pharmaceuticals effective against idiopathic pulmonary fibrosis (see Section 5.5.2). We assess the performance of the algorithm from several different perspectives. First, we confirm that the approach represents a significant improvement over plain Monte Carlo search performed with the proposal generator. Following this, we quantify the learning rate of our approach by measuring how much more likely the approach is to generate desired molecular structures compared to the proposal generator as a function of the budget expended. Having assessed the algorithmic performance of the approach, we proceed to analyze the designed molecules from the perspective of medicinal chemistry. In particular, we discuss some of the designed molecular structures in the context of compounds (Adams et al., 2014) already reported in the literature.
Similar to Section 5.6.2, we evaluate the effectiveness of the approach using the correct- construction curve that shows the cumulative number of discovered molecules exhibiting the target property, i.e., a docking score lower than −11.75, as a function of the budget
0 10 20 30 40 50 0 100 200 300 400 500 round numb er of hits
(a) correct-construction curve
0 10 20 30 40 50 0. 0.6 1.2 1.8 2.4 3.0 round lift
(b) improvement over proposal generator
Figure 5.6: Panel (a) shows the cumulative numbers of hits as a function of the budget expended. The blue curve is the correct-construction curve of our approach (with corresponding confidence interval colored in light blue) and the red curve is the correct-construction curve of the proposal generator. Panel (b) indicates how much more likely it is to see a hit compared to a standard Monte Carlo search performed with the proposal generator. expended. Figure 5.6, which shows correct-construction curves for our approach (blue curve) and the described proposal generator (red curve), confirms that our approach generates more hits than Monte Carlo search with the proposal generator. Moreover, the correct- construction curve of the proposal generator is, apart from a few initial rounds, always below the lower endpoint of the confidence interval for the curve of our approach. The lift of the correct-construction curve for our approach (showed in Figure 5.6, Panel b) indicates that the approach is approximately 2.8 times more likely to output a hit than the proposal generator after 50 rounds of model calibration. The results presented in this study have been generated using simulations consuming approximately 28 hours of cpu time and running on 10 processors in parallel. This offers significant speed up over an exhaustive exploration of the search space specified by the proposal generator that would take more than 8 months of cpu time using the same number of processors. Moreover, while the simulations were relatively short (with a budget of 500 evaluations), the approach managed to discover a number of interesting compounds from the perspective of medicinal chemists.
The experimental knowledge provided to the algorithm was limited to the X-ray crystal structure of the receptor and some basic constraints on the mode of binding applied to the molecular docking. The parent compound and the possible fragments were informed, in a broad sense, by expert knowledge from medicinal chemistry, but no explicit data on the experimental activities of any compounds were used. Previous work (Adams et al., 2014) presented the synthesis and experimental assay (reported as pIC50values, i.e., the negative logarithm of the concentration required for 50% inhibition) of 30 derivatives of the parent compound shown in Figure 5.3. A pIC50 = 6.0 corresponds to a 1µM potency, and the compound might be considered active or worth further investigation. From Table 5.2, we can see that the parent compound has a pIC50 of 5.7 and would, therefore, be considered inactive. Encouragingly, 19 out of the 26 reported active compounds were found by the algorithm. Several of the compounds reported as active were not found, but two of these were not discoverable by the algorithm, because the substitution pattern (forming new ring structures) was not part of the proposal generator. There were two compounds for which the docking score was not sufficiently low, which indicates that there is an opportunity to improve the docking protocol. In total, 20 of the 30 compounds known from previous work (Adams et al., 2014) were discovered.
The described cyclic discovery process (see Section 5.5.2) is a proof of concept and still requires considerable refinement but nonetheless, from a medicinal chemistry and drug dis-
5.7 Discussion 163
covery perspective, the molecules suggested for synthesis are promising for several reasons. First, many of the molecules suggested align with the structure activity relationships (An- derson et al., 2016a,b) which were not part of the input to the algorithm, either in terms of parameters or design. An example is the algorithm predominantly suggests substitutions at the meta position, which indeed appears to be crucial for αvβ6activity. Suggested substitu-
tions also often feature heterocycles which are known to deliver αvβ6activity (Anderson
et al., 2016a,b). Secondly, most of the molecules are drug-like: that is, they resemble both the structures and physico-chemical properties of oral drugs. Thirdly and particularly promising is the speed at which new molecules can be evaluated computationally allowing several iterations to be easily carried out to improve the design quality of the molecules (as detailed earlier). Moving forward, it will be straightforward to incorporate additional constraints, such as scoring molecules against αvβ3, which should improve the selectivity window for
αvβ6over αvβ3, including polar surface area cut-offs (which correlate with several important
drug-like properties) and simple synthetic chemistry considerations.
No. Fragments Docking Score pIC50 No. Hits Compounds independently identified as hits
[4] 3-F −12.16 6.1 25
[22] 3-MeO −11.79 6.5 8
[25] 4-Me −11.96 6.1 7
[32] 3-CN −11.94 6.6 23
Compounds not discovered by the algorithm
[31] 4-Ph −11.92 6.4 -
[38] 3,4-Me2 −12.18 6.7 0
[39] 3,4,-CH2CH2CH2 −11.93 6.8 -
Compounds withpIC50≥ 7.0, but with a docking score above the threshold (−11.75)
[33] 3-CF3 −11.54 7.0 10
[43] 3-CF3-4-Cl −11.39 7.0 1
Parent compound
[15] H −10.26 5.7 0
Table 5.2: Comparison of hits identified by the approach with compounds that have been experimentally assayed (Adams et al., 2014).
5.7
Discussion
In this section, we place our work in the context of machine learning approaches closely related to ours (Section 5.7.1) and discuss some directions for future development of the cyclic discovery process characteristic to drug design (Section 5.7.2).