• No se han encontrado resultados

The previous display methods have all been examples of dimension re- duction; that is to say they are means by which a high dimensional data set can be shown in a lower dimensional plot such as 2- or 3-D. This has been achieved by creating new variables which are linear or non-linear combinations of the original variables. The advantage of this is that all of the information in the original variables is used but the disadvan- tage is that, in combining the variables, the ‘compression’ may obscure some useful information. To avoid this problem what is needed is some means by which all of the variables can be simultaneously displayed but without creating combinations of them. Fortunately, there are some

Figure 4.26 Non-linear map of the compound set shown in Table 4.7, points are numbered as in the table that is 1-4 Aldol-, 5-8 Claisen- and so on. Display generated using the SciFit package – www.scimetrics.com (reproduced from ref. [21] with permission of Elsevier).

ingenious techniques which can be used to display high dimensional data in lower dimensions and these make use of our natural ability to recognize patterns.

Perhaps the most obvious example of human ability to recognize pat- terns is the way that we can identify faces. This was exploited by Herman Chernoff when he devised the method that bears his name, Chernoff faces [23]. In this technique the facial characteristics of a cartoon face, such as size of ears, shape of mouth, slant of eyebrows, size of nose and so on, are assigned to each of the variables in the set. The result is a single face for each case in the set and similar cases can be rapidly identified as they

Figure 4.27 Chernoff faces display of the data in Table 4.7. The data was standard- ized (autoscaled, see Section 3.3) and the cases are shown in the order 1–5, left to right top row, 6–10 l-r second row and so on. Display generated using the statistics package Systat (www.systat.com).

have similar faces. Figure 4.27 shows the data in Table 4.7 displayed as Chernoff faces.

The Aldol- and Claisen-type disconnections are quite clearly distin- guished from the other compounds although they are not so easily sep- arated from one another. The Michael-type disconnections have char- acteristic large noses and the enamines are shown up as having straight mouths, eyebrows and smaller noses than the Michael-type. This type of display can work quite well although it becomes difficult to use for reasonably large numbers of samples.

Chernoff faces are a particular example of a type of plot known as icon plots. Icon plots use a geometric shape or object, the icon, which has sufficient features whose characteristics can be altered so that the variables can be displayed. A popular icon plot is the star plot as shown in Figure 4.28.

The star plot operates by assigning individual ‘rays’ of a symbolic representation of a star to each of the variables. The length of a ray for an individual case represents the magnitude of that variable for the case relative to the maximum magnitude of the variable across all the cases. The star plot in Figure 4.28 quite clearly shows all four classes of com- pounds and easily separates the Aldol- and Claisen-type disconnections. In this respect it has performed better than the Chernoff faces but this is really demonstrating a common property of all data display methods, whether dimension reducing or dimension preserving, and that is that there is no ‘best’ technique. All have their advantages and disadvantages and, depending on the data set, some will work better or worse than others. A final type of icon plot which is quite popular is the flower plot. A flower plot is constructed by assigning variables to ‘petals’ arranged on a circle and, as in the case of other icon plots, each case has a single flower plot symbol. Negative values of variables are shown as the petals going inside the circle and positive going out from the circle. Figure 4.29 shows the now familiar data from Table 4.7 displayed as flower plots where, once again, all four types of compounds can be distinguished.

4.5 SUMMARY

Multivariate display methods are very useful techniques for the inspec- tion of high-dimensional data sets. They allow us to examine the re- lationships between points (compounds, samples, etc.) in both training and test sets, and between descriptor variables. Dimension reduction can be achieved using linear and non-linear methods, both with advantages and disadvantages, and this has proved useful in numerous scientific applications. The linear approach (PCA) forms the basis of a variety of multivariate techniques as described later in this book. Other tech- niques for the display of multidimensional data in fewer dimensions, but without recourse to combinations of the original data, can be useful for moderately sized data sets. Finally, it is not possible to say in advance which, if any, is the best approach to use.

Figure 4.28 Star plot display of the data in Table 4.7. The data was standardized and case labels are shown as R1, R2 and so on. The plot was created using the SciFit package – www.scimetrics.com .

Figure 4.29 Flower plot display of the data in Table 4.7. The data was standardized and case labels are shown as R1, R2 and so on. The plot was created using the SciFit package – www.scimetrics.com .

In this chapter the following points were covered:

1. how the selection of appropriate variables can give the most useful ‘view’ of a data set;

2. the need for the inclusion of more variables in displays of multi- variate data;

3. the way that principal components analysis works and the meaning of scores and loadings;

4. the interpretation of scores plots and loadings plots;

5. how to rotate principal components and the effects of such rota- tions;

6. non-linear methods for reducing the dimensionality of a data set in order to display it in lower dimensions;

7. how artificial neural networks can be used to produce low dimen- sional plots of a multivariate data set;

8. what icon plots are and how they can be used to display multiple variables without data compression.

REFERENCES

[1] Lewi, P.J. (1986). European Journal of Medicinal Chemistry, 21, 155–62. [2] Seal, H. (1968). Multivariate Analysis for Biologists. Methuen, London.

[3] Hudson, B.D., Livingstone, D.J., and Rahr, E. (1989). Journal of Computer-aided

Molecular Design, 3, 55–65.

[4] Dizy, M., Martin-Alvarez, P.J., Cabezudo, M.D., and Polo, M.C. (1992). Journal of

the Science of Food and Agriculture, 60, 47–53.

[5] Ghauri, F.Y., Blackledge, C.A., Glen, R.C., et al. (1992). Biochemical Pharmacology,

44, 1935–46.

[6] Van de Waterbeemd, H., El Tayar, N., Carrupt, P.-A., and Testa, B. (1989). Journal

of Computer-aided Molecular Design, 3, 111–32.

[7] Jackson, J.E. (1991). A User’s Guide to Principal Components, pp. 155–72. John Wiley & Sons, Inc., New York.

[8] Livingstone, D.J. (1991). Pattern recognition methods in rational drug design. In

Molecular Design and Modelling: Concepts and Applications, Part B, Methods in Enzymology, Vol. 203 (ed. J.J. Langone), pp. 613–38. Academic Press, San Diego.

[9] Lin, J.C.C., Nagy, S., and Klim, M. (1993). Food Chemistry, 47, 235–45. [10] Gabriel, K.R. (1971). Biometrika, 58, 453–67.

[11] Shepard, R.N. (1962). Psychometrika, 27, 125–39, 219–46. [12] Kruskal, J.B. (1964). Psychometrika, 29, 1–27.

[13] Sammon, J.W. (1969). IEEE Transactions on Computers, C-18, 401–9.

[14] Kowalski, B.R. and Bender, C.F. (1973). Journal of the American Chemical Society,

[15] Digby, P.G.N. and Kempton, R.A. (1987). Multivariate Analysis of Ecological Com-

munities, pp. 19–22. Chapman & Hall, London.

[16] Clare, B.W. (1990). Journal of Medicinal Chemistry, 33, 687–702.

[17] Varmuza, K. (1980). Pattern Recognition in Chemistry, pp. 106–9. Springer-Verlag, New York.

[18] Kohonen, T. (1990). Proceedings of the IEEE, 78, 1464–80. [19] Zupan, J. (1994). Acta Chimica Slovenica, 41, 327–52.

[20] Bienfait, B. (1994). Journal of Chemical Information and Computer Science, 34, 890–8.

[21] Livingstone, D.J. (1996). Multivariate data display using neural networks. In Neural

Networks in QSAR and Drug Design (ed. J. Devillers), pp. 157–76. Academic Press,

London.

[22] Oja, E. and Kaski, S. (eds) (1999). Kohonen Maps, Elsevier, Amsterdam. [23] Chernoff, H. (1973). Journal of American Statistical Association, 68, 361–8.

5

Documento similar