Having identified pairs of correlated variables, two problems remain in deciding which one of a pair to eliminate. First, is the correlation ‘real’, in other words, has the high correlation coefficient arisen due to a true correlation between the variables or, is it caused by some ‘point and cluster effect’ (see Section 6.2) due to an outlier. The best, and perhaps simplest way to test the correlation is to plot the two variables against one another; effects due to outliers will then be apparent. It is also worth considering whether the two parameters are likely to be correlated with one another. In the case of molecular structures, for example, if one descriptor is electronic and the other steric then there is no reason to expect a correlation, although one may exist, of course. On the other
hand, maximum width and molecular weight may well be correlated for a set of molecules with similar overall shape.
The second problem, having decided that a correlation is real, con- cerns the choice of which descriptor to eliminate. One approach to this problem is to delete those features which have the highest number of correlations with other features. This results in a data matrix in which the maximum number of parameters has been retained but in which the inter-parameter correlations are kept low. Another way in which this can be described is to say that the correlation structure of the data set has been simplified. An alternative approach, where the major aim is to reduce the overall size of a data set, is to retain those features which correlate with a large number of others and to remove the correlating descriptors.
Which of these two approaches is adopted depends not only on the data set but also on any knowledge that the investigator has concerning the samples, the dependent variable(s) and the independent variables. In the case of molecular design it may be desirable to retain some par- ticular descriptor or group of descriptors on the basis of mechanistic information or hypothesis. It may also be desirable to retain a descriptor because we have confidence in our ability to predict changes to its value with changes in chemical structure; this is particularly true for some of the more ‘esoteric’ parameters calculated by computational chemistry techniques. What of the situation where there is a pair of correlated pa- rameters and each is correlated with the same number of other features? Here, the choice can be quite arbitrary but one way in which a decision can be made is to eliminate the descriptor whose distribution deviates most from normal. This is used as the basis for variable choice in a pub- lished procedure for parameter deletion called CORCHOP [1]; a flow chart for this routine is shown in Figure 3.3.
Although the methods which will be used to analyse a data set once it has been treated as described here may not depend on distributional assumptions, deviation from normality is a reasonable criterion to apply. Interestingly, some techniques of data analysis such as PLS (see Chap- ter 7) depend on the correlation structure in a data set and may appear to work better if the data is not pre-treated to remove correlations. For ease of interpretation, and generally for ease of subsequent handling, it is recommended that at least the very high correlations are removed from a data matrix.
Another source of redundancy in a data set, which may be more diffi- cult to identify, is where a variable is correlated with a linear combination of two or more of the other variables in the set. This situation is known
Figure 3.3 Flow diagram for the correlation reduction procedure CORCHOP (reproduced from ref. [2] with permission of Wiley-VCH).
as multicollinearity and may be used as a criterion for removing vari- ables from a data set as part of data pre-treatment. An example of the use of multicollinearity for variable selection is seen in a procedure [2] called UFS (Unsupervised Forward Selection) which is available from the website of the Centre for Molecular Design (www.cmd.port.ac.uk). UFS
constructs a dataset by selecting variables with low multicollinearity in the following way:
r The first step is the elimination of variables that have a standard deviation below some assigned lower limit.
r The algorithm then computes a correlation matrix for the remaining set of variables and chooses the pair of variables with the lowest correlation.
r Correlations between the rest of the variables and these two chosen descriptors are examined and any that exceed some pre-set limit are eliminated.
r Multiple correlations between each of the remaining variables and the two selected ones are examined and the variable with the lowest multiple correlation is chosen.
r The next step is to examine multiple correlations between the re- maining variables and the three selected variables and to select the descriptor with the lowest multiple correlation.
This process continues until some predetermined multiple correlation coefficient limit is reached. The results of the application of CORCHOP and UFS can be quite different as the former only considers pairwise correlations. The aim of the CORCHOP process is to simplify the cor- relation structure of the data set while retaining the largest number of descriptors. The aim of the UFS procedure is to produce a much simpli- fied data set in which both pairwise and multiple correlations have been reduced.
It is desirable to remove multicollinearity from data sets since this can have adverse effects on the results given by some analytical methods, such as regression analysis (Chapter 6). Factor analysis (Chapter 5) is one method which can be used to identify multicollinearity. Finally, a note of caution needs to be sounded concerning the removal of descriptors based on their correlation with other parameters. It is important to know which variables were discarded because of correlations with others and, if possible, it is best to retain the original starting data set. This may seem like contrary advice since the whole of this chapter has dealt with the matter of simplifying data sets and removing redundant information. However, consider the situation where two variables have a correlation coefficient of 0.7. This represents a shared variance of just under 50 %, in other words each variable describes just about half of the information in the other, and this might be a good correlation coefficient cut-off limit for removing variables. Now the correlation coefficient between two
Figure 3.4 Illustration of the geometric relationship between vectors and correla- tion coefficients (reproduced from ref. [2] with permission of Wiley-VCH).
parameters also represents the angle between them if they are considered as vectors, as shown in Figure 3.4.
A correlation coefficient of 0.7 is equivalent to an angle of approxi- mately 45◦. If one of the pair of variables is correlated with a dependent variable with a correlation coefficient of 0.7 this may well be very useful in the description of the property that we are interested in. If the vari- able that is retained in the data set from that pair is one that correlates with the dependent (X1 in Figure 3.4) then all is well. If, however, X1
was discarded and X2 retained then this parameter may now be com-
pletely uncorrelated (θ = 90◦) with the dependent variable. Although this is an idealized case and perhaps unlikely to happen so disastrously in a multivariate data set, it is still a situation to be aware of. One way to approach this problem is to keep a list of all the sets of correlated variables that were in the starting set. Figure 3.5 shows a diagram of the correlations between a set of parameters before and after treatment with the CORCHOP procedure. In this figure the correlation between variables is given by the similarity scale. The correlation between LOG- PRED and Y PEAX, for example, is just over 0.4. It can be seen from the figure that there were 4 other variables with a correlation of∼0.8 with LOGPRED (shown by dotted lines in the adjacent cluster) which have been eliminated by the CORCHOP algorithm. If no satisfactory correlations with activity are found in the de-correlated set, individual variables can be re-examined using a diagram such as Figure 3.5. A list of such correlations may also assist when attempts are made to ‘explain’ correlations in terms of mechanism or chemical features.
Figure 3.5 Dendrogram showing the physicochemical descriptors (for a set of anti- malarials) retained after use of the CORCHOP procedure. Dotted lines indicate parameters that were present in the starting set (reproduced from ref. [3] with permission of Wiley-Blackwell).
A recent review treats the matter of variable selection in some detail [4].
3.7 SUMMARY
Selection of the analytical tools, described in later chapters, which will be used to investigate a set of data should not be dictated by the availability of software on a favourite computer, by what is the current trend, or by personal preference, but rather by the nature of the data within the set.
The statistical distribution of the data should also be considered, both when selecting analytical methods to use and when attempting to inter- pret the results of any analysis. The first stage in data analysis, however, is a careful examination of the data set. It is important to be aware of the scales of measurement of the variables and the properties of their distributions (Chapter 1). The cases (samples, objects, compounds, etc.) need selection for the establishment of training and test sets (Chapter 2) although this may have been done at the outset before the data set was collected. Finally, the variables need examination so that ill-conditioned variables can be removed, missing values identified and treated, redun- dancy reduced and possibly variables selected. All of this is known as pre-treatment and is necessary in order to give the data analysis methods a good chance of success in extracting information.
In this chapter the following points were covered:
1. how an examination of the distribution of variables allows the identification of variables to remove;
2. the need for scaling and examples of scaling methods; 3. ways to treat missing data;
4. the meaning and importance of correlations between variables both simple and multiple;
5. the reasons for redundancy in a data set; 6. the process of data reduction;
7. procedures for variable selection which retain the maximum num- ber of variables in the set or which result in a data set containing minimum multicollinearity.
REFERENCES
[1] Livingstone, D.J. and Rahr, E. (1989). Quantitative Structure–Activity Relationships,
8, 103–8.
[2] Whitley, D.C., Ford, M.G. and Livingstone, D.J. (2000). Journal of Chemical Infor-
mation and Computer Science, 40, 1160–1168.
[3] Livingstone, D.J. (1989). Pesticide Science, 27, 287–304.
[4] Livingstone, D.J. and Salt, D.W. (2005). Variable Selection – Spoilt for Choice? in
Reviews in Computational Chemistry, Vol 21, K. Lipkowitz, R. Larter and T.R.