5 INTERVENCIÓN DEL FACTOR DE RIESGO QUÍMICO PARA PREVENCIÓN
5.5 DISPOSICIÓN FINAL DE PRODUCTOS QUÍMICOS
5.6.2 Preparación Ante Cualquiera De Las Cuatro Emergencias Químicas.
In high-dimensional problems, traditional variable selection methods are no longer sufficient, mainly due to the large value of 𝑝. Therefore, modern procedures such as the least absolute shrinkage and selection operator (Tibshirani, 1996), forward stagewise regression (Hastie et al., 2001) and least angle regression (Efron et al., 2004) were developed to improve stability and predictions (Hesterberg et al., 2008).
This section briefly describes three modern model-building algorithms, namely: the least absolute shrinkage and selection operator (LASSO), least angle regression (LAR) and forward stagewise selection. These three procedures all employ similar strategies, but differ according to the type of direction and the active set of variables used.
3.2.1 THE LASSO
The LASSO is an example of a shrinkage procedure (ridge regression is another familiar example). In such a procedure, a least squares model is fit to all 𝑝 available predictors, however the estimated (least squares) coefficients are shrunk towards zero. Depending on the type of shrinkage performed, some of the estimated least squares coefficients may be shrunk to exactly zero, thereby effectively performing variable selection.
The LASSO is a more recent shrinkage method than ridge regression and it performs continuous shrinkage of the regression coefficients with eventual variable selection. The LASSO uses an ℓ1 penalty which simultaneously performs shrinkage towards zero as well as variable selection. The latter LASSO characteristic stands in contrast to ridge regression, which always includes all 𝑝 predictors in the final model. The LASSO coefficient estimates are defined by the following equation:
𝜷̂𝑙𝑎𝑠𝑠𝑜 = 𝑎𝑟𝑔𝑚𝑖𝑛 𝜷 { 1 2∑ (𝑦𝑖 − 𝛽0− ∑ 𝑥𝑖𝑗𝛽𝑗 𝑝 𝑗=1 ) 2 𝑁 𝑖=1 + 𝜆 ∑|𝛽𝑗| 𝑝 𝑗=1 }.
Here 𝜆 ≥ 0 is a complexity parameter which penalizes the size of the coefficients and thus controls the amount of coefficient shrinkage. The ℓ1 LASSO penalty is
ℓ1 = ∑𝑝𝑗=1|𝛽𝑗| and including this term into the criterion that has to be optimised causes the
to the intercept in the model. In the LASSO the ℓ1 penalty has the effect of shrinking some of the coefficient estimates to be exactly equal to zero when the complexity parameter is sufficiently large. The shrinkage performed by the LASSO is called “soft thresholding”. The LASSO is similar to best-subset selection in the sense that it performs variable selection. However, best-subset selection drops all predictor variables with absolute coefficients smaller than the 𝑀𝑡ℎ largest; this is a form of “hard thresholding” (Hastie et al., 2009:69). A potential disadvantage of the LASSO in high-dimensional problems is that the number of variables retained in the LASSO model is bounded by 𝑚𝑖𝑛{𝑝, 𝑁 − 1}. So, if 𝑝 is much larger than 𝑁 it may happen that the LASSO selects too few variables.
The soft thresholding that forms the basis of the LASSO has links to NSC which was discussed in Section 2.5.
3.2.2 LAR
After its introduction in 1996, the LASSO did not immediately become popular. This was mainly because its computation involved a time-consuming non-linear optimisation approach. Least angle regression, introduced in 2004 by Efron et al., provided an extremely efficient algorithm for computing the entire LASSO solution (hence, for all possible values of the complexity parameter) in regression problems. LAR can be viewed as a “more democratic” (less greedy) version of forward stepwise regression (Efron et al., 2004). It uses a strategy similar to that of forward stepwise regression, which is explained above, but only enters “as much” of a predictor as it deserves which explains why it is viewed as less greedy.
The LAR algorithm consists of five steps and is explained by Hastie et al. (2009:73-74) as follows: The first step in the LAR algorithm is to standardise the predictor variables to have mean zero and unit norm. Start with the current residual, 𝑟 = 𝑦 − 𝑦̅, 𝛽1 = 𝛽2 = ⋯ =
𝛽𝑝 =0. The next step is to identify the predictor 𝑋𝑗, 𝑗 = 1, … , 𝑝 most correlated with 𝑟, and hence the current response. Instead of fitting 𝑋𝑗 completely, LAR moves 𝛽𝑗 continuously
from zero towards its least squares value causing its correlation with the evolving residual to decrease in absolute value. As soon as another variable, say 𝑋ℎ, “catches up” in terms of correlation with the current residual, the process is paused. Therefore, Step 3 in the LAR algorithm is to increase 𝛽𝑗 from 0 in the direction of its least-squares coefficient 〈𝑋𝑗, 𝑟〉, until some other competitor 𝑋ℎ , ℎ ≠ 𝑗 has as much correlation with the current residual as
does 𝑋𝑗. In Step 4, the second variable 𝑋ℎ joins the active set and 𝛽̂𝑗 and 𝛽̂ℎ are moved together in the direction defined by their joint least squares coefficients of the current residual on (𝑋𝑗, 𝑋ℎ). This is done until some other competitor 𝑋𝑙 has as much correlation
with the current residual. This process is continued until all 𝑝 predictors have entered the model. At this point 𝑐𝑜𝑟𝑟(𝑟, 𝑥𝑗) = 0 ∀𝑗 and the full least squares solution has been
reached.
Note that if 𝑝 > 𝑁 − 1, then after min(𝑁 − 1, 𝑝) = 𝑁 − 1 steps the LAR algorithm reaches a zero-residual solution and arrives at the full least-squares solution.
The paper by Hesterberg et al. (2008) describes three remarkable properties of LAR. The first property is the computational efficiency of the LAR algorithm: Efron et al. (2004) state that the entire sequence of LAR steps with 𝑝 < 𝑁 variables requires the same order of computations as an ordinary least squares fit on 𝑝 variables.
The second property is that a simple modification of the basic LAR algorithm can be used as an efficient and relatively simple way to fit the LASSO and the stagewise model, with certain modifications in higher dimensions (Efron et al., 2004).
The third and final remarkable property of the LAR algorithm is the availability of a simple 𝐶𝑝 statistic for choosing the number of steps in the LAR algorithm, namely
𝐶𝑝 = 1 𝜎̂2∑(𝑦𝑖 − 𝑦̂𝑖)2− 𝑁 + 2𝑀, 𝑛 𝑖=1 (3.1)
where 𝑀 is the number of steps and 𝜎̂2 is the estimated residual variance. This property
is based on Theorem 3 in Efron et al. (2004:424) which states that after 𝑀 steps in the LAR algorithm, the degrees of freedom of a LAR estimate is given by
∑cov(𝜇̂, 𝑌𝑖) 𝜎2 , 𝑛
𝑖=1
which is approximately equal to 𝑀. This provides a simple stopping rule for LAR, namely to stop after the number of steps 𝑀 that minimises the 𝐶𝑝 statistic in (3.1).
3.2.3 Forward-stagewise Selection
Forward-stagewise selection is a variable selection technique that is more constrained than forward stepwise regression (Hastie et al., 2009:60). Efron et al. (2004) provide a detailed description of the algorithm for forward-stagewise regression. The forward-
stagewise selection algorithm starts, like forward stepwise selection, with just the intercept in the model. The first step is once again to determine the correlation between each of the predictors and the response. Then the predictor with the largest absolute correlation is selected. Suppose this is variable 𝑋𝑗. Simple linear regression of 𝑦 on 𝑋𝑗
is performed to compute a residual (which is now considered the response). The coefficient from this regression is added to the current coefficient in the model of the chosen variable 𝑋𝑗. The process is repeated, identifying at each step the variable with
the largest absolute correlation with the current residual, 𝑟. At each step the simple linear regression coefficient of the selected variables is computed and this value is then added to the current coefficient for that variable, which may or may not be zero (Hastie et al., 2009:60). The process is repeated until none of the variables are correlated with the residuals, at which point the full least squares solution has been reached. Forward- stagewise selection is referred to as “slow fitting” as it requires more than 𝑝 steps to achieve the least squares fit, making it inefficient and problematic in high-dimensional cases (Hastie et al., 2009:60).