• No se han encontrado resultados

Funciones de la Logística

In document La logística en las empresas virtuales (página 132-140)

Capítulo 2: Distribución comercial: un enfoque estratégico de las Nuevas Tecnologías

2.5 Transporte y distribución: introducción a la cadena logística

2.5.2 Funciones de la Logística

The first problem we have to deal with has to do with the period mismatch, since the PMN data is quarterly and the PSLM data is biennial. One simple approach for filling in the “missing” PSLM quarters is linear interpolation, which uses the first and last

observed values between each missing period to make estimations. But linear interpolation presumes that the data moves in a smooth line, which is not a realistic assumption for the PSLM, which has a large degree of variation. Thus, a more

sophisticated imputation technique has to be used here. The method I choose is referred to as single imputation employing stochastic regression.

Generally speaking, imputation creates a balanced dataset in a way that minimizes bias. It does so by running a multivariate regression to produce conditional group means,

medians or modes. Single imputation’s main limitation is that it assumes that imputed values are the same as observed values, which lowers standard errors and produces generous levels of significance, also referred to as over-fitting the data. In order to overcome this problem I artificially add random noise to the imputed series after running the regressions. The specific steps in the imputation process are detailed below.

D. (ii) Single Imputation for Missing Values

In the first stage, a regression model is developed with each socioeconomic indicator as the dependent variable and time and district serving as independent variables:

68

SEI = α + β1time + β2district + ɛ Equation 2

Where, SEI is the socioeconomic indicator,

α the intercept term and ɛ

the error term. β1

and β2 are the coefficients of the independent variables. Missing data is imputed using

mean values for each socioeconomic indicator from the regression. As mentioned above, using just the mean values to estimate the missing data will result in very high

significance levels for all parameter estimates.

To overcome this, artificial noise or randomness is added to the mean by randomly generating two continuous standard uniform distributions that vary between zero and one. The values generated from this distribution are also referred to as pseudo-random

numbers. The distribution itself is stationary since its mean and variance do not vary overtime. This means that the series does not have a trend, that is, the past and present values of the series are not correlated, a phenomenon that would be called autocorrelation (Chatfield, 2004). This is an important attribute for a time series that is to be used for statistical modeling. Most statistical models require stationary series, and when this assumption is violated there is danger that the regression will show spurious correlations. By using the uniform distribution to create imputed values, we allow for the introduction of artificial noise, which reduces the problem of non-stationarity and over-fitting of the data.

This involves taking the standard errors of time and district and multiplying them by the uniform distribution series and finally, adding them to the mean values estimated from the regression model. These steps are repeated for each of the 11 PSLM indicators included in this analysis.

69 D. (iii) Time Series Differencing

While it can be assumed that we have adequately de-trended the rest of the

socioeconomic variables by using the techniques described above, this may not be enough for the economic perceptions indicator. Recall that the economic perceptions indicator is correlated with the previous period, since it asks respondents to compare their present economic situation to the previous year. This results in the problem of

autocorrelation. Therefore, it is reasonable to assume that for this indicator we need to go one step further in order to make it a stationary series.

One way to do this is to transform the indicator into a period-to-period difference series. If the series is first-differenced the change is calculated from one period to the next. Thus, if t1 denotes period one and t2 denotes period two, the first difference will be the

difference in economic perceptions from t1 to t2. To test whether or not the series still has

the problem of autocorrelation, I use the Levin-Lin-Chu test for unit root testing with panel data. The test confirms that the first-difference economic perceptions series is now stationary (output reproduced in Appendix A).

The next step in the process is principal component analysis.

D. (i) Principal Component Analysis (PCA)

PCA is a data-reducing technique used specifically when correlations between variables are suspected, making it redundant to use each variable separately. There is a high likelihood that the present set of indicators are correlated, for instance it is unlikely that the literacy rate is mutually exclusive of the percentage of women that ever attended

70 school.

The PCA is also effective in reducing the observed number of variables to a smaller set of artificial indicators, called “principal components”. The PCA estimation equation is represented below:

Xi = ai1 F1 + ai12F2 + ……… + aij + Fj Equation 1

Where,

Xi is the ith development indicator,

aij is the factor loading and represents the proportion of variation in Xi accounted for by

factor j.

And, Fj represents the jth factor.

A statistical program, such as Stata, can be used to calculate a factor loaded composite index for the indicators. The first few factors account for the maximum variation.

Eigenvalues measure the amount of variation observed in the actual data that is explained by each principal component. For this paper, I use factors that have eigenvalues of at least 0.5, that is, factors that explain at least 50 percent of the variation in the underlying indicators. I create three sets of principal components or factors: one for education, another for health and a third for availability of housing utilities.

For education, the first factor alone accounts for 82 percent of the variation across all three educational indicators and is the only factor with an eigenvalue greater than 0.5. The factor loadings for this factor, referred to in the subsequent analysis simply as “education”, are provided below:

71

Table 5 Factor Loadings for Education

Education Indicators Factor

Loading

Literacy rate (age: 10 and up) 0.9543 Primary enrollment rate (ages: 5-9) 0.8772 Percentage of females ever attended

school

0.8834

The table indicates that 95.43 percent of the variation in the education factor is accounted for by the literacy rate, 87.72 percent by the primary enrollment rate and 88.34 percent by the female school attendance rate.

For health, I use two factors that together account for 92 percent of the variation in health status, these two factors each have eigenvalues greater than 0.5.

 

Table 6 Factor Loadings for Health

Health Indicators Factor Loadings

Women Children Percentage children fully

immunized

0.1983 0.9599 Percentage pregnant women

receiving tetanus shot 0.6941 0.6116 Percentage women

delivering baby at private or government clinic

0.9556 0.1783

We can see from Table 6 that the first factor titled “women” is defined by women’s pregnancy and delivery related variables, while the second is most strongly associated with immunization rates for children.

For housing utilities, I use the first three factors that together account for 88 percent of the variation in the data and with eigenvalues greater than 0.5 in each case. The factor loadings are represented below:

72

Table 7 Factor Loadings for Housing Utilities

Housing Utility Indicators

Factor Loadings

Toilet & electricity

Cooking fuel Tap water Toilet 0.6432 0.5386 0.0164 Electricity 0.9240 0.1849 0.0834 Cooking – Oil and Gas 0.2103 0.9361 0.1227 Tap water 0.0587 0.0931 0.9927

Table 7 indicates that the first factor, “toilet and electricity” is defined by the availability of toilets and electrification rates; the second, “cooking fuel” by availability of oil and gas for cooking fuel; and the third, by availability of tap water for household use. Finally, we move to estimating the regression model.

In document La logística en las empresas virtuales (página 132-140)