• No se han encontrado resultados

9. PROYECTO DE INTERVENCIÓN PEDAGÓGICA

9.3. EXPERIENCIA PEDAGÓGICA

For the five types of probability distributions described in Section 9.1, skewness and mode tests can be used in combination to uniquely identify each of the five distributions. Both positive and negative deviations from the mean contribute to the variance in the same way since the variance squares both positive and negative deviations. The skewness cubes the deviations from the mean to measure if the deviations are largely symmetric, from the right side of the mean, or from the left side of the mean, as follows:

skewness= E(xμ) 3

σ3 , (9.1)

whereμandσ are the population mean and variance, respectively. The skewness of a normal

distribution or any symmetric distribution is expected to be zero. A left skewed distribution with a long tail to the left of the mean has a negative skewness value. A right skewed distribution with

a long tail to the right of the mean has a positive skewness value. Given a data sample, x1,x2, . . .,

xn, the sample skewness is computed as follows in [1, 2]:

skewness= n

n

i=1(xi¯x)3

(n1)(n2)s3 (9.2)

where ¯x and s are the sample average and standard deviation, and n is the sample size. The skewness value is computed using Statistica [2]. Statistica computes the standard error of the skewness value. If a skewness value is greater than three times of the standard error of the skewness value, the data variable is considered to be right skewed. If a skewness value is smaller than minus three times of the standard error of the skewness value, the data variable is considered to be left skewed.

The mode in the probability density indicates clustering in the data [3]. A probability distribution can have one mode or multiple modes. For examples, a normal distribution has one mode, a skewed distribution has one mode, and a bimodal distribution has two modes. Both the DIP test [3–6] and the mode test [7] are used together to determine the modality of data because through testing the data on only one of the tests is not adequate to distinguish the unimodality and multimodality of data. The DIP test is a test of the unimodality of data as a whole [3, 4]. The DIP test is performed using the diptest package [5] for R statistical software [6] with the significance level set to 0.05. The mode test [7] is a test for significance of each individual potential mode rather than a test on the overall unimodality of data. The mode test program by Minnotte [7] is used. For each mode tested, the program results show its location along with its p-value. The significant level is set to 0.05. Based on the test results, the number of significant modes can be counted.

Through the testing on the Windows performance objects data, the results of the skewness test, the DIP test and the mode test, which indicate the five types of probability distributions, are obtained and summarized in Table 9.1. For example, a data variable is considered to have a uniform distribution if:

Skewness and mode tests to identify five types of probability distributions 147 Table 9.1 The skewness and mode test results used to identify the five types of probability distribution

Probability distribution

(Acronym) DIP test Mode test Skewness test

1. Multimodal distribution (DMM)

Reject the unimodality Any result Any result

2. Uniform distribution (DUF)

Not reject the unimodality Number of significant modes>2

Symmetric

3. Unimodal, symmetric distribution (DUS)

Not reject the unimodality Number of significant modes<2

Symmetric

4. Unimodal, left skewed distribution (DUL)

Not reject the unimodality Number of significant modes<2

Left skewed

5. Unimodal, right skewed distribution (DUR)

Not reject the unimodality Number of significant modes<2

Right skewed

r

the number of significant modes from the mode test is greater than 2;

r

the skewness test indicates that the data is symmetrically distributed.

As discussed in Section 9.1, the data patterns suggest only the five types of probability distribu- tions in the collected data. Hence, the five distributions can be mapped to the five distributions suggested by the data patterns:

r

Left skewed distribution, which corresponds to the 4thdistribution in Table 9.1 and is denoted

as DUL for Distribution, Unimodal, Left skewed.

r

Right skewed distribution, which corresponds to the 5th distribution in Table 9.1 and is

denoted as DUR for Distribution, Unimodal, Right skewed.

r

Normal distribution, which corresponds to the 3rddistribution in Table 9.1 and is denoted

as DUS for Distribution, Unimodal, Symmetric (implying the normal distribution).

r

Multimodal distribution, which corresponds to the 1stdistribution in Table 9.1 and is denoted

as DMM for Distribution, MultiModal.

r

Uniform distribution, which corresponds to the 2nddistribution in Table 9.1 and is denoted

as DUF for Distribution, UniForm.

Based on Table 9.1 and the five distributions suggested by the data patterns, the following test procedure is used to identify the probability distribution of a given data variable in the collected data:

1. Perform the DIP test. If the DIP test rejects the unimodality, the data variable is considered to have a multimodal distribution or DMM.

2. Perform the mode test and the skewness test, and determine the probability distribution based on the test results as follows:

(a) If the mode test indicates more than two significant modes and the skewness test indicates a symmetric distribution, the data variable is considered to have a uniform distribution or DUF.

(b) If the mode test indicates fewer than two significant modes and the skewness test indicates a symmetric distribution, the data variable is considered to have a unimodal, symmetric distribution or DUS.

(c) If the mode test indicates fewer than two significant modes and the skewness test indicates a left skewed distribution, the data variable is considered to have a unimodal, left skewed distribution or DUL.

(d) If the mode test indicates fewer than two significant modes and the skewness test indicates a right skewed distribution, the data variable is considered to have a unimodal, right skewed distribution or DUR.

9.3 PROCEDURE FOR DISCOVERING PROBABILITY

Documento similar