6 RESULTADOS
6.1.2 Categorización final
As opposed to the motivating example of Firth (2003) where inference should be carried out for a contrast that would require the complete variance-covariance matrix, here, quasi-variances are employed to overcome the problem that onlyk−1 standard errors for the DIF analysis of a test are at hand, but allkitems with the corresponding parametersβjare of interest.
In this thesis, the Wald test is used for item-wise DIF detection. In the previous sections, the first anchor item was declared DIF free. The Wald test statistics for all remaining items jof the test included the differences of the estimated and anchored item parameters ˆβrefj −βˆfocj (which strongly
9.2 Quasi-variances 157 depend on the anchor items) in the numerator and the standard errors
q c Var( ˆβref j −βˆ foc j ) = q c Var( ˆβref)
j,j+Var( ˆc βfoc)j,jin the denominator. Therefore, the item parameters are first estimated
separately in each group with the constraint ˜β1 = 0 and then anchored by transforming each group to the restriction reflecting the intended anchor method (see Section 9.2.1 or Chapter 5). The covariances for contrasts between the item parameters of the reference and the focal group are zero.
To allow for testing allkitems, the resulting variance-covariance matrix for each group is now used to calculate quasi-variances for allkitems. These quasi-variances are used to construct the
quasi-Wald testfor each item j, j=1, . . . ,kbased on the quasi-variances s2
j Tj = ˆ βref j −βˆ foc j q c Var( ˆβref) j,j+Var( ˆc βfoc)j,j ≈ ˆ βref j −βˆ foc j q s2j,ref+ s2j,foc . (26)
Note, however, that the anchor problem also occurs in the quasi-Wald test since the differ- ences in the numerator still depend on the anchor method. To evaluate whether the estimated quasi-variances themselves are invariant to the anchor method chosen, the quasi-standard errors were calculated in a short simulation study (results not shown). The data consist of a test of length 40 with 45% balanced or unbalanced DIF-items. Quasi-standard errors were calculated for item parameters that were anchored using either the constant-single anchor (either pure or contaminated) or the constant-four anchor (either pure or contaminated with the degree of con- tamination equal to 50% or 100%). The quasi-standard errors of the quasi-Wald tests for two exemplary items (one simulated DIF-item and one simulated DIF-free item) did not differ in any replication, when machine precision is taken into account. Thus, in the following, it is assumed that quasi-variances are invariant against the employed anchor method. Future research may prove this result from a theoretical perspective.
9.2.3 Simulation study
In order to evaluate the performance of the quasi-Wald test compared to the Wald test, a short simulation study is carried out in the free R system for statistical computing (R Development Core Team, 2011) using the add-on R-package qvcalc by Firth (2012). The following ques- tions are addressed: How do the quasi-Wald tests perform in comparison with the Wald tests for thek−1 items excluding the first anchor candidate? Is the decision of declaring the first anchor as DIF-free appropriate or is the quasi-Wald test result superior?
Data generating process
Each data set, that represents one of 2000 replications from one simulation setting, corresponds to the simulated responses of two groups of subjects (the reference (ref) and the focal (foc) group) in a test withk= 40 items.
• Person and item parameters
The person parameters are generated from a normal ability distribution with a higher mean for the reference group θref ∼ N(0,1) than for the focal group θfoc ∼ N(−1,1) similar to Wang et al. (2012). Values assigned to the item parameters are, again, β = (-2.522, -1.902,. . ., 1.592) used by Wanget al.(2012), see Section 7.3. The responses in each group follow the Rasch model and are generated similar to the previous simulation studies (see e.g. Section 6.4).
• DIF-items
The first 15, 30 or 45 percent of the items (cf. paragraph DIF proportions in the next section) are chosen to display uniform DIF by setting the difference in the item parameters of reference and focal group∆DIF = βrefj −β
foc
j to+.6 or−.6 consistent with the intended
direction of DIF.
Manipulated variables
Similar to previous simulation studies for the comparison of two groups (cf. Chapter 6 and Chapter 7), the manipulated variables were the sample size, the direction of DIF, the percentage of DIF and the anchor methods.
• Sample sizes
The sample sizes in reference and focal group are defined by the following pairs (nref,nfoc) ∈ {(250,250), (500,250), (500,500), (750,500), (750,750), . . ., (1500,1500)}
where, again, equal and different group sizes are considered.
• Directions and proportions of DIF
The direction of DIF is either balancedwhere each DIF-item favors either the reference or the focal group but no systematic advantage for one group remains because the effects cancel out or unbalanced where a systematic disadvantage for the focal group is gen- erated. The proportion of DIF is set to p ∈ {15%,30%,45%}. In order to save space, only the most difficult situation of 45% DIF-items is discussed in detail, but the other proportions are included in the Appendix of this chapter.
• Anchor methods
The quasi-Wald test (indicated by the ending -quasi) and the Wald test are combined with the previously suggested constant4-MPT and forward-MTT anchor methods (see Chapter 7).
Outcome variables
To assess the performance of the quasi-Wald tests, two situations are regarded. Firstly, the quasi-Wald tests are compared to the Wald tests for all items except for the first anchor item.
9.2 Quasi-variances 159 Secondly, the declaration that the first anchor item is DIF-free (that was used throughout this thesis) is compared with the result of the quasi-Wald test for the first anchor item.
• Classification rate
For a single replication theclassification rate indicates the number of correct test deci- sions about whether the item is a DIF-item. The estimated classification rate for each experimental setting is computed as the mean over all 2000 replications. This outcome variable is evaluated for the first anchor candidate.
• False alarm rate and hit rate
For the remainingk−1 items, the false alarm rate (as a measure for the type I error) and the hit rate (that corresponds to the powerof the statistical test) were calculated similar to Section 6.4.
• Further outcome variables
In addition, the proportion where the first anchor item was a DIF-item and also character- istics of the item-wise tests for the first item (a simulated DIF-item) and the last item (a simulated DIF-free item) of the test, such as the estimated (quasi-) standard errors, were computed.