4. ANÁLISIS SOBRE LAS MEJORES TÉCNICAS DISPONIBLES DEL SECTOR
4.7. Gestión medioambiental
4.7.1. MTD en la gestión medioambiental
One way in which generalizability can be determined is through consistency between
multiple methods that score the same traits for the same candidates. Two methods were compared: The Teacher Candidate Final Progress Report33 (PR) and TPA Scores. PRs were completed by
supervisors and mentors as a summative evaluation and recommendation of candidate readiness. For this method, raters scored candidates on twenty-seven separate traits aligned to Standard V. Twenty-one of the twenty-seven PR traits aligned to TPA rubrics. Inter-rater agreement was calculated to determine the degree to which multiple evaluators gave the same score to identical evidence (Graham, Milanowski, & Miller, 2012, p. 5). Because eight candidates had two mentors, differences in rater agreement were considered separately. These scores were compared to the set of scores provided by Pearson. Following a similar process described above (see VE 5), the common indexes for measuring inter-rater agreement were calculated including the percentage of absolute and adjacent agreement and Cohen’s Kappa. In addition to providing the agreement correlations, a three-way Multi-Trait, Multi-Method (MTMM) was developed to compare scores. Finally, a summary of candidate score differences was calculated, focusing on where high-stakes results differed.
33 Sterner revised the PR to align with PESB Standard V criteria in 2008. The PR was aligned to the TPA in 2012, but not significantly modified from its 2008 version.
Findings.
Thirty teacher candidate samples were scored by both methods. In total, 1512 traits were evaluated on the PR using a 5-point Likert scale. Of these, 247 differed (.16). Nearly all (.88) of this difference was within one rubric level (+1/-1). Some (.21) difference is explained by supervisors scoring +1. Most (.67) difference occurred when mentors scored the candidates -1. The total difference in rubric scores within two rubric levels (+2/-2) was 10% and within three rubric levels (+3/-3) was 1%.34Scorer differences were greatest for Standard 3: Knowledge of Teaching and Instructional skills.35 No standards had more than >28% agreement difference (see Appendix S for PR indexes and
Appendix T for agreement percentages). Based on this data, the PR is a consistent scoring tool with
higher levels of scorer agreement than TPA scores (see VE 5).
However, score scatterplots indicate that the PR results demonstrate a ceiling effect with a bunching of scores at the upper level. Studies have indicated that evaluations of teachers can be subjective36 and the ceiling effect on the PR may be a result of such subjectivity. Results with a ceiling effect may also be the result of inherent flaws in the instrument design, for instance the Likert scale may not sufficiently distinguish between upper levels of the scale. It is important to note
34 Much of these differences are accounted for by one outlier in the data. When the outlier data is removed, there are no +3/-3 differences, and +2/-2 differences fall to 2%.
35 Specifically, Standards 3.5, 3.7, 3.8 and 3.9. In addition, Standard 2.1 “Accurately assesses student needs (emotional/academic) and provides feedback” and Standard 2.6 “Communicates clearly and
professionally with families” ranked higher, although 2.6 was ultimately not used because it did not align to the TPA requirements.
36 In fact, researchers have identified a “good subjective” and “bad subjective” in the evaluation of practicing teachers (Rockoff & Speroni, 2011). Others have found that principal evaluation distinguishes between poor teacher quality and excellent teacher quality, but rarely distinguish between teachers who fall in the middle of the distribution (Jacob & Lefgren, 2005).
that only candidates that have passed university benchmarks are allowed to student teach. For these reasons, higher scores of teacher readiness are expected because candidates not ready to teach should have been removed from the sample prior to student teaching. While still useful as a threshold test (for licensure decisions, for instance), PR scores do not allow for a ranking of top performers. Like the TPA results, the purpose of the PR is to provide the university with data to support a recommendation for licensure (a yes or no question). Therefore, its purpose as a threshold test is supported. However, the ceiling effect on the PR limits the variability gathered from any one trait and may reduce the power of statistics on correlations between any two variables.
MTMM.
An item-correlation matrix provides the correlations between aligned traits in the PR and TPA. An MTMM was produced using these correlations and Cronbach’s α. MTMM analysis indicates that reliability for each method is strong (.73-.85) with good internal consistency for each measure. In addition, heterotrait-heteromethod correlations are low demonstrating that the traits do not correlate highly with each other between methods. However, the heterotrait-monomethod correlations are high, suggesting that the traits evaluated by both methods are highly correlated. This could influence Cronbach’s α and method reliability. In other words, high reliability indicators could be the result of item similarity rather than internal consistency. The purpose of an MTMM is to demonstrate the degree to which two different methods, intended to measure the same traits, actually do so. Significantly, for the purposes of this study, the MTMM indicates that these measures do not correlate with each other. Results clearly indicate a method-effect. Surprisingly, there appears to be no discernable pattern between summative PR scores of teacher readiness and summative scores of teacher readiness on the TPA.Table 4.6
Pass Rate Comparisons.
Finally, using TPA cut scores establishing pass rates (35 points), TPA pass rates were compared to PR pass rates. Results indicate that four students who passed PR would have failed the TPA in this field test. Three of those candidates were recommended to pass ST. Table 4.7 lists these four candidates’ TPA subject, TPA score, PR score, and recommendation to pass ST.Table 4.7
TPA Failures Compared to PR/Student Teaching Recommendations
TPA Subject Area TPA Score Score on PR Pass Student Teaching?
1 Elementary literacy 29 89% Yes
2 Elementary literacy 26 39% No
3 Secondary mathematics 31 69% Yes
4 Elementary literacy 34 97% Yes
The three highest TPA scores were compared to PR scores, TPA subject area, and recommendations to pass ST in Table 4.8. A perfect score on the TPA is 80 points. The results indicate that one
candidate passed the TPA with high scores but did not receive a recommendation to pass ST. Table 4.8
TPA Top Scores Compared to PR/Student Teaching Recommendations
TPA Subject Area Score on TPA Score on PR Pass Student Teaching?
1 Secondary social studies 72 79% Yes
2 Secondary science 61 61% No
Finally, Table 4.9 identifies TPA and PR scores for candidates who struggled to meet expectations during ST or required intervention from the supervisor or university faculty during the term.37 Of these nine candidates, two earned the top three highest scores on the TPA and one failed the TPA. Table 4.93839
TPA Scores Compared to Student Teaching Support Intervention Needed
Score on TPA Score on PR Pass Student Teaching?
1 49 66% Yes 2 72 78% Yes 3 41 76% Yes 4 57 93% Yes 5 56 98% Yes 6 40 98% Yes 7 61 61% No 8 47 73% Yes 9 26 39% No
Supports for Validity.
Errors of Measurement (EoM) are not necessarily errors in the assessment instrument but they influence the data and how it might be interpreted. The data from37 Intervention was determined through a midterm progress report, or by interviews with supervisors or faculty. Intervention can occur in many forms, from a single discussion reminding mentors and candidates of university expectations, to a formal contract of required behaviors. Not all interventions are due to
candidate performance. Occasionally, a traumatic life-event, conflict with course requirements, or differences in expectations of mentor support can prompt an intervention. The study investigator facilitated interventions for undergraduate candidates during data collection.
38 Note that candidates 1 and 2 in Table 4.8 also appear in Table 4.9 as candidates 2 and 7. 39 Note that candidate 2 on Table 4.7 is candidate 9 in Table 4.9.
the MTMM may reflect random errors based on statistical fluctuations that exist when measured values are inconsistent and these cannot be predicted. One way EoM are minimized is through standardization (see Chapter two). However, greater standardization in performance assessment are problematic because it requires increased interpretation by candidates and scorers.
Threats to Validity.
Where overall candidate performance is concerned, PR and TPA scores agreed in only one in four cases (.25) for failures, two in three cases (.67) for mastery, and two in nine (.22) marginal cases. In addition, a significant, systematic method-effect was found when analyzing the data using MTMM. Method-effect indicates that score differences are the result of the method of measurement, rather than participant performance. One would expect to see higher levels of correlation between two summative instruments of teacher readiness given to the same population in the same term. Because the analysis of PR data suggests a ceiling effect, and this was a field test, or pilot, for the TPA, it is difficult to be certain whether the method effect belongs to one or both measures.Generally, score levels on the TPA represent candidate performance. However, data indicates that TPA scores could have led to one candidate (.02) receiving licensure who had not yet established readiness and three candidates (.05) to be denied licensure when their ST performance indicated a readiness to teach. In addition, for the nine candidates who required intervention, TPA scores should have reflected candidate difficulty to meet expectations; however, only one of nine of these candidates’ TPA scores suggests potential readiness issues. It is not clear, therefore, from the MTMM and pass rate comparisons that score levels consistently represent candidate performance or whether poor performance on the TPA is a reliable indicator of lack of teaching readiness.