II. Problema de Investigación
2.1. Aproximación temática: observaciones, estudios relacionados, preguntas
The aims of this research study were to identify relative age effects in measures of student achievement and teacher ratings of student performance and behavior, to track the persistence of those age effects from kindergarten through third grade, and to explore the relationship between the measures of student achievement and teacher ratings of student
performance and behavior. To investigate these relationships, I used an instrumental variable to determine the causal effect of relative age and integrated that mechanism into an autoregressive cross-lagged framework to explore the continued effects of relative age and the relationships among the outcome variables. On the basis of the results from those analyses, I made
conclusions about the research questions and related hypotheses.
Research Question 1: How does relative age predict student performance on academic
assessments and teacher assessments of student behaviors in the primary grades after controlling for student SES and gender?
I hypothesized that relative age would have an effect on students’ achievement
assessment scores and teachers’ assessments of students’ behaviors after controlling for SES and gender. The results supported this hypothesis, as relative age had a positive and statistically significant effect on all three of the outcome variables at the end of kindergarten, indicating that at the beginning of kindergarten, being relatively older led to an increase in achievement test scores, ratings of academic performance, and ratings of learning behaviors after controlling for SES and gender. Relative age had a larger effect on students’ achievement assessment scores than on both of the teacher ratings outcomes, and the effect on the ratings of learning behavior was higher than the effect on academic performance ratings.
148
There has been little prior research into how teachers’ judgments of student performance are affected by a student’s relative age. The research on teacher judgments has not focused on student age, only on student grade level as a predictor of accuracy in teachers’ judgments (Paleczek et al., 2017; Valdez, 2013). The results of that research were inconsistent, as findings supported both increased accuracy in elementary and secondary teachers when compared to each other (Karing, 2009; Paleczek et al., 2017). I determined that at kindergarten entry, after
controlling for SES and gender, relative age had a positive, statistically significant effect on teachers’ ratings of student academic performance and learning behavior. Although not specifically addressed in the literature, this result is relatively intuitive, thus not surprising.
The difference in scores on achievement measures at the beginning of kindergarten attributable to the relative age differential has been well-documented (Bedard & Dhuey, 2006; Datar, 2006; Elder & Lubotsky, 2009; Puhani & Weber, 2007). This study adds support to the findings that students who are relatively older when starting kindergarten have an academic advantage, at least initially. Similar results have been used to bolster arguments for academic redshirting and giving students “the gift of time” by delaying entry into kindergarten (Bassok & Reardon, 2013). Historically, researchers have found that students more likely to be redshirted are male, White, and high SES (Bassok & Reardon, 2012, 2013). From the results of this study, I determined that males did have a higher average relative age than females. I also concluded that those identifying as White or Asian, non-Hispanic were more likely to have an increased actual relative age than were students identifying as a racial minority. However, I found no statistically significant effect of SES on initial relative age. Language learner status has not been used previously as a covariate of actual relative age, but I concluded that EL students were less likely to be relatively old than were non-EL students. This exploration of demographic determinants of
149
actual relative age adds clarification to the inconclusive research on whether SES, race, and EL status differentially explain the decision to enter kindergarten on time or delay entry.
Research Question 2: How does the effect of relative age on measures of academic performance and teacher ratings of student behaviors attenuate as students age?
Although the prior research on this aspect of the relative age effect is largely inconsistent and variable, I hypothesized that age-related gaps in achievement scores would decrease through the primary grades. The results of the current study are partially supportive of this hypothesis, as I found positive and statistically significant direct effects of relative age on student achievement, ratings of students’ academic performance, and ratings of students’ learning behaviors in the fall of kindergarten, and with the exception of the effect on ratings of academic performance, the direct effect of relative age on the outcomes dissipated by the Spring of kindergarten. The direct effect of relative age reversed by the end of first grade, indicating that the relatively younger students received more favorable ratings and achievement scores. However, by the end of second grade, the direct effect of relative age was positive for both teacher rating outcomes, but
remained negative for the achievement scores. By the end of third grade, only the effect on ratings of academic performance was statistically significant. Based on the direct effects of relative age, I concluded that for academic assessments and ratings of students’ learning
behaviors, the effect of relative age attenuated by the end of third grade, but persisted, although not as strongly, through third grade for ratings of students’ academic performance.
While direct effects are quite informative, they are calculated while everything else in the model is held constant. The total effects of relative age on the outcomes at each point provide a picture of the overall effect across time. From the calculated total effects, I determined that with respect to achievement scores, relative age’s effect had decreased but remained statistically
150
significant at the end of third grade. The total effects of relative age on ratings of academic performance did not dissipate across the grades with increases and decreases between the waves. While the direct effect of relative age at each wave had attenuated by third grade for achievement scores and ratings of learning behaviors, the total effect of entering kindergarten as relatively older decreased only for the achievement scores and actually increased for the learning behavior ratings. Although seemingly a distinction without a difference, the direct effects of relative age and the total effects of relative age are separate concepts, with the former often referred to as “age at test effects” (Black et al., 2011; Datar, 2006; Elder & Lubotsky, 2009; McEwan &
Shapiro, 2008). The age at test effect is a function of how old the student is at the time of the test, meaning if an assessment is taken by all children at the same time rather than at the same age then there will always be younger and older children (McEwan & Shapiro, 2008). The age at test effects are what I have referred to as direct effects of relative age throughout Chapters 3 and 4. The total effects of relative entry age discussed in this study are more specifically the total effect of having entered schools at an older or younger age and are related to a student’s absolute age. It is in this section of literature on relative age effect persistence that the current study has the most to add. To avoid introducing a confusion due to terminology, I will refer to the age at test effects as AATE and the total effects of relative entry age as TERA.
In the past, researchers have had difficulty parsing out the AATE from the TERA as they intertwine statistically and overlap conceptually but are often interpreted in combination simply as relative age effects. Economists have a particular interest in the delineation of the two
concepts, as they have different implications for lifetime earnings, child care expenses, and human capital (McEwan & Shapiro, 2008). The AATE theoretically decline over time because they are susceptible to the influence of time spent in school (Bedard & Dhuey, 2006; Elder &
151
Lubotsky, 2006). In other words, the more time students spend in school (i.e., they progress through the grades), the less important their relative age standing becomes. The AATE are represented in this study as the measurements of the direct effects of relative age at each wave. In applying this conceptual framework to my results, I concluded that the AATE portion of the relative age effect on achievement scores and teacher ratings of students’ learning behaviors attenuates as students progress through elementary school. The AATE on teacher ratings of student academic performance lessens by the end of third grade, but remains statistically significant.
The TERA have increased longevity compared to AATE, and their persistence is reminiscent of the Matthew Effect, in which students who enter school older receive
quantitatively and qualitatively improved opportunities beginning in kindergarten because of their increased age and continue to receive such benefits throughout their education because of prior standing. The TERA are represented in this study by the total effect of the Wave 1 relative age on the outcomes at each wave. Analyzing the results through this lens I determined that the effect of entering school as a relatively older or younger student remains significant, but
decreases by the end of third grade on measures of student achievement. For measures of teachers’ perception of student performance, the TERA have a differential effect depending on the grade level, but show no overall decrease at the end of third grade.
Because I was able to assign students an ARA at each wave, I was able to disentangle AATE from TERA. There are only a handful of published research studies delineating the two effects, as researchers have thought the two effects to be inseparable, arguing that a student’s age at test and relative kindergarten entry age were “perfectly collinear” (Black et al., 2011).
152
move between schools with different cutoff dates and can be retained or accelerated in grade level, both of which lead to changes in actual relative age. This practical separation is evident in the path coefficients of the autoregressive actual relative age pathway (Chapter 4, Figure 4.2). While the coefficients are indicative of strong predictability between the ages, they (with the exception of relative age at Wave 2) are not collinear. As such, I was able to differentiate the effect of being relatively young at an assessment point from the effect of entering kindergarten as a relatively older or younger student. Prior researchers attempting to make a distinction between the two effects did not account for changes in relative age due to moving or grade level retention and acceleration, which complicated their ability to separate the effects statistically.
To provide a definitive answer to this research question, the effects of a student’s age at test (AATE) begin to decline as soon as first grade, but the total effects of a student’s relative entry age (TERA) persist through third grade, only showing a decline for student achievement scores. The decline of the AATE is supportive of prior research that found initial effects of being relatively older on achievement scores and academic measures begin to attenuate starting in the primary grades (Elder & Lubotsky, 2009). The decrease in TERA on achievement measures and lack of change on teacher ratings of academic performance and learning behaviors are supportive of the inconsistent results of prior research on relative age, as it appears the effects are dependent on the grade level when examined and type of outcome measure, although research on teacher perception measures is scant (Bassok & Reardon, 2013; Crawford et al., 2017; McEwan & Shapiro, 2008; Puhani & Weber, 2007).
To provide students with appropriate educational opportunities and services, schools often use assessment data early in a student’s educational career to evaluate student performance strengths and weaknesses. Educators and other decision makers use assessments of student
153
ability and achievement combined with observation of student behaviors to place a student in specific programs, advocate for educational services, and to assess student progress toward learning goals (Cao, Jung, & Lee, 2017; Gredler, 2000). Special education and gifted education are need-based interventions in which students must display and be identified with a need to receive the services and supports. The current study’s results have important implications for need-based programs in education in which teachers’ observations of student performance play a large role. If teacher judgments and referrals are a main component of the identification for services process and teacher judgments are affected by a student’s relative age, relatively older students would be identified by teachers for gifted or high potential services while the relatively younger students would be referred for special education services more often. A student’s relative age must be taken into account when teacher judgments are used as part of the referral process for need-based services.
The kindergarten entry age effect is an additional phenomenon not explicitly a focus of this study, but worth mentioning nonetheless. It is a function of schools having an earlier or later cutoff date for school entry and theoretically would affect everyone entering school under a particular cutoff date guideline in the same way. It is represented by the effect of the entry age variable on the outcomes. Based on the results of the current study, this cluster-level
kindergarten entry age effect on student outcomes starts out positive in the fall of kindergarten (i.e., schools with cutoff dates early in the year have students that are relatively older), becomes negative in first and second grades (i.e., schools with cutoff dates later in the year have students that are relatively older), and then positive at the end of third grade. However, these results are not statistically significant, and as such should not be used as a basis for policy implications.
154
Research Question 3: What is the magnitude of the relationship between teacher assessments of student behaviors and student performance on academic assessments in the primary grades?
Large portions of the current research on teachers’ judgments focuses on the comparison of the judgments to measures of academic performance rather than an analysis of the stability or predictability of those judgments, which was the focus of this research question. From the results of the current study I concluded that measures of student achievement are highly predictive of student achievement at the next grade, measures of teacher perceptions of students’ learning behaviors are moderately predictive of teacher perceptions of students’ learning behaviors at the next grade, and measures of teacher perceptions of students’ academic performance are not very predictive of teacher perceptions of students’ academic performance at the next grade. Student achievement scores are predictive of academic performance ratings at the next grade level, but the relationship does not work in the reverse direction (i.e., academic performance ratings are not predictive of academic achievement at the next grade level). Ratings of academic performance negatively predict ratings of student learning behaviors at the next grade level, and ratings of student learning behaviors positively predict ratings of academic performance at the next grade level. These results show that the relationships between teacher perceptions of student behaviors and performance and student achievement scores are complicated, and the conclusion here is not that one form of assessment is better than the other, but more that the assessments are capturing different aspects of behavior or performance. If the goal of a program was to identify students for some type of academic services based on achievement, it would be more advantageous to use student achievement scores than teacher judgments of student achievement.
155
Previous research on teacher ratings has had two areas of focus. The first is exploring whether teachers rate students differently based on student, teacher, and assessment
characteristics. That research has produced evidence that teachers’ perceptions of student skill, behavior, and performance vary by student race, gender, and SES, generally finding that teachers assigned lower ratings to students who were racial minorities and from low SES backgrounds than to White and Asian students and those from high SES backgrounds (Ready & Wright, 2011; Valdez, 2013). The results of the current study show that the relationships between the teacher ratings and SES were strongest in the fall of kindergarten where high SES students received higher ratings. However, in the spring of kindergarten, first, and second grade lower SES students receive higher teacher ratings of academic performance. The relationship of student gender to teachers’ ratings differed by content area (Meissel et al., 2017; Pigott & Cowen, 2000; Ready & Wright, 2011; Valdez, 2013). Additionally, researchers found significant differences in how teachers rated student performance with students with special needs and ELs receiving lower scores (Meissel et al., 2017). In the results of this study, teachers rated EL students higher than non-EL students in academic performance and learning behaviors in the spring of
kindergarten, first, and second grades. Although statistically significant, the strength of the relationships between the covariates and teacher judgment variables is low with the exception of the SES relationship in the fall of kindergarten.
The second focus of teacher rating/judgment research is determining whether teachers are accurate in their judgments of students’ performance. The research studies with this focus
usually compare teacher ratings of student performance to student performance on achievement tests to determine whether there is systematic over- or under-estimation of student ability as defined by a student’s performance on an achievement test. The overwhelming conclusion for
156
these studies is that there is variability in accuracy by student demographic characteristics (Feinberg & Shapiro, 2003, 2009; McKown & Weinstein, 2008; Rubie-Davies et al., 2012; Südkamp et al., 2012). Researchers also found that the type of instrument used mediates
teachers’ accuracy. Teachers showed increased accuracy when using assessments that provided a standard for comparison and have skill-specific questions (Hoge & Coladarci, 1989; Südkamp et al., 2012). The achievement assessments used in this study were content and standards-based with multiple-choice options for the student to indicate a definitively correct or incorrect answer. The teacher ratings of academic performance consisted of questions asking the teacher to rate the student’s general ability in a content area. The indicators for the teacher ratings of learning behaviors variables were composite scores of questions on the frequency with which a student displayed a specific behavior. The results of this study showed high consistency and stability between years for the two measures that had either questions with correct answers (student achievement assessments) or questions relating to very specific behaviors (teacher ratings of student learning behaviors). Combined with evidence from prior research I conclude that a reason for the decreased consistency in the teacher ratings is the lack of specificity in the
questions about students’ academic performance, which is more of a conclusion about the type of assessment than the teacher’s ability to assess a student’s performance accurately. In this
research study, I analyzed the relationship between achievement scores and ratings of academic performance and learning behaviors not to determine whether teachers were accurate in their judgments, but rather to explore the relatedness of the measures being used as standard for accuracy.
The results showed that the three measures were moderately to highly correlated with one