
The Adapted Measure of Math Engagement (AM-ME) uses a scale to measure student engagement in math. This scale was developed as part of a multi-year community participatory action research project. More details about the project can be found on the project webpage. This post reviews the quantitative methods used to evaluate the construct validity of the scale.
Data collection
After an iterative, multi-year research process including Rasch item analysis and item response theory, literature review, interviews, focus groups, and meetings with student and teacher researchers, we began the validation stage with a pool of 70 items to capture various facets of math engagement. The survey was administered to 2,227 middle and high school students as an internet-based survey in spring 2025. Given the length of the survey, a planned missingness design was used to reduce respondent burden. This design works by dividing the candidate survey items for the AM-ME into four blocks; six versions of the survey are then created so that each survey form contains two out of the four blocks. This is designed so that each block is presented once (and only once) with each other block which ensures all items will co-occur with all other times for at least some respondents.
After data was collected, we randomly split the sample into two halves. We used the first half to explore and identify the latent structure of the scale using exploratory factor analysis (EFA) and reserved the second half to confirm that structure using confirmatory factor analysis (CFA) on data the exploratory step had never seen. This approach is common in scale validation to avoid overfitting: Testing the CFA on data not used to generate the factor structure provides more robust evidence for the patterns being identified.
Exploratory factor analysis
EFA is a data-driven technique used to uncover the underlying structure of a set of observed variables. It assumes that the correlations among items can be explained by a smaller number of unobserved, or latent, variables called factors. Each factor represents a coherent dimension that items covary with, or “load onto,” and items that vary together are treated as indicators of the same underlying construct. The model was fit using maximum likelihood estimation. The model was fit in R using the lavaan package.
A central question in any EFA is how many factors emerge among the items. The goal of factor analyses is to explain variations in students’ responses. When items vary similarly, or co-vary, we assume that an underlying concept explains that variance. To identify how many factors best fit the data, we answered it using two complementary sources of evidence. The first was a scree plot—a graph that displays each factor’s eigenvalue (a measure of how much variance that factor accounts for) in descending order. Reading a scree plot involves locating the “elbow,” the point at which the curve flattens, and additional factors explain little further variance; factors above the elbow are retained. The second source was substantive theory: We asked if the factors that emerged were interpretable and consistent with what is known about math engagement. This was a critical part of the community-engaged aspect of this work: The AM-ME research group played a major role in defining the factors after initial data analysis. Combing findings from the scree plot with theory, we concluded that eight factors best described the data. Items with low factor loadings (<.3) or those that cross-loaded on multiple factors were removed. This -led to a final scale with 34 items.
We hypothesized that the eight factors are themselves elements of a single, broader construct, or higher order factor, which we labeled “math engagement.” A higher-order factor is a second-level latent variable that explains the correlations among the first-order factors, just as those first-order factors explain the correlations among the individual items. This structure expresses the idea that math engagement is a unifying disposition that gives rise to several distinguishable but related components (the factors).
Confirmatory factor analysis
To test this hypothesized structure, we turned to the second half of the sample and conducted a confirmatory factor analysis (CFA). Where EFA discovers structure, CFA evaluates a structure specified in advance. We told the model exactly which items belonged to which factors and that the eight factors loaded onto a single higher-order factor; then we assessed how well that model reproduced the relationships observed in the data. This was fit using a structural equation model framework in R using the lavaan package.
Because the survey used Likert response options (in this case, strongly disagree through strongly agree), we treated the responses as ordinal rather than continuous. Accordingly, we estimated the model with a robust mean- and variance-adjusted weighted least squares estimator (WLSMV / WLSMVS). This estimator is built for categorical, non-normally distributed indicators: instead of assuming multivariate normality, it works from the polychoric correlations among the ordinal items and adjusts the model’s test statistic so that fit is evaluated appropriately.
To handle incomplete responses, we used a pairwise approach to missing data. Pairwise estimation uses all the data available for each pair of variables when computing the associations the model is built on, thereby retaining more information from the sample. This method is well suited to surveys with planned missingness designs. The resulting model showed acceptable fit, supporting our hypothesized higher-order structure. We evaluated fit by comparing a variety of fit statistics including CFI, TLI, and RMSEA to commonly accepted criteria of good fit.
Measurement invariance
A scale that fits well in the full sample may still behave differently across subgroups of respondents. Before scores can be compared fairly across groups, we need evidence of measurement invariance—that is, evidence for whether the instrument measures the same construct in the same way for all groups of students. We evaluated invariance through a standard, increasingly strict sequence of nested models.
Configural invariance is the least restrictive level. It tests whether the same overall pattern holds across groups—the same items load on the same factors—while allowing the strength of the relationship between each item and the factor. Establishing configural invariance shows that respondents share the same conceptual map of the construct, and it provides the baseline against which the stricter models are compared.
Metric invariance (also called weak invariance) goes further by constraining the factor loadings to be equal across groups. Loadings describe how strongly each item is tied to its factor, so equal loadings indicate that the construct has the same meaning and is measured on the same scale across groups. When metric invariance holds, relationships between the factor and other variables can be compared meaningfully.
Scalar invariance (also called strong invariance) adds the requirement that the item intercepts—or, with ordinal data, the item thresholds—also be equal across groups on top of the equal loadings. Thresholds govern the response level at which a person tends to cross from one category to the next. When this constraint holds, observed differences in scores reflect genuine differences in the underlying construct rather than group-specific response tendencies, which is what licenses fair comparison of latent means across groups.
At each step we compared the more constrained model with the one before it, checking whether the added equality constraints meaningfully worsened fit. We found that model fit did not meaningfully decrease as these constraints were added, which suggest this scale works well for all groups. Holding up under this sequence provides strong evidence that the math-engagement scale functions equivalently across the groups examined.
Suggested citation: Kelley, C., Holquist, S., Crowder, M., Hsieh, D., Scott, A., Yu, M., & the Adapted Measure of Math Engagement Research Group. Quantitative methods used to develop and validate the Adapted Measure of Math Engagement (AM-ME). Child Trends. DOI: 10.56417/7698f6256u
References
Graham, J. W., Taylor, B. J., Olchowski, A. E., & Cumsille, P. E. (2006). Planned missing data designs in psychological research. Psychological Methods, 11(4), 323–343. https://doi.org/10.1037/1082-989X.11.4.323
Rhemtulla, M., & Little, T. D. (2012). Planned missing data designs for research in cognitive development. Journal of Cognition and Development, 13(4), 425–438. https://doi.org/10.1080/15248372.2012.717340
Vandenberg, R. J., & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organizational Research Methods, 3(1), 4–70. https://doi.org/10.1177/109442810031002
Worthington, R. L., & Whittaker, T. A. (2006). Scale development research: A content analysis and recommendations for best practices. The Counseling Psychologist, 34(6), 806–838. https://doi.org/10.1177/0011000006288127


