Integrating Ensemble Clustering and Text Embeddings for Estimating the Factor Loadings of Self-Report Scales

0Citations
Citations of this article
2Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Advances in large language models can provide opportunities to evaluate the characteristics of scales prior to data collection. In this study, we explore if item text can be used to predict a scale’s psychometric properties. Specifically, we examine if clustering consensus (i.e., the frequency by which items are grouped with other items from the same underlying factor across multiple clustering algorithms), and a cosine similarity metric (i.e., the semantic similarity of items to other items from the same factor), can be used to predict exploratory factor analysis (EFA) factor loadings. Across six scales with varying sample sizes, number of factors/items, we found that both the cosine similarity and ensemble clustering consensus methods predicted factor loading values. While the methods share some conceptual and empirical overlap, and results vary by scale, the ensemble clustering approach explains incremental variance above and beyond cosine similarity in predicting factor loadings. Using both methods in conjunction can be a useful way to identify problematic items prior to data collection and help researchers develop more optimal scales from the onset, thereby potentially saving time, resources, and increasing the likelihood of developing sound measures.

Cite

CITATION STYLE

APA

Voss, N. M., Wu, F. Y., Javalagi, A. A., & Kell, H. J. (2026). Integrating Ensemble Clustering and Text Embeddings for Estimating the Factor Loadings of Self-Report Scales. Educational and Psychological Measurement. https://doi.org/10.1177/00131644261430762

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free