Developing and Evaluating an English Language Proficiency Test Using Generalizability Theory: A One-Facet Design

Authors

  • Hossein Salarian University of Tehran, Tehran, Iran

DOI:

https://doi.org/10.61227/arji.v8i3.843

Keywords:

English language proficiency assessment, Generalizability Theory, G-study, D-study, score dependability, person-by-item design, test development

Abstract

The present study aimed to develop and evaluate an English language proficiency test using Generalizability Theory (G-Theory) and a one-facet person-by-item (p × i) design. The study addressed the need for a systematic examination of the sources of variability underlying English proficiency scores and the identification of an efficient test length. Sixty Iranian English language learners participated in the study and completed a researcher-developed 30-item multiple-choice English language proficiency test assessing vocabulary, grammar, and reading comprehension. The test was administered under standardized conditions, and participants’ responses were analyzed using an analysis of variance within a G-Theory framework. In the Generalizability Study (G-study), persons were treated as the objects of measurement and items as the single measurement facet. The analysis estimated the variance components associated with persons, items, and the person-by-item/residual component. The findings indicated that the person component accounted for a substantial proportion of the observed score variability (39.95%), whereas item-related variance was very small (0.13%). The person-by-item/residual component represented the largest proportion of variance (59.92%), suggesting that learners’ performance varied across items and that measurement error remained an important source of score variability. The 30-item test produced a high Generalizability coefficient (G = .952), indicating strong score dependability for relative decisions concerning learners’ English proficiency. A Decision Study (D-study) further examined alternative test lengths of 10, 15, 20, 25, and 30 items. The results showed that the G-coefficient increased systematically from .870 for 10 items to .952 for 30 items, while relative error variance decreased as the number of items increased. Although the 30-item test yielded the highest dependability, the 20- and 25-item versions also demonstrated high levels of score dependability, suggesting a practical balance between measurement precision and testing efficiency. The study concludes that G-Theory provides a useful framework for identifying sources of score variability and making evidence-based decisions about English proficiency test design and length.

Downloads

Download data is not yet available.

References

Akindahunsi, O., & Afolabi, E. R. I. (2021). Using Generalizability Theory to investigate the reliability of scores assigned to students in English language examination in Nigeria. Journal of Measurement and Evaluation in Education and Psychology, 12(2), 147–162. https://doi.org/10.21031/epod.820989

Andersen, S. A. W., Nayahangan, L. J., Park, Y. S., & Konge, L. (2021). Use of Generalizability Theory for exploring reliability of and sources of variance in assessment of technical skills: A systematic review and meta-analysis. Academic Medicine, 96(11), 1609–1619. https://doi.org/10.1097/ACM.0000000000004150

Aryadoust, V., Zakaria, A., Lim, M. H., & Chen, C. (2020). An extensive knowledge mapping review of measurement and validity in language assessment and SLA research. Frontiers in Psychology, 11, Article 1941. https://doi.org/10.3389/fpsyg.2020.01941

Chen, D., Hebert, M., & Wilson, J. (2022). Examining human and automated ratings of elementary students’ writing quality: A multivariate Generalizability Theory application. American Educational Research Journal, 59(6), 1122–1156.

Ellis, J. L. (2021). A test can have multiple reliabilities. Psychometrika, 86(4), 869–876. https://doi.org/10.1007/s11336-021-09800-2

Güvendir, M. A. (2022). Generalizability Theory in testing L2 speaking. In J. I. Liontas (Ed.), The

TESOL encyclopedia of English language teaching. Wiley. https://doi.org/10.1002/9781118784235.eelt1034

Jackson, D. J. R., Michaelides, G., Dewberry, C., & Englert, P. (2022). Clarifying the scope of Generalizability Theory for multifaceted assessment. New Zealand Journal of Psychology, 51(2), 53–64.

Kocaoğlu, S., & Şahin, M. G. (2024). Investigating the effect of testlets consisting of open-ended and multiple-choice items on reliability via Generalizability Theory. Journal of Measurement and Evaluation in Education and Psychology, 15(1), 65–78. https://doi.org/10.21031/epod.1429423

Li, G. (2023). Which method is optimal for estimating variance components and their variability in Generalizability Theory? Evidence from a set of unified rules for bootstrap method. PLOS ONE, 18(7), Article e0288069. https://doi.org/10.1371/journal.pone.0288069

Liao, R. J. T. (2023). The use of Generalizability Theory in investigating the score dependability

of classroom-based L2 reading assessment. Language Testing, 40(1), 86–106. https://doi.org/10.1177/02655322211070840

Norris, J. M., & Lee, J. (2023). The effectiveness of the TOEFL Essentials test for distinguishing English proficiency levels (ETS Research Memorandum No. RM-23-07). Educational Testing Service.

O’Sullivan, B. (2023). Reflections on the application and validation of technology in language testing. Language Assessment Quarterly, 20(4–5), 501–511. https://doi.org/10.1080/15434303.2023.2291486

Polat, M., & Turhan, N. S. (2021). Applying Generalizability Theory in language testing: Comparing nested and crossed scoring designs in the assessment of speaking skills. International Journal of Curriculum and Instruction, 13(3), 3344–3358.

Salarian, H. Ahmadi Shirazi, M., & Alavi, S. M. (2019). An investigation into item types and text types of reading comprehension section of Iranian Ph.D. entrance exams using G-theory. Journal of Modern Research in English Language Studies, 6(1), 1-29.

Sari, E., & Han, T. (2022). Using Generalizability Theory to investigate the variability and reliability of EFL composition scores by human raters and e-rater. Porta Linguarum, 38, 161– 177. https://doi.org/10.30827/portalin.vi38.18056

Schmidgall, J., Cid, J., Carter Grissom, E., & Li, L. (2021). Making the case for the quality and use of a new language proficiency assessment: Validity argument for the redesigned TOEIC Bridge tests. ETS Research Report Series, 2021, 1–22. https://doi.org/10.1002/ets2.12335

Shin, J. (2022). Investigating and optimizing score dependability of a local ITA speaking test across language groups: A Generalizability Theory approach. Language Testing, 39(2), 313– 337. https://doi.org/10.1177/02655322211052680

Wolf, M. K., Bailey, A. L., & Ballard, L. (2023). Aligning English language proficiency assessments to standards: Conceptual and technical issues. TESOL Quarterly, 57(2), 670–685. https://doi.org/10.1002/tesq.3199

Additional Files

Published

2026-09-16

 


How to Cite

Salarian, H. . (2026). Developing and Evaluating an English Language Proficiency Test Using Generalizability Theory: A One-Facet Design. Action Research Journal Indonesia (ARJI), 8(3), 1098–1115. https://doi.org/10.61227/arji.v8i3.843

Similar Articles

<< < 9 10 11 12 13 14 15 16 17 > >> 

You may also start an advanced similarity search for this article.