An Exploratory Evaluation of GPT-4's Consistency as an English Essay Rater: A Many-Facet Rasch Model Analysis of AI versus Human Rating Patterns
- 1 Brigham Young University-Hawaii, United States
- 2 Florida State University, United States
- 3 Brigham Young University, United States
Abstract
This study examined the defensibility of using GPT-4 for automated essay scoring, using a ManyFacet Rasch Model analysis. Forty English for academic purposes student essays were rated by GPT- 4 and four trained educators to assess nuances in rubric application, severity, leniency, and bias. Findings suggest that while GPT-4 tended to avoid the use of extreme scores, exhibiting a moderate central tendency rating, it does show a high level of consistency in its scoring behavior. This study contributes to understanding the extensions and limitations of using Generative AI tools in scoring essays, and provides insights into the use of AI tools in assessing writing.
Keywords
Full text
The full text of this article is available as a PDF download.
Download full article (PDF)How to cite
Pack, A., Carter, S., Barrett, A., Escalante, J., & Wolfersberger, M. (2026). An Exploratory Evaluation of GPT-4's Consistency as an English Essay Rater: A Many-Facet Rasch Model Analysis of AI versus Human Rating Patterns. International Journal of TESOL Studies, 1–22. Advance online publication. https://doi.org/10.58304/ijts.2560213

