Web of Science (ESCI)Scopus · SJR Q1ISSN 2632-6779 (Print)/2633-6898 (Online)
● Online First

An Exploratory Evaluation of GPT-4's Consistency as an English Essay Rater: A Many-Facet Rasch Model Analysis of AI versus Human Rating Patterns

Austin Pack1,Steven Carter1,Alex Barrett2,Juan Escalante1,Mark Wolfersberger3
  • 1 Brigham Young University-Hawaii, United States
  • 2 Florida State University, United States
  • 3 Brigham Young University, United States
Published: Online Firstpp. 1–22International Journal of TESOL StudiesOpen Access

Abstract

This study examined the defensibility of using GPT-4 for automated essay scoring, using a ManyFacet Rasch Model analysis. Forty English for academic purposes student essays were rated by GPT- 4 and four trained educators to assess nuances in rubric application, severity, leniency, and bias. Findings suggest that while GPT-4 tended to avoid the use of extreme scores, exhibiting a moderate central tendency rating, it does show a high level of consistency in its scoring behavior. This study contributes to understanding the extensions and limitations of using Generative AI tools in scoring essays, and provides insights into the use of AI tools in assessing writing.

Keywords

Full text

The full text of this article is available as a PDF download.

Download full article (PDF)

How to cite

Citation · APA

Pack, A., Carter, S., Barrett, A., Escalante, J., & Wolfersberger, M. (2026). An Exploratory Evaluation of GPT-4's Consistency as an English Essay Rater: A Many-Facet Rasch Model Analysis of AI versus Human Rating Patterns. International Journal of TESOL Studies, 1–22. Advance online publication. https://doi.org/10.58304/ijts.2560213