-
Bui (2025)
ChatGPT as an automated essay scoring tool in the writing classrooms: How it compares wit…
Education and Information Technologies
↗
-
Bui (2025)
Using generative artificial intelligence as an automated essay scoring tool: A comparativ…
Innovation in Language Learning and Teaching
↗
-
Chang (2024)
Making each point count: Revising a local adaptation of the Jacobs et al.’s (1981) ESL CO…
-
Chapelle (2008)
Building a validity argument for the test of English as a foreign language
-
Chen (2025)
A systematic review and meta-analysis of AI-enabled assessment in language learning: Desi…
Journal of Computer Assisted Learning
↗
-
DeVore (2025)
Exploring the ability of LLMs to classify written proficiency levels
Computer Speech & Language
↗
-
Fielding, R.T. (2000). Architectural styles and the design of network-based software architectures. [Unpublis…
-
Gao, T., Fisch, A., & Chen, D. (2021). Making pre-trained language models better few-shot learners. arXiv pre…
-
Gao (2025)
Assessing the reliability and relevance of DeepSeek in EFL writing evaluation: A generali…
Lang Test Asia
-
Giray (2023)
Prompt engineering with ChatGPT: A guide for academic writers
Annals of Biomedical Engineering
↗
-
Goldshtein (2024)
Automating bias in writing evaluation: Sources, barriers, and recommendations
The Routledge international handbook of automated essay evaluation
-
Hannah (2023)
Validity Arguments for Automated Essay Scoring of Young Students’ Writing Traits
Language Assessment Quarterly
↗
-
Huang (2023)
Trends, research issues and applications of artificial intelligence in language education
Educational Technology & Society
-
Huawei (2023)
A systematic review of automated writing evaluation systems
Education and Information Technologies
↗
-
Jin (2025)
When AI meets source use: Exploring ChatGPT's potential in L2 summary writing assessment
-
Kane (2013)
Validating the interpretations and uses of test scores
Journal of Educational Measurement
↗
-
Kim (2025)
Automated essay scoring with GPT-4 for a local placement test: Investigating prompting st…
TESOL Quarterly
-
Lan et al. (2025)
Assessing Writing
-
Landis (1977)
The measurement of observer agreement for categorical data
-
Linacre (2000)
Comparing “partial credit” and “rating scale” models
Rasch Measurement Transactions
-
Liu, X., Ji, K., Fu, Y., Tam, W.L., Du, Z., Yang, Z., & Tang, J. (2021). P-tuning v2: Prompt tuning can be co…
-
Manning (2025)
Human versus machine: The effectiveness of ChatGPT in automated essay scoring
Innovations in Education and Teaching International
↗
-
Marvin (2024)
Prompt engineering in large language models
International conference on data intelligence and cognitive informatics
-
Mayer (2023)
Prompt text classifications with transformer models: An exemplary introduction to prompt-…
Journal of Research on Technology in Education
↗
-
McNamara (2019)
Fairness, Justice, and Language Assessment
-
Mizumoto (2023)
Exploring the potential of using an AI language model for automated essay scoring
Research Methods in Applied Linguistics
↗
-
OpenAI (n.d.). Reponses API: Create a model response–temperature. OpenAI Platform. Retrieved from 〈https://pl…
-
OpenAI. (2023, November 6). ChatGPT—Release Notes: Introducing GPTs. Retrieved from 〈https://help.openai.com/…
-
Page (1966)
The imminence of grading essays by computer
Phi Delta Kappan
-
Poole (2024)
Can ChatGPT Reliably and Accurately Apply a Rubric to L2 Writing Assessments? The Devil i…
Journal of Technology & Chinese Language Teaching
-
Ramesh (2021)
An automated essay scoring system: a systematic literature review
Artificial Intelligence Review
↗
-
Rudner (2006)
An evaluation of IntelliMetric™ essay scoring system
The Journal of Technology, Learning and Assessment
-
(2024)
The Routledge international handbook of automated essay evaluation
-
Shin (2024)
Exploratory study on the potential of ChatGPT as a rater of second language writing
Education and Information Technologies
↗
-
Shrout (1979)
Intraclass correlations: uses in assessing rater reliability
-
Tate (2024)
Can AI provide useful holistic essay scoring?
Computers and Education: Artificial Intelligence
-
Wang (2024)
Effectiveness of large language models in automated evaluation of argumentative essays: f…
Computer Assisted Language Learning
-
Weigle (2013)
Assessing Writing
-
Williamson (2012)
A framework for evaluation and use of automated scoring
Educational Measurement: Issues and Practice
↗
-
Wilson (2021)
Elementary teachers’ perceptions of automated feedback and automated scoring: Transformin…
-
Xiao, H., Liu, Y., Li, J., & Zhou, Z. (2024). Human-AI collaborative essay scoring: A dual-process framework.…
-
Yamashita (2024)
An application of many-facet Rasch measurement to evaluate automated essay scoring: A cas…
Research Methods in Applied Linguistics
↗
-
Yamashita (2025)
Exploring potential biases in GPT-4o’s ratings of English language learners’ essays
-
Yancey (2023)
Rating short L2 essays on the CEFR scale with GPT-4
Proceedings of the 18th Workshop on innovative Use of NLP for Building Educational Applications (BEA 2023)
↗
-
Yin, W., Hay, J., & Roth, D. (2019). Benchmarking zero-shot text classification: Datasets, evaluation and ent…
-
Yun (2023)
Meta-Analysis of Inter-Rater Agreement and Discrepancy Between Human and Automated Englis…
-
Zhang, C., He, S., Li, L., Qin, S., Kang, Y., Lin, Q., Rajmohan, S. & Zhang, D. (2025). API agents vs. GUI ag…