Anchor is the key: Toward accessible automated essay scoring with large language model through prompting

Jaeyoon Choi University of California, Irvine ; Tamara Tate University of California, Irvine ; Mark Warschauer University of California, Irvine

Abstract

Automated Essay Scoring (AES) offers a scalable solution to the time-intensive and inconsistent nature of human scoring. While traditional AES systems require large sets of prompt-specific scored essays, large language models (LLMs) provide a powerful, adaptable alternative, capable of evaluating essays holistically without an extensive amount of pre-scored essays. However, most research on LLM-based AES focuses on resource-intensive optimization methods that are impractical for educators. In this study, we examine prompting – the most practical and accessible way for teachers to interact with LLMs – and its impact on holistic essay scoring. Using argumentative essays from secondary school students, we evaluate the effectiveness of incorporating grading rubrics, source materials, and anchor papers into prompts. Our results show that providing anchor papers significantly improves LLM-human agreement, bringing it closer to human-human scoring reliability. Moreover, while GPT-4o outperforms other models, GPT-4o mini achieves comparable results at a substantially lower cost. These findings highlight the potential of structured prompting strategies in enhancing the accuracy and accessibility of LLM-based AES in education. • Anchored prompts improve scoring reliability, nearing human-human reliability. • Rubric and exemplar prompts offer a low-resource alternative to AES. • GPT-4o is strongest; but GPT-4o mini gives similar accuracy at lower cost.

Journal
Assessing Writing
Published
2026-07-01
DOI
10.1016/j.asw.2026.101053
CompPile
Open Access
OA PDF Hybrid
Topics
Export

Citation context

Cited by in this index (1)

  1. Assessing Writing

References (40) · 1 in this index

  1. Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altma…
  2. Using automated essay scores as an anchor when equating constructed response writing tests
    International Journal of Testing  ↗
  3. Automated bahasa indonesia essay evaluation with latent semantic analysis
    Journal of Physics: Conference Series
  4. Nondeterminism of “deterministic” llm system settings in hosted environments
    Proceedings of the 5th workshop on evaluation and comparison of NLP systems
  5. Validity and reliability of automated essay scoring
    Handbook of automated essay evaluation
Show all 40 →
  1. Automated scoring in the era of artificial intelligence: An empirical study with turkish essays
    System  ↗
  2. Chatgpt as an automated essay scoring tool in the writing classrooms: How it compares wit…
    Education and Information Technologies  ↗
  3. An overview of automated scoring of essays
    The Journal of Technology, Learning and Assessment
  4. Rater types in writing performance assessments: A classification approach to rater variability
    Language Testing  ↗
  5. College Composition and Communication
  6. Tabllm: Few-shot classification of tabular data with large language models
    International conference on artificial intelligence and statistics
  7. Automated essay scoring systems
    Handbook of open, distance and digital education
  8. Automated essay scoring: A survey of the state of the art
    IJCAI
  9. Labrak, Y., Rouvier, M., & Dufour, R.(2024). A zero-shot and few-shot study of instruction-finetuned large la…
     ↗
  10. An application of hierarchical kappa-type statistics in the assessment of majority agreem…
    Biometrics  ↗
  11. Learning to write in middle school? insights into adolescent writers’ instructional exper…
    Journal of Adolescent & Adult Literacy  ↗
  12. Automated essay scoring: A reflection on the state of the art
    Proceedings of the 2024 conference on empirical methods in natural language processing
  13. A comprehensive review of automated essay scoring (aes) research and development
    Pertanika Journal of Science & Technology  ↗
  14. Can large language models automatically score proficiency of written essays?
    Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LRECCOLING 2024)
  15. Exploring the potential of using an ai language model for automated essay scoring
    Research Methods in Applied Linguistics  ↗
  16. Disciplinary literacy in history: An exploration of the historical nature of adolescents’…
    The Journal of the Learning Sciences  ↗
  17. Using writing tasks to elicit adolescents’ historical reasoning
    Journal of Literacy Research  ↗
  18. The critical role of anchor paper selection in writing assessment
    Applied Measurement in Education  ↗
  19. Large language models and automated essay scoring of english language learner writing: In…
    Computers and Education: Artificial Intelligence
  20. The analytic writing continuum: A comprehensive writing assessment system
  21. An automated essay scoring systems: A systematic literature review
    Artificial Intelligence Review  ↗
  22. Investigating neural architectures for short answer scoring
    Proceedings of the 12th workshop on innovative use of NLP for building educational applications
  23. Seßler, K., Fürstenberg, M., Bühler, B., & Kasneci, E. (2025, March). Can AI grade your essays? A comparative…
     ↗
  24. Shermis, M.D., & Burstein, J.C. (2003). Automated essay scoring: A crossdisciplinary perspective.
     ↗
  25. Automated essay scoring and revising based on open-source large language models
    IEEE Transactions on Learning Technologies  ↗
  26. Harnessing llms for multidimensional writing assessment: Reliability and alignment with h…
    Heliyon  ↗
  27. Can AI provide useful holistic essay scoring?
    Computers and Education: Artificial Intelligence
  28. The effects of prior computer use on computer-based writing: The 2011 NAEP writing assessment
    Computers & Education  ↗
  29. Effectiveness of large language models in automated evaluation of argumentative essays: F…
    Computer Assisted Language Learning
  30. Xiao, C., Ma, W., Xu, S.X., Zhang, K., Wang, Y., & Fu, Q. (2024). From automation to augmentation: Large lang…
  31. On protecting the data privacy of large language models (llms) and llm agents: A literatu…
    High-Confidence Computing  ↗
  32. 15 students initiating feedback. Feedback in second language writing
    Contexts and issues
  33. Yoon, S.-Y., Miszoglad, E., & Pierce, L.R. (2023). Evaluation of chatgpt feedback on ell writers’ coherence a…
  34. The impact of example selection in few-shot prompting on automated essay scoring using gp…
    International conference on artificial intelligence in education
  35. Advances in the field of automated essay evaluation
    Informatica