Abstract

Automated writing evaluation (AWE) has traditionally emphasized textual features such as vocabulary and syntax, while often overlooking writers’ social identities and cognitive behaviors – factors central to understanding writing as a multidimensional construct. With the increasing integration of large language models (LLMs) into AWE, questions remain about how their assessments align with human judgments and the sources of potential divergences. This study investigates how linguistic (e.g., lexical diversity), cognitive (e.g., pausing behavior), and social (e.g., gender) factors covary with essay scores assigned by human raters and LLMs. We analyzed 4245 argumentative essays paired with demographic metadata and keystroke-logging data, using correlation analyses, random forest models, and regression-based approaches to examine relationships among writer characteristics, writing-process features, textual features, and essay scores. Results showed moderate agreement between human and LLM scores, but the two scoring systems exhibited different patterns of association with linguistic, cognitive, and social variables. These findings suggest that human and LLM evaluations rely on partially different cues and demonstrate how socio-cognitive metadata can be used to examine the factors associated with writing assessment decisions. By moving beyond text-only comparisons, this approach provides a complementary lens for understanding why and how human and machine judgments converge or diverge.

Journal
Assessing Writing
Published
2026-07-01
DOI
10.1016/j.asw.2026.101090
CompPile
Open Access
OA PDF Hybrid
Topics
Export

Citation context

Cited by in this index (0)

No articles in this index cite this work.

References (87) · 12 in this index

  1. The secret life of connectives: A taxonomy to study individual differences in mid-adolesc…
    Reading and Writing  ↗
  2. What can L2 writers’ pausing behavior tell us about their L2 writing processes?
    Studies in Second Language Acquisition  ↗
  3. Modeling local coherence: An entity-based approach
    Computational Linguistics  ↗
  4. Fitting linear mixed-effects models using lme4
    Journal of Statistical Software  ↗
  5. Automated evaluation of writing – 50 years and counting
    Proceedings of the 58th annual meeting of the association for computational linguistics
Show all 87 →
  1. Contrast coding choices in a decade of mixed models
    Journal of Memory and Language  ↗
  2. Random forests
    Machine Learning  ↗
  3. Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and …
    Applied Measurement in Education  ↗
  4. ChatGPT as an automated essay scoring tool in the writing classrooms: How it compares wit…
    Education and Information Technologies  ↗
  5. Contact, the feature pool and the speech community: The emergence of Multicultural London…
    Journal of Sociolinguistics  ↗
  6. Negative concord in child African American English: Implications for specific language im…
    Journal of Speech, Language, and Hearing Research  ↗
  7. Developing linguistic constructs of text readability using Natural Language Processing
    Scientific Studies of Reading  ↗
  8. Plagiarism detection using keystroke logs
    Proceedings of the 17th international conference on educational data mining
  9. Assessing Writing
  10. Using writing process and product features to assess writing quality and explore how thos…
    ETS Research Report Series
  11. Three waves of variation study: The emergence of meaning in the study of sociolinguistic …
    Annual Review of Anthropology  ↗
  12. Complex dynamic systems theory and L2 writing development
  13. Aligning keystrokes with cognitive processes in writing
    Observing writing
  14. Ethnic variation in its social context: Evidence from the English of Chinese Australians
    PhD Thesis
  15. A multi-dialectal, longitudinal corpus of human-AI hybrid language production
    Proceedings of the fifteenth language resources and evaluation conference (LREC 2026)  ↗
  16. Scaling regression inputs by dividing by two standard deviations
    Statistics in Medicine  ↗
  17. Statistics for Linguistics with R: A practical introduction
  18. Writing process differences in subgroups reflected in keystroke logs
    Journal of Educational and Behavioral Statistics  ↗
  19. Halliday's Introduction to Functional Grammar
  20. The elements of statistical learning: Data mining, inference, and prediction
  21. Beyond final products: Multi-dimensional essay scoring using keystroke logs and deep learning
    LAK25: The 15th international learning analytics and knowledge conference
  22. The Hewlett Foundation automated essay scoring
    The William and Flora Hewlett Foundation
  23. Genre pedagogy: Language, literacy and L2 writing instruction
    Journal of Second Language Writing  ↗
  24. Capturing the diversity in lexical diversity
    Language Learning  ↗
  25. Ke, Z., & Ng, V. (2019). Automated essay scoring: A survey of the state of the art. Proceedings of the 28th I…
     ↗
  26. A new measure of rank correlation
    Biometrika  ↗
  27. Do large language models produce texts with "human-like" lexical diversity? Evidence from…
    International Journal of Applied Linguistics, Online first
  28. Understanding the dynamics of second language writing through keystroke logging and compl…
    Proceedings of the twelfth language resources and evaluation conference
  29. Assessing syntactic sophistication in L2 writing: A usage-based approach
    Language Testing  ↗
  30. Assessing the validity of lexical diversity using direct judgements
    Language Assessment Quarterly  ↗
  31. Negative attraction and negative concord in English grammar
    Language  ↗
  32. The intersection of sex and social class in the course of linguistic change
    Language Variation and Change  ↗
  33. Principles of linguistic change
    Volume 2: Social factors
  34. Assessing Writing
  35. A complex dynamic systems perspective to researching language classroom dynamics
    Research Questions in Language Education and Applied Linguistics
  36. Larsen-Freeman, D., & Cameron, L. (2008). Complex systems and applied linguistics. Oxford University.
  37. A human-centric automated essay scoring and feedback system for the development of ethica…
    Educational Technology & Society
  38. Written Communication
  39. Lenth, R.V. (2025). emmeans: Estimated marginal means, aka least-squares means. R package version 1.11.1. 〈ht…
  40. Written Communication
  41. Assessing Writing
  42. The many dimensions of algorithmic fairness in educational applications
    Proceedings of the fourteenth workshop on innovative use of NLP for building educational applications
  43. Automatic analysis of syntactic complexity in second language writing
    International Journal of Corpus Linguistics  ↗
  44. A corpus-based evaluation of syntactic complexity measures as indices of college-level ES…
    TESOL Quarterly  ↗
  45. Lexical diversity and language development: Quantification and assessment
  46. Can large language models automatically score proficiency of written essays?
    Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation
  47. vocd: A theoretical and empirical evaluation
    Language Testing  ↗
  48. MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversi…
    Behavior Research Methods  ↗
  49. Written Communication
  50. Assessing Writing
  51. Exploring the potential of using an AI language model for automated essay scoring
    Research Methods in Applied Linguistics  ↗
  52. Journal of Writing Research
  53. The gender-linked language effect: an empirical test of a general process model
    Language Sciences  ↗
  54. The effectiveness of automated writing evaluation in EFL/ESL writing: A three-level meta-…
    Interactive Learning Environments  ↗
  55. Dissecting racial bias in an algorithm used to manage the health of populations
    Science  ↗
  56. Can the probability distribution of dependency distance measure language proficiency of s…
    Journal of Quantitative Linguistics
  57. Assessing Writing
  58. The imminence of… grading essays by computer
    The Phi Delta Kappan
  59. Validating automated essay scoring: A (modest) refinement of the “gold standard”
    Applied Measurement in Education  ↗
  60. Exploring second language writers' pausing and revision behaviors: A mixed-methods study
    Studies in Second Language Acquisition  ↗
  61. Exploring the relationship of working memory to the temporal distribution of pausing and …
    Studies in Second Language Acquisition  ↗
  62. Estimating the dimension of a model
    Annals of Statistics  ↗
  63. Prediction of essay scores from writing process and product features using data mining methods
    Applied Measurement in Education  ↗
  64. Revising in two languages: A multi-dimensional comparison of online writing revisions in …
    Journal of Second Language Writing  ↗
  65. An introduction to recursive partitioning: Rationale, application, and characteristics of…
    Psychological Methods  ↗
  66. A comparative study of the human, automated scoring model, and GPT-4 ratings of young EFL…
    Language Testing  ↗
  67. Journal of Writing Research
  68. Exploring the application of keystroke logging techniques to research in second language …
    Research Methods in Applied Linguistics  ↗
  69. The intersection of ethnicity and social class in language variation and change
    Language Variation and Change  ↗
  70. Uto, M., Xie, Y., & Ueno, M. (2020). Neural automated essay scoring incorporating hand crafted features. Proc…
     ↗
  71. Computers and Composition
  72. Voeten, C.C. (2025). buildmer: Stepwise Elimination and Term Reordering for Mixed-Effects Regression. R packa…
  73. Best practice in statistics: The use of log transformation
    Annals of Clinical Biochemistry  ↗
  74. Trimming and winsorization
    Encyclopedia of Biostatistics  ↗
  75. ranger: A fast implementation of Random Forests for highdimensional data in C++ and R
    Journal of Statistical Software  ↗
  76. Writing in foreign language contexts: Learning, teaching, and research
    Multilingual Matters
  77. An accurate computation of the hypergeometric distribution function
    ACM Transactions on Mathematical Software  ↗
  78. Human-AI collaborative essay scoring: A dual-process framework with LLMs
    LAK25: The 15th international learning analytics and knowledge conference (LAK 2025)
  79. Exploring potential biases in GPT-4o’s ratings of English language learners’ essays
    Language Testing  ↗
  80. Rating short L2 essays on the CEFR scale with GPT-4
    Proceedings of the 18th workshop on innovative use of NLP for building educational applications (BEA 2023)
  81. Assessing Writing
  82. The effectiveness of automated writing evaluation on writing quality: A meta-analysis
    Journal of Educational Computing Research  ↗