Amery D. Wu

1 article
University of British Columbia

Loading profile…

Publication timeline

Co-author network

Research topics

Who reads Wu

Amery D. Wu's work travels primarily in Composition & Writing Studies (100% of indexed citations) · 1 indexed citations.

By cluster

  • Composition & Writing Studies — 1

Top citing journals

Counts include only citations from indexed journals that deposit reference lists with CrossRef. Authors whose readers publish primarily in venues without reference deposits will appear less central than they are. See coverage notes →

  1. Assessing fairness in AI-assisted writing scoring: Developing fairness measures to detect predictive bias in automated essay scoring ↗
    Abstract

    Automated essay scoring (AES) is increasingly utilized in educational settings, yet concerns about its fairness persist. This study reviews current fairness measures in AES and summarizes their respective strengths and weaknesses. Drawing on principles from educational and psychological testing, we introduce two measures for detecting potential predictive bias: conditional disparity ratio and conditional disparity difference. Our method emphasizes two key principles: first, that bias should be assessed among students with comparable proficiency levels, and second, that evaluations should be conducted on a test set independent of the AES training set. We demonstrated this approach using writing samples from the Facial Action Coding System task within the PERSUADE 2.0 corpus to assess potential predictive bias related to sex and race. Four AES models were evaluated for predictive bias: ordinal logistic regression using TF–IDF features, fine-tuned BERT, and ChatGPT in both zero-shot and few-shot settings. The findings indicated that, without accounting for proficiency, subgroup differences remained ambiguous, making it difficult to detect potential predictive bias. In contrast, conditioning on proficiency revealed clearer and more interpretable patterns of bias. The discussion addresses key factors and the extension of the bias detection framework and outlines future directions for bias mitigation. • Introduces two fairness measures for automated writing scoring. • Distinguishes predictive bias from real proficiency differences. • Uses multiple scoring models, from machine learning to large language models. • Shows fairness varies by demographic group and proficiency level. • Offers practical guidance for bias detection in AI writing assessment.

    doi:10.1016/j.asw.2026.101066