Assessing Writing
1,070 articlesOctober 2026
-
Abstract
This study investigates developmental changes in grammatical subjects in the English writing of Japanese learners, focusing on subject position as an index of syntactic growth. Previous L2 assessment research has emphasized clausal subordination and global complexity indices; noun phrase complexity has attracted growing attention, but the subject position itself has rarely been isolated as a unit of analysis. Drawing on a longitudinal corpus of 26,399 texts written by 1,086 Japanese high school students over three years, subjects were automatically extracted and annotated using SubjEx, a dependency-parsing system developed for this purpose (accuracy: 93.5%). Mean subject length increased steadily across grades and CEFR proficiency bands, with clearer differentiation from B1 onward; a mixed-effects model with a random slope further indicates increasing divergence in learners’ growth trajectories. Subjects shifted from pronoun-dominant forms toward more complex noun phrases, with greater use of adjectival modification, possessives, compound nouns, and prepositional phrases, while lexical analyses showed a decline in first-person subjects and a rise in informationally dense nominal patterns. Subject position thus emerges as a sensitive indicator of L2 writing development, reflecting a shift from personal expression toward condensed informational packaging, with implications for complexity research and proficiency assessment.
-
Effects of varying time constraints on quality, linguistic features, and fluency behaviors in L2 writing ↗
Abstract
Although previous research has explored how L2 writers allocate their time and how time limits influence writing scores, many of these studies focus solely on final texts, providing only a partial understanding of the writing process. A more comprehensive investigation is essential for enhancing the construct validity of writing assessments. Drawing on empirical research on timed writing and on theoretical models of the writing process, this study explores how time constraints affect L2 learners' written texts and fluency behaviors. A total of 123 participants (60 high-intermediate and 63 advanced EFL learners) completed an argumentative writing task with keystroke logging software under either a 30-minute or 60-minute condition, and stimulated recalls were conducted with a subset. The texts were analyzed for syntactic and lexical complexity, fluency, and overall quality. Results indicate that time constraints did not significantly affect linguistic complexity or writing quality, although fluency behaviors differed across conditions. Time constraints thus affected learners' writing processes more than the linguistic quality of their texts. The core implication for L2 writing assessment is that, for test takers experienced with timed writing, shortening the time allotted to an argumentative task changes composing processes without compromising the product-based scores on which assessment decisions rest.
-
Abstract
Previous studies on L2 phrase complexity predominantly focused on noun phrases, while the verb phrase complexity are few and mostly simple verb phrase structures, they have limited expressive capacity. Complex verb phrase structures (e.g., n-grams, VACs) also have limitations. The present study proposes a set of 41 Chinese complex verb phrase structures and are meaningful. A total of 246 verb phrase complexity measures are calculated from the dimensions of account, frequency, diversity, and density. The ability of these measures to predict writing quality is compared with that of 10 large-grained syntactic complexity measures. The results show that large-grained measures (Average sentence length) and verb phrase complexity measures (two verb-object structures and three adverbial-head structures) can respectively explain 14.4% and 41.9% of variance in writing scores. Our results illustrate the importance of Chinese complex verb phrase structures in assessing L2 writing quality.
-
Replicating and expanding a validity argument for eRevise as a formative assessment: Empirical support for theories of conceptual change ↗
Abstract
As automated writing evaluation (AWE) systems proliferate, it is important to assess them for the extent to which they serve an authentic formative assessment purpose. We developed eRevise as an AWE system to engage upper elementary students in text-based writing, receive automated feedback, and revise their essay based on that feedback. In past work, we presented a validity argument for a response-to-text formative assessment. Here, we replicate evidence for the mediational processes we identified with a second response-to-text formative assessment. Beyond replication, we expand the validity argument in several ways: We examine multiple writing outcomes to understand whether the relationships in the data generalize. We also explore patterns of relationships to better understand which students are most likely to benefit from eRevise . Furthermore, we expand the investigation from feature score improvement alone to also consider students’ conceptual development over time. Our findings provide further evidence for sociocultural mechanisms supporting students’ conceptual development of evidence use in writing. We discuss the implications of our findings for future design of eRevise and AWE systems, in general. We also discuss the need for a recursive process where our findings also contribute to refined theory development and design of future AWE systems. (1). Formative Assessment; (2) Argumentative Writing; (3) Adaptive Expertise; (4) Conceptual Change; (5) Validity Argument.
-
Student-AI collaboration in peer feedback: Effects on perceived feedback quality, emotional responses, and feedback literacy development ↗
Abstract
This study investigates how English as a foreign language (EFL) student reviewers engage in open-ended, dialogic interactions with generative artificial intelligence (AI) during the feedback generation process in peer assessment within EFL writing classrooms. It examines the impact of these interactions on feedback quality perceived by recipients, emotional responses (task enjoyment and anxiety), and feedback literacy. A quasi-experimental design was employed with 60 Chinese undergraduate students, divided into an experimental group (EG) that used generative AI (Doubao) for support and a control group (CG) that did not. Over three intervention cycles, data from chat histories, feedback quality ratings by recipients, and pre/post questionnaires on emotions and feedback literacy were analyzed. The results indicated that EG students primarily employed AI for linguistic refinement of their comments, with limited use for enhancing the content or structure. Nevertheless, AI support led to significant, progressive improvements in the perceived quality of feedback, particularly in affect, description, justification, and constructiveness. Furthermore, EG students reported significantly higher task enjoyment and lower anxiety compared to the CG. The intervention also positively enhanced all dimensions of feedback literacy: knowledge and abilities, willingness to participate, cooperative learning, and appreciation of peer feedback. The findings suggest that generative AI can serve as a powerful scaffold, reducing the emotional and cognitive burdens of peer assessment while fostering a more supportive and effective feedback environment. This study underscores the value of integrating AI into peer feedback practices to develop students’ feedback literacy and improve the overall quality of peer learning experiences.
-
Abstract
Feedback agency is a key concept in enhancing student writing performance. While growing attention has been paid to student feedback agency, existing research remains largely theoretical and qualitative. As a result, there is a lack of psychometrically supported instruments to measure this construct. To address this gap, the present two-phase study aimed to develop and evaluate a Student Feedback Agency Scale (SFAS), drawing on social cognitive theory. In the development stage, Principal Component Analysis (PCA) was conducted on a sample of 235 Chinese international postgraduate students. The results yielded a 23-item SFAS comprising six components: Action Taking, Goal Setting, Processing, Generating, Self-efficacy, and Seeking. Using an independent sample of 349 participants from the same population, Confirmatory Factor Analysis (CFA) was conducted in the evaluation stage. The results supported a good model fit (RMSEA = .055, IFI = .923, TLI = .908, and CFI = .922). Multi-group CFAs further confirmed the structural invariance across gender, academic level, and discipline. Overall, the findings provide psychometric evidence to support the interpretation and use of the SFAS scores to measure student agency in writing feedback processes. Based on these results, the factor structure and subscales of the SFAS are discussed, and implications are outlined.
July 2026
-
Abstract
With the advent of artificial intelligence, large language model (LLM) based Automated Essay Scoring (AES) systems have been developed that can consistently make human-like decisions that do not depend fully on surface level linguistic features. However, research into the use of LLM-based AES systems is limited and little is known about the reliability, agreement, or validity of the systems. The goal of this study was to provide evidence for the reliability, agreement, and validity of LLM-based AES systems in a standardized writing assessment used for secondary school students. Both representation and generative LLM-based AES systems were developed to score persuasive essays and assessed for reliability. Then the agreement of the developed AES systems with human raters was assessed through correlational analyses. We used extrinsic convergent validation approaches to examine if the human and LLM scores correlated with linguistic components. Results indicate strong reliability and agreement for the LLM scores. In terms of convergent validity, initial correlational analyses indicated that the representation LLM AES system showed differential correlations with the human scores in terms of a text length and type-token ratio component. This result contrasts with the correlational results from the generative LLM AES model, which indicated no differences in associations between the model and human scores with regards to the linguistic components.
-
From keystrokes to scores: Toward a multidimensional predictive model of writing evaluation by humans and large language models across linguistic, cognitive, and social dimensions ↗
Abstract
Automated writing evaluation (AWE) has traditionally emphasized textual features such as vocabulary and syntax, while often overlooking writers’ social identities and cognitive behaviors – factors central to understanding writing as a multidimensional construct. With the increasing integration of large language models (LLMs) into AWE, questions remain about how their assessments align with human judgments and the sources of potential divergences. This study investigates how linguistic (e.g., lexical diversity), cognitive (e.g., pausing behavior), and social (e.g., gender) factors covary with essay scores assigned by human raters and LLMs. We analyzed 4245 argumentative essays paired with demographic metadata and keystroke-logging data, using correlation analyses, random forest models, and regression-based approaches to examine relationships among writer characteristics, writing-process features, textual features, and essay scores. Results showed moderate agreement between human and LLM scores, but the two scoring systems exhibited different patterns of association with linguistic, cognitive, and social variables. These findings suggest that human and LLM evaluations rely on partially different cues and demonstrate how socio-cognitive metadata can be used to examine the factors associated with writing assessment decisions. By moving beyond text-only comparisons, this approach provides a complementary lens for understanding why and how human and machine judgments converge or diverge.
-
Examining linguistic reasoning and metalinguistic strategies in L2 writing: Insights from teacher- and AI-feedback revisions ↗
Abstract
Second language (L2) writers’ capacity to reason about language choices during revision is theoretically central to advanced writing development, yet few studies trace how this linguistic reasoning emerges in repeated feedback-revision cycles that integrate teacher and AI suggestions. This multi-case, longitudinal mixed-methods study examined how ten L2 learners enacted metalinguistic strategies and developed linguistic reasoning across a 12-week online Academic English program. Triangulated data sources included 1246 feedback revision episodes, pre-post reasoning-task responses, and semi-structured exit interviews. We operationalized linguistic reasoning as the processual construct of interest and treated metalinguistic strategies as observable indicators; reasoning responses were scored with a three-dimensional analytic rubric while revisions and interaction logs were coded thematically and by strategy. Results show a systematic shift from early, correctness-focused edits toward later revisions characterized by coordinated structural comparison, more explicit self-explanation, and integrative argumentation; composite reasoning scores increased, with the largest gains observed on depth of explanation. Process analyses identified recurrent patterns in which alternative-rich feedback coincided with comparison moves, partially divergent suggestions often coincided with more explicit self-explanation, and integrative argumentation involving discourse-level revision became more visible over time. Pedagogical implications are also discussed.
-
Examining the relevance of three TOEFL Essentials writing tasks to the accounting profession: The role of domain experts ↗
Abstract
Large-scale English language proficiency tests are increasingly used to make decisions about professional registration despite not originally being developed to make predictions about language use in the workplace. Domain experts can play a valuable role as informants in establishing the relevance of test tasks to a specific TLU domain. However, this practice has seldom been critically examined. In particular, no studies to date have examined the nature of the interview questions used when engaging with domain experts in this type of research. The current study was designed with two aims: to (1) explore the relevance of three writing tasks (i.e., Build a Sentence, Write for an Academic Discussion, Write an Email) from the TOEFL Essentials test to the accounting profession and (2) evaluate the judgements of domain expert participants. Twenty accountants from non-English speaking backgrounds as well as three accounting educators were interviewed for the study, drawing on a methodology with broad, general questions. The data was analysed qualitatively to identify (a) to what extent the participants considered the three tasks relevant and (b) what task features they attended to when commenting on the relevance of the tasks. The findings showed that the participants generally found the Build a Sentence the least relevant of the three writing tasks, and the Write an Email task the most relevant. When reviewing the task features the participants judged as relevant to writing demands in their workplace, it was shown they focused on a small/narrow range of features. They also engaged in ‘misinterpretations’, comparing aspects of test and workplace tasks that did not align (e.g., comparing a writing task to events in a spoken meeting). The findings are discussed in terms of domain expert involvement in validation research.
-
Abstract
Existing scholarship abounds in pedagogical efforts to leverage the advantages of GenAI to support argumentative writing. Yet, these efforts have mainly positioned learners as consumers of GenAI feedback and support on writing, which is often associated with concerns of cognitive atrophy and metacognitive laziness. The emergence of AI-assisted vibe coding makes it possible for learners to play a more active role in creating personalized AI solutions. The repositioning of learners’ roles in relation to GenAI may induce deeper cognitive processing and greater metacognitive engagement. This article introduces AI-assisted vibe coding, a novel no-code development approach that enables learners as technology creators by designing and developing their own personalized argumentative writing tools. We present Base44, a full-stack, AI-powered platform that facilitates this process, and illustrate how students can leverage it to create personalized tools for addressing authentic argumentative writing challenges through concrete workflow examples. We explore the potential of integrating AI-assisted vibe coding into argumentative writing instruction, examine its limitations, and offer directions for future teaching practice and research.
-
Abstract
GPTZero is an AI detection platform that scans written text for statistical signatures of machine generation and returns a probability score estimating whether it was produced by a human or an AI. In higher education, many teachers have turned to AI detection as a first-line response to the integrity crisis triggered by large language models. However, empirical findings on GPTZero’s efficacy are notably mixed. Some studies report strong diagnostic value under controlled conditions, while others document substantial false-negative rates, near-random performance on certain AI-generated essays, and frequent misclassification of AI-translated texts across several languages. Multilingual and L2 writers often bear the greatest cost, as their carefully constructed English is sometimes assigned high AI-likelihood scores because their linguistic profiles may appear less natural to models trained predominantly on standard or formulaic patterns of written English. In developing countries, where students commonly write in English as a second or third language, these limitations represent more than minor technical issues; they raise concerns about equity, potentially placing disproportionate burdens on writers working to meet academic language expectations. This article argues that GPTZero is unsuitable as a definitive tool for high-stakes assessment of writing. Instead, it proposes a shift toward postplagiarism frameworks that recognize responsible AI use. Within this approach, AI detection outputs serve as formative resources for developing critical AI literacy rather than surveillance tools. Flagged content becomes a starting point for metacognitive dialogue, which supports trust-based pedagogies that emphasize student agency and intellectual accountability.
-
Abstract
This study investigates intraindividual variability (IIV) in curriculum-based measurements of writing (CBM-W) among primary school children. Inspired by the idea that such fluctuations may not only represent measurement error but also warrant description in relation to writing performance and development, the study examines patterns of IIV in this context. Data were collected from 345 children (51.6% female) in Grades 3–6 in Switzerland (mean age 10;5), including both monolinguals and multilinguals learning German. Students produced 10 writing samples at two time points (fall, spring), scored for correct writing sequences. The coefficient of variation quantified IIV at both times. Results show that IIV is systematically associated with performance level, with higher variability among lower-performing students. No clear age-related pattern was found once performance level was taken into account. IIV did not emerge as a predictor of subsequent writing development among students with comparable initial performance. In addition, students’ home language background did not moderate the association between IIV and subsequent writing development. These findings suggest that IIV in CBM-W is closely related to performance level and highlight the importance of interpreting variability in relation to students’ performance, calling for caution when using CBM-W in educational decision-making.
-
Assessing fairness in AI-assisted writing scoring: Developing fairness measures to detect predictive bias in automated essay scoring ↗
Abstract
Automated essay scoring (AES) is increasingly utilized in educational settings, yet concerns about its fairness persist. This study reviews current fairness measures in AES and summarizes their respective strengths and weaknesses. Drawing on principles from educational and psychological testing, we introduce two measures for detecting potential predictive bias: conditional disparity ratio and conditional disparity difference. Our method emphasizes two key principles: first, that bias should be assessed among students with comparable proficiency levels, and second, that evaluations should be conducted on a test set independent of the AES training set. We demonstrated this approach using writing samples from the Facial Action Coding System task within the PERSUADE 2.0 corpus to assess potential predictive bias related to sex and race. Four AES models were evaluated for predictive bias: ordinal logistic regression using TF–IDF features, fine-tuned BERT, and ChatGPT in both zero-shot and few-shot settings. The findings indicated that, without accounting for proficiency, subgroup differences remained ambiguous, making it difficult to detect potential predictive bias. In contrast, conditioning on proficiency revealed clearer and more interpretable patterns of bias. The discussion addresses key factors and the extension of the bias detection framework and outlines future directions for bias mitigation. • Introduces two fairness measures for automated writing scoring. • Distinguishes predictive bias from real proficiency differences. • Uses multiple scoring models, from machine learning to large language models. • Shows fairness varies by demographic group and proficiency level. • Offers practical guidance for bias detection in AI writing assessment.