All Journals

1,585 articles
Year – Topic Clear
Export
assessment×

October 2026

  1. Generative artificial intelligence in writing assessment: University professors’ perspectives and recommendations ↗
    doi:10.1016/j.asw.2026.101125
  2. Collaborative agency in rating scale development for L2 academic writing assessment ↗
    doi:10.1016/j.asw.2026.101126
  3. Developmental trends in subject length and complexity in Japanese learners’ English writing ↗
    Abstract

    This study investigates developmental changes in grammatical subjects in the English writing of Japanese learners, focusing on subject position as an index of syntactic growth. Previous L2 assessment research has emphasized clausal subordination and global complexity indices; noun phrase complexity has attracted growing attention, but the subject position itself has rarely been isolated as a unit of analysis. Drawing on a longitudinal corpus of 26,399 texts written by 1,086 Japanese high school students over three years, subjects were automatically extracted and annotated using SubjEx, a dependency-parsing system developed for this purpose (accuracy: 93.5%). Mean subject length increased steadily across grades and CEFR proficiency bands, with clearer differentiation from B1 onward; a mixed-effects model with a random slope further indicates increasing divergence in learners’ growth trajectories. Subjects shifted from pronoun-dominant forms toward more complex noun phrases, with greater use of adjectival modification, possessives, compound nouns, and prepositional phrases, while lexical analyses showed a decline in first-person subjects and a rise in informationally dense nominal patterns. Subject position thus emerges as a sensitive indicator of L2 writing development, reflecting a shift from personal expression toward condensed informational packaging, with implications for complexity research and proficiency assessment.

    doi:10.1016/j.asw.2026.101124
  4. Effects of varying time constraints on quality, linguistic features, and fluency behaviors in L2 writing ↗
    Abstract

    Although previous research has explored how L2 writers allocate their time and how time limits influence writing scores, many of these studies focus solely on final texts, providing only a partial understanding of the writing process. A more comprehensive investigation is essential for enhancing the construct validity of writing assessments. Drawing on empirical research on timed writing and on theoretical models of the writing process, this study explores how time constraints affect L2 learners' written texts and fluency behaviors. A total of 123 participants (60 high-intermediate and 63 advanced EFL learners) completed an argumentative writing task with keystroke logging software under either a 30-minute or 60-minute condition, and stimulated recalls were conducted with a subset. The texts were analyzed for syntactic and lexical complexity, fluency, and overall quality. Results indicate that time constraints did not significantly affect linguistic complexity or writing quality, although fluency behaviors differed across conditions. Time constraints thus affected learners' writing processes more than the linguistic quality of their texts. The core implication for L2 writing assessment is that, for test takers experienced with timed writing, shortening the time allotted to an argumentative task changes composing processes without compromising the product-based scores on which assessment decisions rest.

    doi:10.1016/j.asw.2026.101123
  5. From confetti-like products to languaging ontologies: A bibliometric retrospective on two decades of L2 writing assessment (2006–2025) ↗
    doi:10.1016/j.asw.2026.101122
  6. AI and writing assessment: Innovations, validity, and ethical considerations ↗
    doi:10.1016/j.asw.2026.101120
  7. Evaluating the role of AI in writing assessment development through the lens of the validity argument ↗
    doi:10.1016/j.asw.2026.101107
  8. PEUMO: An AI-based writing assessment and support tool for engineering report writing in Spanish ↗
    doi:10.1016/j.asw.2026.101106
  9. Replicating and expanding a validity argument for eRevise as a formative assessment: Empirical support for theories of conceptual change ↗
    Abstract

    As automated writing evaluation (AWE) systems proliferate, it is important to assess them for the extent to which they serve an authentic formative assessment purpose. We developed eRevise as an AWE system to engage upper elementary students in text-based writing, receive automated feedback, and revise their essay based on that feedback. In past work, we presented a validity argument for a response-to-text formative assessment. Here, we replicate evidence for the mediational processes we identified with a second response-to-text formative assessment. Beyond replication, we expand the validity argument in several ways: We examine multiple writing outcomes to understand whether the relationships in the data generalize. We also explore patterns of relationships to better understand which students are most likely to benefit from eRevise . Furthermore, we expand the investigation from feature score improvement alone to also consider students’ conceptual development over time. Our findings provide further evidence for sociocultural mechanisms supporting students’ conceptual development of evidence use in writing. We discuss the implications of our findings for future design of eRevise and AWE systems, in general. We also discuss the need for a recursive process where our findings also contribute to refined theory development and design of future AWE systems. (1). Formative Assessment; (2) Argumentative Writing; (3) Adaptive Expertise; (4) Conceptual Change; (5) Validity Argument.

    doi:10.1016/j.asw.2026.101095
  10. Student-AI collaboration in peer feedback: Effects on perceived feedback quality, emotional responses, and feedback literacy development ↗
    Abstract

    This study investigates how English as a foreign language (EFL) student reviewers engage in open-ended, dialogic interactions with generative artificial intelligence (AI) during the feedback generation process in peer assessment within EFL writing classrooms. It examines the impact of these interactions on feedback quality perceived by recipients, emotional responses (task enjoyment and anxiety), and feedback literacy. A quasi-experimental design was employed with 60 Chinese undergraduate students, divided into an experimental group (EG) that used generative AI (Doubao) for support and a control group (CG) that did not. Over three intervention cycles, data from chat histories, feedback quality ratings by recipients, and pre/post questionnaires on emotions and feedback literacy were analyzed. The results indicated that EG students primarily employed AI for linguistic refinement of their comments, with limited use for enhancing the content or structure. Nevertheless, AI support led to significant, progressive improvements in the perceived quality of feedback, particularly in affect, description, justification, and constructiveness. Furthermore, EG students reported significantly higher task enjoyment and lower anxiety compared to the CG. The intervention also positively enhanced all dimensions of feedback literacy: knowledge and abilities, willingness to participate, cooperative learning, and appreciation of peer feedback. The findings suggest that generative AI can serve as a powerful scaffold, reducing the emotional and cognitive burdens of peer assessment while fostering a more supportive and effective feedback environment. This study underscores the value of integrating AI into peer feedback practices to develop students’ feedback literacy and improve the overall quality of peer learning experiences.

    doi:10.1016/j.asw.2026.101100
  11. Modeling the reading-to-writing pipeline: Knowledge graph and LLM-based assessment framework for source-based writing ↗
    doi:10.1016/j.asw.2026.101098
  12. Development and evaluation of a student feedback agency scale ↗
    Abstract

    Feedback agency is a key concept in enhancing student writing performance. While growing attention has been paid to student feedback agency, existing research remains largely theoretical and qualitative. As a result, there is a lack of psychometrically supported instruments to measure this construct. To address this gap, the present two-phase study aimed to develop and evaluate a Student Feedback Agency Scale (SFAS), drawing on social cognitive theory. In the development stage, Principal Component Analysis (PCA) was conducted on a sample of 235 Chinese international postgraduate students. The results yielded a 23-item SFAS comprising six components: Action Taking, Goal Setting, Processing, Generating, Self-efficacy, and Seeking. Using an independent sample of 349 participants from the same population, Confirmatory Factor Analysis (CFA) was conducted in the evaluation stage. The results supported a good model fit (RMSEA = .055, IFI = .923, TLI = .908, and CFI = .922). Multi-group CFAs further confirmed the structural invariance across gender, academic level, and discipline. Overall, the findings provide psychometric evidence to support the interpretation and use of the SFAS scores to measure student agency in writing feedback processes. Based on these results, the factor structure and subscales of the SFAS are discussed, and implications are outlined.

    doi:10.1016/j.asw.2026.101096

September 2026

  1. Learning How to Write as a Cyborg ↗
    Abstract

    Discussions of generative artificial intelligence (GenAI) focus on describing AI as either a tool or a collaborator. Such discussions do not fully grasp GenAI's transformative impact. This article proposes a model of cyborg learning that outlines the skills necessary for using GenAI effectively: background knowledge, critical evaluation of AI output, and rhetorical integration of AI output. Presenting multiple use cases from the literature in technical and professional communication and the author's own experience, the article illustrates how to write as a cyborg.

    doi:10.1177/10506519261485275
  2. The Implicit Audience Problem in Computational Argument Quality Assessment: A Pragmatic Critique and the ASAQ Framework ↗
    doi:10.1007/s10503-026-09734-y
  3. Gender Differences in Nonverbals of Business Undergraduates: Self-Assessment Insights of Perceived Credibility ↗
    Abstract

    This study investigates how perceived credibility can serve as self-feedback to enhance speaker credibility. Using a critical literacy-informed self-assessment tool, 30 business undergraduates in an advanced business communication course observed the first minute of their recorded presentation at 5-s intervals for 5 non-verbal cues. Coded findings revealed two key gender differences in positioning for speaker credibility: females used fewer body movements, more facial expressiveness and eye contact while males demonstrated greater fluency and more confident hand gestures. These findings may allow instructors and learners to mitigate gender-based differences in the classroom to enhance oral presentation competence critical for workplace success.

    doi:10.1177/23294906261479672
  4. Human-AI feedback ecologies in L2 writing: A Q-methodological study of evaluation, positioning, and regulation ↗
    doi:10.1016/j.compcom.2026.103026
  5. Navigating ethical collaboration with machine translation: An exploratory study on the role of L2 development in GenAI writing ↗
    Abstract

    We explore how language development, and training informed by development, impact L2 students’ navigation of ethical collaboration with machine translation (MT). Our interdisciplinary approach integrated the Writing Studies’ concept of ‘ethical collaboration’ – an ethically guided interface between students and GenAI – with Applied Linguistics, which highlights the need to address GenAI/MT ethics and L2 language. We created a model synthesizing the “Student Guide to AI Literacy” (MLA, 2024) with a theory of L2 development. Data were collected by 1) an established protocol – direct writing in English, self-translation from the L1 and machine translation from the L1, 2) developmentally focused and acknowledgement training, and 3) post-editing of self-translation. An emergence (onset) criterion and frequencies measured written development and evaluation of MT output, frequencies measured monitoring after training and a statistical analysis measured impact of the training types. The analysis of development and evaluation demonstrated that development impacted un/ethical evaluation of MT output. The post-edits indicated that developmentally focused training encouraged ethical monitoring when development permitted and had more impact than acknowledgement training. We conclude that guidelines and training on ethical collaboration with GenAI should be informed by L2 development, not only GenAI literacy, and that research on this topic should continue.

    doi:10.1016/j.compcom.2026.103021

August 2026

  1. The Discourse-Directive Function of Loci in the Argument Model of Topics ↗
    Abstract

    Abstract In argumentation theory, loci - also referred to as topoi - are commonly viewed as abstract sources of inference that guide the invention of arguments. Within the Argument Model of Topics (AMT), loci are analysed as sources of inference underlying individual arguments. This paper extends said view by arguing that loci perform a discourse-directive function beyond the level of individual arguments. This function is understood in an analytical sense and refers to how loci orient argumentative interaction by indicating the inferential relations along which issues become arguable in particular ways, such as procedural conditions, goals, constrained choices, conceptual boundaries, or source evaluation. Situated at the meso-level of analysis, this function exceeds the micro-level reconstruction of individual arguments while preceding the identification of recurring argumentative patterns at the macro-level. This paper shows how the reappearance of premises in extended debates, when repeatedly connected to conclusions through the same loci , recurrently instantiates the same inferential relations that orient argumentative interaction. These recurrent inferential relations do not themselves constitute argumentative patterns, but repetition makes these relations empirically traceable across argumentative contributions before the stabilization of argumentative patterns. The analytical framework is illustrated by examples from the controversy of nuclear energy production.

    doi:10.1007/s10503-026-09727-x
  2. Why So Much Writing? I Thought This Was an IT class ↗
    Abstract

    Annual assessment scores from the Association to Advance Collegiate Schools of Business (AACSB), local business leader feedback, and extensive classroom experience all indicate that students need additional practice writing, particularly in business and/or technical programs in which students often actively avoid courses featuring writing (and to a lesser extent, public speaking). Moreover, experiential learning activities tend to aid students in their eventual transition into the workplace—and they often generate increased levels of student enthusiasm and engagement. However, even as students commonly engage in group-work in business and tech programs, they typically operate within insular groups that rarely if ever interact with other groups. On the other hand, real-world workgroups depend functionally upon contributions from other groups and, in turn, produce deliverables of their own within cycles of inter- and intra-group feedback and cooperation. Therefore, this semester-long, full-class assignment has been designed to address these issues within an Information Technology course setting that, quite unexpectedly for the participants, forces students to practice all of these skills within a simulated work environment.

    doi:10.1177/23294906261472047

July 2026

  1. “It Tends to Remove Things I Originally Wanted to Emphasize”: Effects of ChatGPT Revision on Rhetorical Move-Steps in Personal Statements ↗
    Abstract

    As ChatGPT is increasingly used in second language (L2) writing practice and research, its potential to provide feedback and revision has attracted much scholarly attention. However, it remains largely unknown whether and how ChatGPT revision can influence rhetorical move-steps. This study investigates the effects of ChatGPT revision on rhetorical move-steps in English personal statements (PSs) written by L2 English undergraduate students, using a combination of corpus data and stimulated recall interviews. Based on an unstructured prompt, our analysis revealed significant reductions in the rhetorical efforts devoted to five rhetorical steps. Students’ responses highlighted both benefits and concerns regarding these revisions, illustrating how AI-generated changes can alter textual features and affect writer-reader communication from the writers’ perspective. The findings highlight the importance of students’ critical evaluation of AI-generated revisions and iterative engagement with them.

    doi:10.1177/07410883261461562
  2. The Dynamics Between the Biographical Factors, Professional Identities, and Teaching Practices of Second Language Writing Instructors ↗
    Abstract

    This qualitative study investigates the biographical factors that influence the professional identities and teaching practices of second language (L2) writing instructors. Previous L2 writing studies have tended to rely on teacher interviews to examine such identities and practices and have failed to include the important perspective of classroom observations. This study conducted teacher interviews and classroom observations with three graduate teaching assistants (GTAs) of a first-year writing course at a U.S. university. The findings suggest that although teachers’ disciplinary backgrounds, teaching philosophies, and prior experiences shape their professional identities and instructional approaches, their classroom practices in language instruction do not always align with their beliefs and stated values. Teacher education programs should integrate language pedagogy and grammar instruction into writing pedagogy, while encouraging teachers to engage in consistent self-evaluation and reflection on their practices.

    doi:10.1177/07410883261461552
  3. Bridging the Gap Between Numbers and Experience: Identifying Community Risk Thresholds in Extreme Weather Communication ↗
    Abstract

    In weather risk communication, there is a disconnect between numerical models and community members’ experiences. To help resolve this issue, the authors propose a topological method to identify community risk thresholds, arguing that they offer a middle ground that enables more robust communication about risk. The authors surveyed 1,429 participants, including decision-makers and public audience members, who responded to probabilistic visualizations from five National Weather Service offices. The results reveal that audiences engage in sophisticated risk assessment by drawing on multiple dimensions that can be communicated as community risk thresholds. Embedding these thresholds in the forecast can better link numerical and experiential risk.

    doi:10.1177/10506519261455873
  4. Assessing the reliability and validity of large language models in automatic essay scoring ↗
    Abstract

    With the advent of artificial intelligence, large language model (LLM) based Automated Essay Scoring (AES) systems have been developed that can consistently make human-like decisions that do not depend fully on surface level linguistic features. However, research into the use of LLM-based AES systems is limited and little is known about the reliability, agreement, or validity of the systems. The goal of this study was to provide evidence for the reliability, agreement, and validity of LLM-based AES systems in a standardized writing assessment used for secondary school students. Both representation and generative LLM-based AES systems were developed to score persuasive essays and assessed for reliability. Then the agreement of the developed AES systems with human raters was assessed through correlational analyses. We used extrinsic convergent validation approaches to examine if the human and LLM scores correlated with linguistic components. Results indicate strong reliability and agreement for the LLM scores. In terms of convergent validity, initial correlational analyses indicated that the representation LLM AES system showed differential correlations with the human scores in terms of a text length and type-token ratio component. This result contrasts with the correlational results from the generative LLM AES model, which indicated no differences in associations between the model and human scores with regards to the linguistic components.

    doi:10.1016/j.asw.2026.101082
  5. “Investigating the impact of ChatGPT-assisted self-assessment on college students’ writing development: Insights from diverse linguistic backgrounds” ↗
    doi:10.1016/j.asw.2026.101092
  6. From keystrokes to scores: Toward a multidimensional predictive model of writing evaluation by humans and large language models across linguistic, cognitive, and social dimensions ↗
    Abstract

    Automated writing evaluation (AWE) has traditionally emphasized textual features such as vocabulary and syntax, while often overlooking writers’ social identities and cognitive behaviors – factors central to understanding writing as a multidimensional construct. With the increasing integration of large language models (LLMs) into AWE, questions remain about how their assessments align with human judgments and the sources of potential divergences. This study investigates how linguistic (e.g., lexical diversity), cognitive (e.g., pausing behavior), and social (e.g., gender) factors covary with essay scores assigned by human raters and LLMs. We analyzed 4245 argumentative essays paired with demographic metadata and keystroke-logging data, using correlation analyses, random forest models, and regression-based approaches to examine relationships among writer characteristics, writing-process features, textual features, and essay scores. Results showed moderate agreement between human and LLM scores, but the two scoring systems exhibited different patterns of association with linguistic, cognitive, and social variables. These findings suggest that human and LLM evaluations rely on partially different cues and demonstrate how socio-cognitive metadata can be used to examine the factors associated with writing assessment decisions. By moving beyond text-only comparisons, this approach provides a complementary lens for understanding why and how human and machine judgments converge or diverge.

    doi:10.1016/j.asw.2026.101090
  7. When does GenAI feedback support learning? Trust calibration and verification in L2 writing assessment ↗
    doi:10.1016/j.asw.2026.101087
  8. Young L2 students’ use of an AI-assisted writing assessment and feedback tool: An exploratory study in multiple settings ↗
    doi:10.1016/j.asw.2026.101083
  9. Beyond the grade: Reimagining second language (L2) writing assessment through ungrading ↗
    doi:10.1016/j.asw.2026.101074
  10. GPTZero and the challenges of AI detection in assessing writing ↗
    Abstract

    GPTZero is an AI detection platform that scans written text for statistical signatures of machine generation and returns a probability score estimating whether it was produced by a human or an AI. In higher education, many teachers have turned to AI detection as a first-line response to the integrity crisis triggered by large language models. However, empirical findings on GPTZero’s efficacy are notably mixed. Some studies report strong diagnostic value under controlled conditions, while others document substantial false-negative rates, near-random performance on certain AI-generated essays, and frequent misclassification of AI-translated texts across several languages. Multilingual and L2 writers often bear the greatest cost, as their carefully constructed English is sometimes assigned high AI-likelihood scores because their linguistic profiles may appear less natural to models trained predominantly on standard or formulaic patterns of written English. In developing countries, where students commonly write in English as a second or third language, these limitations represent more than minor technical issues; they raise concerns about equity, potentially placing disproportionate burdens on writers working to meet academic language expectations. This article argues that GPTZero is unsuitable as a definitive tool for high-stakes assessment of writing. Instead, it proposes a shift toward postplagiarism frameworks that recognize responsible AI use. Within this approach, AI detection outputs serve as formative resources for developing critical AI literacy rather than surveillance tools. Flagged content becomes a starting point for metacognitive dialogue, which supports trust-based pedagogies that emphasize student agency and intellectual accountability.

    doi:10.1016/j.asw.2026.101078
  11. QuillBot in L2 writing: Implications for assessment in the age of AI ↗
    doi:10.1016/j.asw.2026.101079
  12. Assessing fairness in AI-assisted writing scoring: Developing fairness measures to detect predictive bias in automated essay scoring ↗
    Abstract

    Automated essay scoring (AES) is increasingly utilized in educational settings, yet concerns about its fairness persist. This study reviews current fairness measures in AES and summarizes their respective strengths and weaknesses. Drawing on principles from educational and psychological testing, we introduce two measures for detecting potential predictive bias: conditional disparity ratio and conditional disparity difference. Our method emphasizes two key principles: first, that bias should be assessed among students with comparable proficiency levels, and second, that evaluations should be conducted on a test set independent of the AES training set. We demonstrated this approach using writing samples from the Facial Action Coding System task within the PERSUADE 2.0 corpus to assess potential predictive bias related to sex and race. Four AES models were evaluated for predictive bias: ordinal logistic regression using TF–IDF features, fine-tuned BERT, and ChatGPT in both zero-shot and few-shot settings. The findings indicated that, without accounting for proficiency, subgroup differences remained ambiguous, making it difficult to detect potential predictive bias. In contrast, conditioning on proficiency revealed clearer and more interpretable patterns of bias. The discussion addresses key factors and the extension of the bias detection framework and outlines future directions for bias mitigation. • Introduces two fairness measures for automated writing scoring. • Distinguishes predictive bias from real proficiency differences. • Uses multiple scoring models, from machine learning to large language models. • Shows fairness varies by demographic group and proficiency level. • Offers practical guidance for bias detection in AI writing assessment.

    doi:10.1016/j.asw.2026.101066
  13. Investigating the impact of ChatGPT-assisted self-assessment on college students' writing development: Insights from diverse linguistic backgrounds ↗
    doi:10.1016/j.asw.2026.101061
  14. LAWE-CL2: Multi-agent LLM-based automated writing evaluation system integrating linguistic features with fine-tuning for Chinese L2 writing assessment ↗
    doi:10.1016/j.asw.2026.101051
  15. Developing a rating scale for written intralinguistic mediation in a local context ↗
    Abstract

    Intralinguistic mediation as a task type offers testing bodies an opportunity to expand the test construct to include authentic ability for use test tasks and is of great relevance as a representation of English lingua franca usage in the higher education context. This study reports on the development of an analytic rating scale for one particular mediation task specification, that of facilitating communication in delicate situations and disagreements , which was co-constructed by raters based on theoretical considerations, example scripts, and CEFR/CV descriptors. Raters used the scale to mark 80 performances across three tasks, and results were analyzed using a many-facet Rasch hybrid partial credit model, which pays particular attention to the functioning of an analytic scale and its categories. Findings show that despite the complexities of this multidimensional construct, it can be operationalized through well-designed tasks and a strategic scale-development process. Following focus group feedback and further reflection on the quantitative results, minor refinements were made to the scale. Findings indicate potential for future scale development for scoring authentic, context-specific task types, and the study has clear implications for other testing bodies hoping to include mediation in their proficiency exams. • European university test re-development project in a lingua franca context. • Developed a rating scale for an intra-linguistic written mediation task specification. • Scale co-constructed by raters using theory, example scripts, and CV descriptors. • Iterative mixed-method approach informed the rating scale development. • New scale shown to be valid, contribution to future mediation assessment.

    doi:10.1016/j.asw.2026.101049

June 2026

  1. Résumé Redesign: An Experiential Assignment Using AI Applicant Tracking Systems ↗
    Abstract

    The Résumé Redesign exercise equips business communication students with essential job search competencies through an experimental approach to résumé design. Grounded in contemporary research on AI literacy and résumé design, this assignment incorporates an AI-powered résumé scanner alongside faculty feedback to strengthen students’ understanding of formatting, mechanics, and audience adaption. Survey results from 136 student respondents indicate the AI feedback to be clear and helpful. This exercise offers flexibility for in-person and online learning environments to enhance a résumé assignment and introduce students to AI-mediated recruitment practices, while maintaining the critical role of human review in the résumé evaluation.

    doi:10.1177/23294906261454315
  2. Hype Circulation: Shaping Public Engagement with Emerging Technologies Through Anticipatory Communities and Embodiment ↗
    Abstract

    This article investigates how the public circulation of hype functions rhetorically in the context of emerging immersive technologies. Through an analysis of YouTube reviews of Apple’s Vision Pro headset, we examine how rhetorical strategies operate at the intersection of wearable and embodied technology discourse, affect theory, and public circulation. The analysis focuses on four dominant rhetorical strategies that function to intensify anticipation while simultaneously constraining critical evaluation: excitement, exploration, exemplification, and exaggeration. These findings are supplemented with extended-use experiences from multiple sources, illustrating the tension between anticipatory rhetoric and lived experience. By theorizing the embodied and affective dimensions of VR discourse, this article contributes to our understanding of how rhetorical practices in digital environments reshape critical evaluation about technologies.

    doi:10.1080/02773945.2026.2684344
  3. Teaching grammar and writing: A randomised controlled trial and implementation process evaluation of Englicious ↗
    Abstract

    Very few research studies of the teaching of grammar and writing had been carried out with children younger than eight-years-old prior to the research reported in this paper. The research evaluated a new approach to teaching grammar and writing called Englicious. A Randomized Controlled Trial (RCT) and Implementation and Process Evaluation (IPE) research design, featuring 1,246 pupils aged six to seven-years-old in 70 primary school classes, was used to evaluate the effectiveness of Englicious for improving children’s writing. The approach was implemented in the context of the programs of study for grammar teaching in England’s national curriculum. The research found that there was no effect of the grammar teaching intervention on pupils’ narrative writing. The effect size for pupils’ sentence generation was sufficient to merit reflections about potential impacts of aspects of the intervention although this outcome also did not reach statistical significance. It is hypothesised that the manipulation of words, phrases and sentences, combined with practice at writing, may have contributed to any positive effect, although this would need to be confirmed in future research. It is concluded that until more research is done to investigate the effectiveness of different approaches to teaching grammar and writing with young children, existing evidence-based approaches are more likely to be effective to help young pupils’ narrative writing. England’s national curriculum specifications for teaching grammar and writing could usefully be reviewed to more closely reflect the evidence base from this field of writing research.

    doi:10.17239/jowr-2026.17.04.04
  4. Teaching Writing in Secondary School English Language Classrooms in Côte d'Ivoire: An Exploratory Study ↗
    Abstract

    English language teachers in many countries around the world teach large classes of 40 or more students. Since the 18th century, many English teaching methods have focused on speaking skills. However, in recent years, globalization has heightened the importance of writing instruction. Writing skills have become increasingly important for English language learners in professional, academic, and personal contexts. Despite this increasing importance, to date, there has been limited research that shows how writing instruction is implemented in large, secondary school English language classes although large-class contexts represent the majority of English language classes globally. Therefore, this qualitative, exploratory study aimed to explore and understand writing instruction in one particular large-class context—public secondary school English language classes in Côte d'Ivoire. Data were collected through autobiographical essays, interviews, and teaching artifacts and analyzed through narrative profiles. Major findings show that Ivorian secondary school English language teachers might not have much training for teaching writing. Additionally, Ivorian secondary school English language teachers face numerous challenges teaching writing, including student reluctance to write. Implications suggest the need to give more consideration to writing assessments and instruction and the need to ensure that secondary school English language teachers receive adequate training for teaching writing that fits their large-class contexts.

    doi:10.3138/wap-2025-0009
  5. The Impact of Statutory Assessments on Writing Instruction in English Primary Schools: An Exploration of Teacher Perceptions and Practices ↗
    Abstract

    This article outlines the curriculum and assessment regime related to writing in England that affects children in their final year of primary school (Year 6: ages 10–11). The assessment in England is colloquially known as SATs and, for writing, consists of a grammar, punctuation, and spelling test, alongside a separate teacher assessment of writing using a set of criteria. Both these assessments are statutory and are integral components of the accountability system for primary schools in England. From an interpretive perspective, this article explores how teachers perceive the assessment and how it influences their instructional practices through data collected from interviews with 10 Year 6 teachers. A broad discussion centres around the potential negative consequences of the assessment regime on the teaching of writing and includes the teachers’ perceived impact of these assessments on writing instruction and the effects this has on the teachers’ teaching practices. There is a particular focus on issues around “teaching-to-the-test” with some implications for policy.

    doi:10.3138/wap-2024-0006

May 2026

  1. Do You Want to Build a Straw Man?: Evidentiary and Argumentative Modes of the Straw Man Fallacy ↗
    Abstract

    Abstract I propose a distinction between evidentiary and argumentative modes of the straw man fallacy. Traditional studies of this fallacy have focused on the changes that occur when a discussant represents another’s speech acts. This places undue normative significance on the question of “How much change is too much change?”, a question that has constantly eluded theorizing. I argue that in a dualist framework, when the evidentiary mode of the fallacy is taken as a starting point, the evaluation can begin with the phenomenon of deceit (as a marker of fallaciousness) and construct critical responses without the need to demonstrate that the change induced was in some sense disproportionate. To make this point, I give both imaginary and real-life examples of such evaluations.

    doi:10.1007/s10503-026-09710-6
  2. Developing an Integrated Logging Method for Process-Based Assessment and Feedback in Technical Writing ↗
    Abstract

    Traditional assessments of technical writing privilege final products or task outcomes and provide limited warrant for predicting writing competence across task types and workplace contexts. This article proposes a methodological shift toward process-based assessment, which evaluates the effectiveness of writers’ strategic actions during text production using observable process indicators and translates that evaluation into individualized, strategy-focused feedback benchmarked against group norms. To address alphabetic bias and genre/context blindness in keystroke logging, the article develops an integrated logging method for analyzing technical writing processes. To demonstrate feasibility and analytic affordances, it presents a case study adapting Perrin’s progression analysis for professional writing: Processes from 24 technical writers were captured and analyzed to generate performance feedback, strategy instruction, and process-based competence measures. By specifying analytic procedures, key process indicators, and principles for inferring writing competence from writing processes, this article advances process-oriented professional writing research and offers a transferable, scalable framework for workplace evaluation.

    doi:10.1177/07410883261440254
  3. Beyond Co-Regulation: Interplay as a Methodological Framework for Examining Self-Regulation in Generative AI-Assisted Writing ↗
    Abstract

    As generative artificial intelligence (GenAI) tools become embedded in writing practices, researchers must refine methodologies for studying self-regulation in AI-assisted composition. While sociocognitive and co-regulation frameworks have effectively captured self-regulatory processes in human collaboration, they are insufficient for understanding how writers manage the dynamic and probabilistic nature of AI-generated text. This article introduces interplay as a methodological framework to analyze the recursive process of initiating, responding, adapting, and revising in human–AI writing interactions. Unlike co-regulation, where collaborators share communicative intent, interplay highlights the writer’s active role in interpreting and steering AI-generated content. Drawing on self-regulation theory, we propose an analytical framework that integrates traditional self-regulation categories (goal-setting, monitoring, and reflection) with interplay-specific coding (initiation, evaluation, acceptance, and adaptation). Through case analyses of human–AI writing exchanges, we demonstrate how interplay provides a systematic approach to studying agency, decision making, and regulatory strategies in AI-assisted writing. We argue that recognizing interplay as a distinct dimension of self-regulation advances both empirical research and pedagogical approaches to AI-mediated composition.

    doi:10.1177/07410883261440232

April 2026

  1. Examining Automated Writing Evaluation Error Coverage in Relation to Uptake and Retention ↗
    Abstract

    Despite the current widespread use of Automated Writing Evaluation (AWE) feedback, many issues regarding its efficacy still remain unresolved. Recent studies mainly focus on correctly detected errors with a lack of attention on the comprehensiveness of error detection, or error coverage. Error coverage is interesting because little is known about the capacity of AWE systems to fully detect common second language (L2) errors. It is also important to investigate the potential effect of such capacity on student uptake and retention, which are important constructs in fostering L2 writing development. To this end, the present study compared teacher feedback and AWE error coverage in L2 writing classes. The findings suggest that both the AWE system and the teacher demonstrated low error coverage across grammar, usage, and mechanics error categories. However, they indicated differences in the types of errors they identified most frequently. The AWE system flagged more mechanical errors, whereas the teacher provided twice as many corrections for grammar errors, including wrong/missing words, prepositions, and incorrect word forms. While the AWE system performed moderately in flagging articles and comma errors, it struggled with more nuanced grammatical errors, suggesting it may not be a reliable standalone tool for addressing specific needs of L2 learners’ writing challenges. Interestingly, coverage was positively associated with successful uptake, with students utilizing a wider variety of revision acts (i.e., change, add, delete, remove) on AWE errors identified compared to errors not identified. However, error coverage did not correlate with short- or long-term retention of accuracy, implying that retention may result from the interplay of error coverage with other factors. Findings provide implications for writing teachers regarding the employment of AWE systems and for AWE developers regarding the future optimizations of the AWE systems.

  2. Generative artificial intelligence for automated writing evaluation: A systematic review of trends, efficacy, and challenges ↗
    doi:10.1016/j.asw.2026.101041
  3. Pursuing fair writing assessment: Halo effects in primary school foreign language writing in grade six ↗
    Abstract

    Assessing the writing competence of pupils learning English as a foreign language (EFL) at primary school is associated with specific challenges because of learners’ limited language resources. This study investigates the extent to which characteristics of their texts trigger so-called halo effects. Halo effects are an assessment bias where the quality of one feature unintentionally influences the evaluation of other aspects. The study examines halo effects across nine aspects of text quality (communicative effect, level of detail, coherence, cohesion, complexity of syntax and grammar, correctness of syntax and grammar, vocabulary, orthography and punctuation), based on a random sample of narrative texts from a sixth-grade corpus. 200 pre-service teachers assessed four randomly assigned texts. Halo effects were calculated by comparison to expert ratings using multi-level regression analyses. Results show that orthography and vocabulary were the two main triggers of halo effects. Punctuation also triggered some halo effects, but to a smaller extent. The assessment of communicative effect, complexity and correctness of syntax and grammar was not determined by the corresponding text quality but dominated by other criteria. Results highlight the importance of being aware of halo effects when assessing young EFL learners’ texts and emphasise the need for suitable training measures. • Analysis of halo effects across nine aspects of text quality. • Random sample of narrative texts from a sixth-grade EFL corpus. • Orthography and vocabulary are the two main triggers of halo effects. • Punctuation also triggers halo effects but to a smaller extent. • Halo effects call for awareness and targeted training.

    doi:10.1016/j.asw.2026.101036
  4. From spelling to content: The influence of spelling quality on text assessment ↗
    doi:10.1016/j.asw.2026.101014
  5. How do L2 writing subskills interact hierarchically? Insights from diagnostic classification models ↗
    Abstract

    This study examined the hierarchical structure among second/foreign language (L2) writing subskills using a Hierarchical Diagnostic Classification Model (HDCM). A pool of 500 essays composed by English as a Foreign Language (EFL) students was assessed by four experienced EFL teachers using the Empirically-derived Descriptor-based Diagnostic (EDD) checklist. Based on a literature review and the expertise of three content experts, several models were developed to reflect various hierarchical interactions among L2 writing subskills, including linear, divergent, convergent, independent, unstructured, mixed, and higher-order. The comparison of the models showed the presence of an unstructured interaction among L2 writing subskills, indicating that content is the foundational subskill for the mastery of vocabulary, grammar, organization, and mechanics. Higher mastery classes were also associated with higher educational levels, greater frequency of English use, and longer exposure to L2. Understanding the hierarchical relationships among L2 writing subskills can improve targeted instructional strategies and assessment practices. • A constrained version of existing DCMs is represented by hierarchical DCMs. • Models were developed to show hierarchical interactions among L2 writing subskills. • An unstructured interaction among L2 writing subskills was identified. • Higher mastery classes were associated with higher educational levels. • The classes were associated with greater English use and longer L2 exposure.

    doi:10.1016/j.asw.2026.101029
  6. Assessing GenAI-assisted digital multimodal composing: Reconceptualizing a genre-based framework through self-assessment and peer assessment ↗
    doi:10.1016/j.asw.2026.101017
  7. Assessing fairness in finetuned scoring models with demographically restricted training data ↗
    Abstract

    The increasing adoption of automated essay scoring (AES) in high-stakes educational contexts necessitates careful examination of potential biases within the systems. This study investigates how the demographic composition of training data influences fairness in AES systems developed from finetuned large language models (LLMs). Using the PERSUADE corpus of 26,000 student essays, we conducted a systematic analysis using demographically restricted training sets to isolate the impact of training data demographics on LLM-AES performance. Each demographically restricted training set comprised essays written by one racial/ethnic group. Four variants of a Longformer-based AES were developed: one trained on demographically balanced data and three trained on demographically restricted datasets. An initial analysis of the human ratings indicated that demographic factors significantly predict human essay scores (marginal R² = 0.125), a pattern that is paralleled in national writing assessment data. LLM-AES systems trained on demographically restricted data exhibited small systematic biases (marginal R² = 0.043). However, the LLM trained on balanced data showed minimal demographic bias, suggesting that representative training data can effectively prevent amplification of demographic disparities beyond those present in human ratings. These results highlight both the importance and limitations of training data diversity in achieving fair assessment outcomes. • 12.5% of variance in human essay ratings was explained by demographics. • We construct demographically restricted training sets to isolate bias. • Balanced training data minimized LLM-AES bias across demographic groups. • LLM-AES trained on demographically restricted data showed more bias.

    doi:10.1016/j.asw.2026.101032
  8. Aligning ACTFL writing proficiency guidelines with CEFR descriptors: Insights from Chinese writing assessment ↗
    doi:10.1016/j.asw.2026.101033