Assessing Writing

22 articles
Year – Topic Clear
Export Subscribe (Atom)
qualitative research×

October 2026

  1. Generative artificial intelligence in writing assessment: University professors’ perspectives and recommendations ↗
    Abstract

    Research has recognized the potential of Generative AI (GenAI) to support writing assessment. However, the responsible and effective application of GenAI in writing assessment remains underexplored, and consequently, university instructors are uncertain about how to integrate these tools into their assessment practices. Drawing on Mills et al.’s (2024) three-dimensional GenAI literacy framework, this qualitative study examined professors’ beliefs about incorporating GenAI into L2 writing assessment and their recommendations for responsible and effective integration. Twenty-four professors from universities in Chinese Mainland, Hong Kong, Macau, and Singapore were recruited. Data were collected through semi-structured interviews and stimulated recall tasks. The findings revealed that while professors recognized benefits of incorporating GenAI in writing assessment, they expressed concerns about certain critical issues. They emphasized the essential role of human agency and offered practical recommendations for effective, responsible, and supportive use. Based on these findings, the study proposes seven evidence-based and actionable recommendations, including enhancing instructors’ GenAI literacy, adopting multiple sources of feedback, prioritizing teacher judgment and modification for reliability, validity, and discriminative power, ensuring transparent and ethical use of GenAI, foster socio-emotional awareness, as well as establishing and implementing institutional policies and training for GenAI adoption in writing assessment. These recommendations help guide the development of institutional policies and enable university instructors of L2 writing and other disciplines to integrate GenAI into assessment practices with responsibility, ethics, and effectiveness.

    doi:10.1016/j.asw.2026.101125
  2. Collaborative agency in rating scale development for L2 academic writing assessment ↗
    Abstract

    Calls for greater learner involvement in assessment have highlighted a need for rating scales that are not only technically sound but also pedagogically meaningful. However, in second language (L2) writing, students are largely absent from scale design, with most instruments being developed by teachers or test developers. Previous co-construction studies have often positioned students as the primary agents, with comparative underuse of teacher expertise and limited curricular alignment. This study addresses this gap by investigating how collaborative agency between teachers and students can inform the development of a rating scale for academic writing in a Saudi higher education context. Drawing on a multi-source framework, four meshed teacher–student focus groups and a follow-up focus group were conducted to negotiate the scale criteria, properties, and level descriptors. Qualitative content analysis revealed both convergence and divergence: while grammar, organization, and development were prioritized, students argued for a stronger recognition of ideas and discourse-level features, whereas teachers emphasized accuracy and feasibility. Through cycles of debate and compromise, the process yielded a five-level analytic rating scale with six criteria grounded in curricular objectives and stakeholder perspectives. This study shows how collaborative agency was enacted through a dialogic process of negotiating and integrating teacher and student perspectives in rating scale development. It extends participatory assessment research in L2 writing by illustrating a context-sensitive, theoretically informed, and curriculum-aligned approach to analytic rating scale co-construction for classroom use.

    doi:10.1016/j.asw.2026.101126

July 2026

  1. Epistemic cognition and feedback engagement: A case study of Chinese L2 writers ↗
    doi:10.1016/j.asw.2026.101071
  2. Developing a rating scale for written intralinguistic mediation in a local context ↗
    Abstract

    Intralinguistic mediation as a task type offers testing bodies an opportunity to expand the test construct to include authentic ability for use test tasks and is of great relevance as a representation of English lingua franca usage in the higher education context. This study reports on the development of an analytic rating scale for one particular mediation task specification, that of facilitating communication in delicate situations and disagreements , which was co-constructed by raters based on theoretical considerations, example scripts, and CEFR/CV descriptors. Raters used the scale to mark 80 performances across three tasks, and results were analyzed using a many-facet Rasch hybrid partial credit model, which pays particular attention to the functioning of an analytic scale and its categories. Findings show that despite the complexities of this multidimensional construct, it can be operationalized through well-designed tasks and a strategic scale-development process. Following focus group feedback and further reflection on the quantitative results, minor refinements were made to the scale. Findings indicate potential for future scale development for scoring authentic, context-specific task types, and the study has clear implications for other testing bodies hoping to include mediation in their proficiency exams. • European university test re-development project in a lingua franca context. • Developed a rating scale for an intra-linguistic written mediation task specification. • Scale co-constructed by raters using theory, example scripts, and CV descriptors. • Iterative mixed-method approach informed the rating scale development. • New scale shown to be valid, contribution to future mediation assessment.

    doi:10.1016/j.asw.2026.101049

January 2026

  1. Assessing the effects of task complexity on cognitive demands in L2 writing ↗
    Abstract

    The assessment of task-generated cognitive demands has been receiving increasing attention in task complexity research. However, scant attention has been paid to assessing cognitive demands when task complexity is manipulated along both resource-directing and resource-dispersing dimensions. To address this gap, the present study aimed to investigate the relative effects of reasoning demands and prior knowledge on cognitive demands in L2 writing. Eighty-eight EFL students completed two letter-writing tasks with varying reasoning demands under one of two conditions, that is, either with prior knowledge available or without prior knowledge available. Cognitive demands were assessed by the post-task questionnaire, the dual-task method and the open-ended questions. The results revealed that reasoning demands and prior knowledge were strong determinants of cognitive demands, which provided empirical evidence for Robinson’s Cognition Hypothesis. Moreover, the post-task questionnaire, the dual-task method and open-ended questions were found to assess distinct aspects of cognitive demands, which highlighted the importance of data triangulation in exploring task complexity effects. The study provides language teachers and assessors with implications for task design and implementation. • How reasoning demands and prior knowledge affect cognitive demands was underexplored. • Cognitive demands were assessed by both quantitative and qualitative methods. • Findings supported some assumptions underlying Robinson’s framework. • The independent measures assessed distinct aspects of cognitive demands.

    doi:10.1016/j.asw.2025.100998

October 2024

  1. Validating an integrated reading-into-writing scale with trained university students ↗
    Abstract

    Integrated tasks are often used in higher education (HE) for diagnostic purposes, with increasing popularity in lingua franca contexts, such as German HE, where English-medium courses are gaining ground. In this context, we report the validation of a new rating scale for assessing reading-into-writing tasks. To examine scoring validity, we employed Weir’s (2005) socio-cognitive framework in an explanatory mixed-methods design. We collected 679 integrated performances in four summary and opinion tasks, which were rated by six trained student raters. They are to become writing tutors for first-year students. We utilized a many-facet Rasch model to investigate rater severity, reliability, consistency, and scale functioning. Using thematic analysis, we analyzed think-aloud protocols, retrospective and focus group interviews with the raters. Findings showed that the rating scale overall functions as intended and is perceived by the raters as valid operationalization of the integrated construct. FACETS analyses revealed reasonable reliabilities, yet exposed local issues with certain criteria and band levels. This is corroborated by the challenges reported by the raters, which they mainly attributed to the complexities inherent in such an assessment. Applying Weir’s (2005) framework in a mixed-methods approach facilitated the interpretation of the quantitative findings and yielded insights into potential validity threads. • FACET analyses show reasonable reliabilities and scale functioning. • Mixed-methods approach facilitates interpreting the quantitative findings. • Raters perceive rating scale as valid operationalization of integrated construct. • Applying Weir’s socio-cognitive framework reveals potential validity threads. • Raters attribute challenges to the complexities inherent in integrated writing.

    doi:10.1016/j.asw.2024.100894

January 2024

  1. Reading, receiving, revising: A case study on the relationship between peer review and revision in writing-to-learn ↗
    doi:10.1016/j.asw.2024.100808

October 2023

  1. Understanding EFL students’ feedback literacy development in academic writing: A longitudinal case study ↗
    doi:10.1016/j.asw.2023.100770

July 2023

  1. Peer-feedback of an occluded genre in the Spanish language classroom: A case study ↗
    Abstract

    Learning how to write occluded genres is an elusive task (Swales, 1996) – even more so in the case of students writing in a second or additional language. To achieve discourse competence in the use of one of these genres, in this case the ‘statement of purpose’ typical of post-graduate programme admission forms, it is first necessary to fully understand its features at both the macrotextual and microlinguistic levels (Gillaerts, 2003; Bhatia, 2004). This qualitative study focuses on the writing of learners of Spanish as an additional language to analyse whether feedback provided by peers impacts the quality of the statements of purpose they write. Through a dual discourse analysis of their written work and in-class interactions during peer- feedback sessions, our study finds that, when properly trained and using tailored assessment tools, students can use peer-assessment profitably to improve the quality of their statements of purpose, as well as to acquire appropriate metalanguage to guide others. Our results thus reconfirm the beneficial effects of helping students to achieve feedback literacy.

    doi:10.1016/j.asw.2023.100756
  2. Developing EFL teachers’ feedback literacy for research and publication purposes through intra- and inter-disciplinary collaborations: A multiple-case study ↗
    doi:10.1016/j.asw.2023.100751

April 2023

  1. Genre pedagogy: A writing pedagogy to help L2 writing instructors enact their classroom writing assessment literacy and feedback literacy ↗
    Abstract

    As part of a larger case study, this single exploratory case study aims to explore the potential of genre-based pedagogy (GBP) to allow L2 writing instructors to enact their writing assessment literacy and feedback literacy. The findings demonstrate that GBP afforded the participating writing instructor of a genre-based EAP writing course to carry out effective writing classroom assessment practices and thus enact their2 writing assessment literacy and feedback literacy. GBP allowed effective writing classroom assessment practices such as diagnostic assessment and learner involvement in assessment. More specifically, genre exploration tasks led to diagnostic assessment and helped the instructor coordinate effective classroom discussions to elicit evidence of the students’ knowledge of the target genre that they would study. Second, students’ production of texts in target genres not only allowed the instructor to collect evidence of the students’ specific genre knowledge, but it also afforded learner involvement through self-reflection. The instructor could also efficiently interpret this evidence and provide formative feedback through pre-established genre specific assessment criteria.

    doi:10.1016/j.asw.2023.100717

October 2022

  1. Assessing pragmatic performance in advanced L2 academic writing through the lens of local grammars: A case study of ‘exemplification’ ↗
    doi:10.1016/j.asw.2022.100668
  2. Validity evidences for scoring procedures of a writing assessment task. A case study on consistency, reliability, unidimensionality and prediction accuracy ↗
    Abstract

    Scoring is a fundamental step in the assessment of writing performance. The choice of the scoring procedure as well as the adoption of a discrepancy resolution method can impact the psychometric properties of the scores and therefore the final pass/fail decision. In a comprehensive framework which considers scoring as part of the validation process of the scores, the aim of this paper is to evaluate the impact of rater mean, parity and tertium quid procedures on score properties. Using data from a writing assessment task applied in a professional context, the paper analyses score reliability, dependability, unidimensionality and decision accuracy on two sets of data; complete data and subsample of discrepant data. The results show better performance of the tertium quid procedure in terms of reliability indicators but a lower quality in defining construct unidimensionality.

    doi:10.1016/j.asw.2022.100669

April 2022

  1. The mediating effects of student beliefs on engagement with written feedback in preparation for high-stakes English writing assessment ↗
    Abstract

    Research in L2 writing contexts has shown developing writers’ beliefs exert a powerful mediating effect on how they respond to written feedback. The mediating role of beliefs is magnified in preparation for high-stakes English writing assessment contexts, where tangible outcomes pivot on successful test performance. The present qualitative case study utilises data from semi-structured interviews to investigate how the beliefs of three self-directed IELTS preparation candidates mediated their affective, behavioural, and cognitive engagement with electronic teacher written feedback across three multi-draft Task 2 rehearsal essays. Utilising a metacognitive conceptual approach (Wenden, 1998), the study identified seven themes: 1) self-concept beliefs regulated engagement, 2) reliance on the expertise of a quality teacher, 3) engagement was mediated by individuals’ learning-to-write beliefs, 4) belief in comprehensive, critical written feedback, 5) feedback deemed transferable was more comprehensively engaged with, 6) entrenched test-taking strategy beliefs hindered engagement, and 7) supplementary self-directed learning activities were considered of limited value. The implications for practitioners of IELTS Writing preparation and the IELTS co-owners are discussed.

    doi:10.1016/j.asw.2022.100611

April 2020

  1. Student engagement with automated written corrective feedback (AWCF) provided by Grammarly: A multiple case study ↗
    doi:10.1016/j.asw.2020.100450

July 2018

  1. Student engagement with teacher written corrective feedback in EFL writing: A case study of Chinese lower-proficiency students ↗
    doi:10.1016/j.asw.2018.03.001

July 2017

  1. Understanding university students’ peer feedback practices in EFL writing: Insights from a case study ↗
    doi:10.1016/j.asw.2017.03.004

January 2017

  1. How students' ability levels influence the relevance and accuracy of their feedback to peers: A case study ↗
    doi:10.1016/j.asw.2016.07.002

October 2016

  1. Rubrics and corrective feedback in ESL writing: A longitudinal case study of an L2 writer ↗
    doi:10.1016/j.asw.2016.06.003

October 2015

  1. Developing rubrics to assess the reading-into-writing skills: A case study ↗
    doi:10.1016/j.asw.2015.07.004

April 2011

  1. Academic tutors’ beliefs about and practices of giving feedback on students’ written assignments: A New Zealand case study ↗
    doi:10.1016/j.asw.2011.02.004

February 2000

  1. The student, the text, and the classroom context: A case study of teacher response ↗
    doi:10.1016/s1075-2935(00)00017-9