Discovering the Predictive Power of Five Baseline Writing Competences

Vivekanandan Kumar ; Shawn N. Fraser Athabasca University ; David Boulanger Athabasca University

Abstract

Background: A shift of focus has been marked in recent years in the development of automated essay scoring systems (AES) passing from merely assigning a holistic score to an essay to providing constructive feedback over it. Despite all the major advances in the domain, many objections persist concerning their credibility and readiness to replace human scoring in high-stakes writing assessments. The purpose of this study is to shed light on how to build a relatively simple AES system based on five baseline writing features. The study shows that the proposed AES system compares very well with other state-of-the-art systems despite its obvious limitations. Literature Review: In 2012, ASAP (Automated Student Assessment Prize) launched a demonstration to benchmark the performance of state-of-the-art AES systems using eight hand-graded essay datasets originating from state writing assessments. These datasets are still used today to measure the accuracy of new AES systems. Recently, Zupanc and Bosnic (2017) developed and evaluated another state-of-the-art AES system, called SAGE, which enclosed new semantic and consistency features and provided for the first time an automatic semantic feedback. SAGE’s agreement level between machine and human scores for ASAP dataset #8 (the dataset also of interest in this study) was measured and had a quadratic weighted kappa of 0.81, while it ranged for 10 other state-of-the-art systems between 0.60 and 0.73 (Chen et al., 2012; Shermis, 2014). Finally, this section discusses the limitations of AES, which come mainly from its omission to assess higher-order thinking skills that all writing constructs are ultimately designed to assess. Research Questions: The research questions that guide this study are as follows: RQ1: What is the power of the writing analytics tool’s five-variable model (spelling accuracy, grammatical accuracy, semantic similarity, connectivity, lexical diversity) to predict the holistic scores of Grade 10 narrative essays (ASAP dataset #8)? RQ2: What is the agreement level between the computer rater based on the regression model obtained in RQ1 and the human raters who scored the 723 narrative essays written by Grade 10 students (ASAP dataset #8)? Methodology: ASAP dataset #8 was used to train the predictive model of the writing analytics tool introduced in this study. Each essay was graded by two teachers. In case of disagreement between the two raters, the scoring was resolved by a third rater. Basically, essay scores were the weighted sums of four rubric scores. A multiple linear regression analysis was conducted to determine the extent to which a five-variable model (selected from a set of 86 writing features) was effective to predict essay scores. Results: The regression model in this study accounted for 57% of the essay score variability. The correlation (Pearson), the percentage of perfect matches, the percentage of adjacent matches (±2), and the quadratic weighted kappa between the resolved scores and predicted essay scores were 0.76, 10%, 49%, and 0.73, respectively. The results were measured on an integer scale of resolved essay scores between 10-60. Discussion: When measuring the accuracy of an AES system, it is important to take into account several metrics to better understand how predicted essay scores are distributed along the distribution of human scores. Using average ranking over correlation, exact/adjacent agreement, quadratic weighted kappa, and distributional characteristics such as standard deviation and mean, this study’s regression model ranks 4th out of 10 AES systems. Despite its relatively good rank, the predictions of the proposed AES system remain imprecise and do not even look optimal to identify poor-quality essays (binary condition) smaller than or equal to a 65% threshold (71% precision and 92% recall). Conclusions: This study sheds light on the implementation process and the evaluation of a new simple AES system comparable to the state of the art and reveals that the generally obscure state-of-the-art AES system is most likely concerned only with shallow assessment of text production features. Consequently, the authors advocate greater transparency in the development and publication of AES systems. In addition, the relationship between the explanation of essay score variability and the inter-rater agreement level should be further investigated to better represent the changes in terms of level of agreement when a new variable is added to a regression model. This study should also be replicated at a larger scale in several different writing settings for more robust results.

Journal
Journal of Writing Analytics
Published
2017-01-01
DOI
10.37514/jwa-j.2017.1.1.08
CompPile
Open Access
OA PDF Gold
Topics
Export

Citation context

Cited by in this index (3)

  1. Journal of Writing Analytics
  2. Journal of Writing Analytics
  3. Journal of Writing Analytics

References (36) · 7 in this index

  1. University of Toronto, Department of Computer Science (2017). Computational linguistics. Retrieved June 27, 2…
  2. Aluthman, E. S. (2016). The effect of using automated essay evaluation on ESL undergraduate students' writing…
     ↗
  3. Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater v.2. Journal of Technology, Learning,…
  4. Bennett, R. E. (2011). CBAL: Results from piloting innovative K-12 assessments. ETS Research Report Series, 2…
     ↗
  5. Brenner, H., & Kliebsch, U. (1996). Dependence of weighted kappa coefficients on the number of categories. Ep…
     ↗
Show all 36 →
  1. Chen, H., He, B., Luo, T., & Li, B. (2012). A ranked-based learning approach to automated essay scoring. In 2…
     ↗
  2. Clemens, C. (2017). A causal model of writing competence (Master's thesis). Retrieved from https://dt.athabas…
  3. Crossley, S. A., Kyle, K., & McNamara, D. S. (2016a). The development and use of cohesive devices in L2 writi…
     ↗
  4. Crossley, S. A., Kyle, K., & McNamara, D. S. (2016b). The tool for the automatic analysis of text cohesion (T…
     ↗
  5. Crossley, S. A., & McNamara, D. S. (2011). Text coherence and judgments of Vivekanandan Kumar, Shawn N. Frase…
  6. Assessing Writing
  7. El Ebyary, K., & Windeatt, S. (2010). The impact of computer-based feedback on students' written work. Intern…
     ↗
  8. Elliot, N., Rudniy, A., Deess, P., Klobucar, A., Collins, R., & Sava, S. (2016). ePortfolios: Foundational me…
  9. Fazal, A., Hussain, F. K., & Dillon, T. S. (2013). An innovative approach for automatically grading spelling …
     ↗
  10. Foltz, P. W. (1996). Latent semantic analysis for text-based research. Behavior Research Methods, 28(2), 197-…
     ↗
  11. Gebril, A., & Plakans, L. (2016). Source-based tasks in academic writing assessment: Lexical diversity, textu…
     ↗
  12. Gregori-Signes, C., & Clavel-Arroitia, B. (2015). Analysing lexical density and lexical diversity in universi…
     ↗
  13. Huck, S. (2009). Statistics misconceptions. New York, NY: Taylor & Francis.
  14. Journal of Writing Analytics
  15. Larkey, L. S. (1998). Automatic essay grading using text categorization techniques. In Proceedings of the 21s…
     ↗
  16. Latifi, S., Gierl, M. J., Boulais, A.-P., & De Champlain, A. F. (2016). Using automated scoring to evaluate w…
     ↗
  17. Manning, C. D., Raghavan, P., Schütze, H., & others. (2008). Introduction to information retrieval (Vol. 1). …
     ↗
  18. McNamara, D. S., Crossley, S. A., & Roscoe, R. (2013). Natural language processing in an intelligent writing …
     ↗
  19. Assessing Writing
  20. Miłkowski, M. (2010). Developing an open-source, rule-based proofreading tool. Software: Practice and Experie…
     ↗
  21. Naber, D. (2003). A rule-based style and grammar checker. Retrieved from https://www.researchgate.net/publica…
  22. Perelman, L. (2013). Critique of Mark D. Shermis & Ben Hammer, Contrasting state-of-the-art automated scoring…
  23. Assessing Writing
  24. Rudner, L. M., Garcia, V., & Welch, C. (2006). An evaluation of IntelliMetric TM essay scoring system. The Jo…
  25. Assessing Writing
  26. Assessing Writing
  27. Villalon, J., & Calvo, R. A. (2013). A decoupled architecture for scalability in text mining applications. Jo…
  28. Warschauer, M., & Ware, P. (2006). Automated writing evaluation: Defining the classroom research agenda. Lang…
     ↗
  29. Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. E…
     ↗
  30. Assessing Writing
  31. Zupanc, K., & Bosnić, Z. (2017). Automated essay evaluation with semantic analysis. Knowledge-Based Systems, …
     ↗