Developing an e-rater Advisory to Detect Babel-generated Essays

Aoife Cahill Educational Testing Service ; Martin Chodorow The Graduate Center, CUNY ; Michael Flor Educational Testing Service

Abstract

Background: It is important for developers of automated scoring systems to ensure that their systems are as fair and valid as possible. This commitment means evaluating the performance of these systems in light of construct-irrelevant response strategies. The enhancement of systems to detect and deal with these kinds of strategies is often an iterative process, whereby as new strategies come to light they need to be evaluated and effective mechanisms built into the automated scoring systems to handle them. In this paper, we focus on the Babel system, which automatically generates semantically incohesive essays. We expect that these essays may unfairly receive high scores from automated scoring engines despite essentially being nonsense. Literature Review: We discuss literature related to gaming of automated scoring systems. One reason that Babel essays are so easy to identify as nonsense by human readers is that they lack any semantic cohesion. Therefore, we also discuss some literature related to cohesion and detecting semantic cohesion. Research Questions: This study addressed three research questions:Can we automatically detect essays generated by the Babel system?Can we integrate the detection of Babel-generated essays into an operational automated essay scoring system while making sure not to flag valid student responses?Does a general approach for detecting semantically incohesive essays also detect Babel-generated essays?Research Methodology: This article describes the creation of two corpora necessary to address the research questions: (1) a corpus of Babel-generated essays and (2) a corresponding corpus of good-faith essays. We built a classifier to distinguish Babel-generated essays from good-faith essays and investigated whether the classifier can be integrated into an automated scoring engine without adverse effects. We also developed a measure of lexical-semantic cohesion and examined its distribution in Babel and in good-faith essays.Results: We found that the classifier built on Babel-generated essays and good-faith essays and using features from the automated scoring engine can distinguish the Babel-generated essays from the good-faith ones with 100% accuracy. We also found that if we integrated this classifier into the automated scoring engine it flagged very few responses that were submitted as part of operational submissions (76 of 434,656). The responses that were flagged had previously been assigned a score of Null (non-scorable) or a score of 1 by human experts. The measure of lexical-semantic cohesion shows promise in being able to distinguish Babel-generated essays from good-faith essays.Conclusions: Our results show that it is possible to detect the kind of gaming strategy illustrated by the Babel system and add it to an automated scoring engine without adverse effects on essays seen during real high-stakes tests. We also show that a measure of lexical-semantic cohesion can separate Babel-generated essays from good-faith essays to a certain degree, depending on task. This points to future work that would generalize the capability to detect semantic incoherence in essays. Directions for Further Research: Babel-generated essays can be identified and flagged by an automated scoring system without any adverse effects on a large set of good-faith essays. However, this is just one type of gaming strategy. It is important for developers of automated scoring systems to continue to be diligent about expanding the construct coverage of their systems in order to prevent weaknesses that can be exploited by tools such as Babel. It is also important to focus on the underlying linguistic reasons that lead to nonsense sentences. Successful identification of such nonsense would lead to improved automated scoring and feedback.

Journal
Journal of Writing Analytics
Published
2018-01-01
DOI
10.37514/jwa-j.2018.2.1.08
CompPile
Open Access
OA PDF Gold
Topics
Export

Citation context

Cited by in this index (1)

  1. Journal of Writing Analytics

References (38) · 2 in this index

  1. Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater® V. 2. The Journal of Technology, Lea…
  2. BABEL Generator (2014). Retrieved from http://babel-generator.herokuapp.com
  3. College Composition and Communication
  4. Beigman Klebanov, B., Madnani, N., Burstein, J., & Somasundaran, S. (2014). Content Importance models for sco…
     ↗
  5. Assessing Writing
Show all 38 →
  1. Bennett, R. E. (2015). The changing nature of educational assessment. Review of Research in Education, 39(1),…
     ↗
  2. Breiman, L. (2001). Random forests. Machine Learning, 45, 5-32.
     ↗
  3. Bridgeman, B., Trapani, C., & Attali, Y. (2012). Comparison of human and machine scoring of essays: Differenc…
     ↗
  4. Burstein, J., Tetreault, J., & Madnani, N. (2013). The E-rater® automated essay scoring system. Handbook of a…
  5. Carrell, P. L. (1982). Cohesion is not coherence. TESOL Quarterly, 16(4), 479-488.
     ↗
  6. Flor, M., & Beigman Klebanov, B. (2014). Associative lexical cohesion as a factor in text complexity. Interna…
     ↗
  7. Greene, P. (2018, July 2). Automated essay scoring remains an empty dream. Retrieved from Forbes: https://www…
  8. Halliday, M. A., & Hasan, R. (1976). Cohesion in English. London: Longman.
  9. Halliday, M. A., & Matthiessen, C. (2004). An introduction to Functional Grammar (3rd edition). London: Arnold.
  10. Heilman, M., Cahill, A., Madnani, N., Lopez, M., Mulholland, M., & Tetreault, J. (2014). Predicting grammatic…
     ↗
  11. Higgins, D., & Heilman, M. (2014). Managing what we can measure: Quantifying the susceptibility of automated …
     ↗
  12. Hoey, M. (1991). Patterns of lexis in text. Oxford University Press.
  13. Hoey, M. (2005). Lexical priming: A new theory of words and language. London: Routledge.
  14. Huang, L., Joseph, A. D., Blaine, N., Rubinstein, B. I., & Tygar, J. (2011). Adversarial machine learning. Pr…
     ↗
  15. Kane, M. T. (2013). Validating the interpretation and uses of test scores. Journal of Educational Measurement…
     ↗
  16. Klobucar, A., Deane, P., Elliot, N., Chaitanya, C., Deess, P., & Rudniy, A. (2012). Automated essay scoring a…
     ↗
  17. Levy, O., & Goldberg, Y. (2014). Linguistic regularities in sparse and explicit word representations. Proceed…
     ↗
  18. Lochbaum, K. E., Rosenstein, M., Foltz, P. W., & Derr, M. A. (2013, April). Detection of gaming in automated …
  19. Mandler, J. M., & Johnson, N. S. (1977). Remembrance of things parsed: Story structure and recall. Cognitive …
     ↗
  20. Marathe, M., & Hirst, G. (2010). Lexical chains using distributional measures of concept distance. Internatio…
     ↗
  21. Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., & Dean, J. (2013). Distributed representations of words…
  22. Morris, J., & Hirst, G. (1991). Lexical cohesion computed by thesaural relations as an indicator of the struc…
  23. Powers, D., Burstein, J., Chodorow, M., Fowles, M., & Kukich, K. (2001). Stumping E-Rater: Challenging the va…
     ↗
  24. Robinson, N. (2017, October 12). Push to have robots mark school tests under fire from prominent US academic.…
  25. Robinson, N. (2018, January 29). Robot marking of NAPLAN tests scrapped. Retrieved from ABC News: http://www.…
  26. Silber, H. G., & McCoy, K. F. (2002). Efficiently computed lexical chains as an intermediate representation f…
     ↗
  27. Smith, T. (2018, June 30). More states opting to 'robo-grade' student essays by computer. Retrieved from NPR:…
  28. Somasundaran, S., Burstein, J., & Chodorow, M. (2014). Lexical chaining for measuring discourse coherence qua…
  29. The Dada Engine. (2000). Retrieved from http://dev.null.org/dadaengine/
  30. Van Dijk, T. A. (1980). Macrostructures: An interdisciplinary study of global structures in discourse, intera…
  31. Williamson, D. M., Bejar, I. I., & Hone, A. S. (2005). 'Mental Model'™ Comparison of Automated and human scor…
     ↗
  32. Yoon, S.-Y., Cahill, A., Loukina, A., Zechner, K., Riordan, B., & Madnani, N. (2018). Atypical Inputs in educ…
     ↗
  33. Zhang, M., Chen, J., & Ruan, C. (2016). Evaluating the advisory flags and machine scoring difficulty in the e…
     ↗