Assessing Writing Constructs: Toward an Expanded View of Inter-Reader Reliability

Valerie Ross California University of Pennsylvania ; Rodger LeGrand Massachusetts Institute of Technology

Abstract

Background: This study focuses on construct representation and inter-reader agreement and reliability in ePortfolio assessment of 1,315 writing portfolios. These portfolios were submitted by undergraduates enrolled in required writing seminars at the University of Pennsylvania (Penn) in the fall of 2014.  Penn is an Ivy League university with a diverse student population, half of whom identify as students of color. Over half of Penn’s students are women, 12% are international, and 12% are first-generation college students. The students’ portfolios are scored by the instructor and an outside reader drawn from a writing-in-the-disciplines faculty who represent 24 disciplines. The portfolios are the product of a shared curriculum that uses formative assessment and a program-wide multiple-trait rubric. The study contributes to scholarship on the inter-reader reliability and validity of multiple-trait portfolio assessments as well as to recent discussions about reconceptualizing evidence in ePortfolio assessment.  Research Questions: Four questions guided our study: What levels of interrater agreement and reliability can be achieved when assessing complex writing performances that a) contain several different documents to be assessed; b) use a construct-based, multi-trait rubric; c) are designed for formative assessment rather than testing; and d) are rated by a multidisciplinary writing faculty?   What can be learned from assessing agreement and reliability of individual traits? How might these measurements contribute to curriculum design, teacher development, and student learning? How might these findings contribute to research on fairness, reliability, and validity; rubrics; and multidisciplinary writing assessment? Literature Review: There is a long history of empirical work exploring the reliability of scoring highly controlled timed writings, particularly by test measurement specialists. However, until quite recently, there have been few instances of applying empirical assessment techniques to writing portfolios.  Developed by writing theorists, writing portfolios contain multiple documents and genres and are produced and assessed under conditions significantly different from those of timed essay measurement. Interrater reliability can be affected by the different approaches to reading texts depending on the background, training, and goals of the rater. While a few writing theorists question the use of rubrics, most quantitatively based scholarship points to their effectiveness for portfolio assessment and calls into question the meaningfulness of single score holistic grading, whether impressionistic or rubric-based. Increasing attention is being paid to multi-trait rubrics, including, in the field of writing portfolio assessment, the use of robust writing constructs based on psychometrics alongside the more conventional cognitive traits assessed in writing studies, and rubrics that can identify areas of opportunity as well as unfairness in relation to the background of the student or the assessor. Scholars in the emergent field of empirical portfolio assessment in writing advocate the use of reliability as a means to identify fairness and validity and to create great opportunities for portfolios to advance student learning and professional development of faculty.  They also note that while the writing assessment community has paid attention to the work of test measurement practitioners, the reverse has not been the case, and that conversations and collaborations between the two communities are long overdue. Methodology: We used two methods of calculating interrater agreement: absolute and adjacent percentages, and Cohen’s Unweighted Kappa, which calculates the extent to which interrater agreement is an effect of chance or expected outcome. For interrater reliability, we used the Pearson product-moment correlation coefficient. We used SPSS to produce all of the calculations in this study.  Results: Interrater agreement and reliability rates of portfolio scores landed in the medium range of statistical significance.  Combined absolute and adjacent percentages of interrater reliability were above the 90% range recommended; however, absolute agreement was below the 70% ideal.  Furthermore, Cohen’s Unweighted Kappa rates were statistically significant but very low, which may be due to “kappa paradox.” Discussion: The study suggests that a formative, rubric-based approach to ePortfolio assessment that uses disciplinarily diverse raters can achieve medium-level rates of interrater agreement and reliability. It raises the question of the extent to which absolute agreement is a desirable or even relevant goal for authentic feedback processes of a complex set of documents, and in which the aim is to advance student learning. At the same time, our findings point to how agreement and reliability measures can significantly contribute to our assessment process, teacher training, and curriculum. Finally, the study highlights potential concerns about construct validity and rater training.  Conclusion: This study contributes to the emergent field of empirical writing portfolio assessment that calls into question the prevailing standard of reliability built upon timed essay measurement rather than the measurement, conditions, and objectives of complex writing performances.  It also contributes to recent research on multi-trait and discipline-based portfolio assessment.  We point to several directions for further research:  conducting “talk aloud” and recorded sessions with raters to obtain qualitative data on areas of disagreement; expanding the number of constructs assessed; increasing the range and granularity of the numeric scoring scale; and investigating traits that are receiving low interrater reliability scores. We also ask whether absolute agreement might be more useful for writing portfolio assessment than reliability and point to the potential “kappa paradox,” borrowed from the field of medicine, which examines interrater reliability in assessment of rare cases. Kappa paradox might be useful in assessing types of portfolios that are less frequently encountered by faculty readers. These, combined with the identification of jagged profiles and student demographics, hold considerable potential for rethinking how to work with and assess students from a range of backgrounds, preparation, and abilities.  Finally, our findings contribute to a growing effort to understand the role of rater background, particularly disciplinarity, in shaping writing assessment. The goals of our assessment process are to ensure that we are measuring what we intend to measure, specifically those things that students have an equal chance at achieving and that advance student learning.  Our findings suggest that interrater agreement and reliability measures, if thoughtfully approached, will contribute significantly to each of these goals.

Journal
Journal of Writing Analytics
Published
2017-01-01
DOI
10.37514/jwa-j.2017.1.1.09
CompPile
Open Access
OA PDF Gold
Topics
Export

Citation context

Cited by in this index (9)

  1. Across the Disciplines
  2. Journal of Writing Analytics
  3. Journal of Writing Analytics
  4. Pedagogy
  5. Journal of Writing Analytics
Show all 9 →
  1. Journal of Writing Analytics
  2. Journal of Writing Analytics
  3. Journal of Writing Analytics
  4. Journal of Writing Analytics

References (82) · 23 in this index

  1. Altman, D. (1991). Practical statistics for medical research (reprint 1999). Boca Raton, FL: CRC Press.
  2. American Educational Research Association (AERA), American Psychological Association (APA), & National Counci…
  3. Andrade, H. L. (2006). The trouble with a narrow view of rubrics. The English Journal, 95(6), 9-9.
     ↗
  4. Andrade, H. L., Du, Y., & Wang, X. (2008). Putting rubrics to the test: The effect of a model, criteria gener…
     ↗
  5. Anson, C. M., Dannels, D. P., Flash, P., & Gaffney, A. L. H. (2012). Big rubrics and weird genres: The futili…
Show all 82 →
  1. Attali, Y. (2016). A Comparison of newly-trained and experienced raters on a standardized writing assessment.…
     ↗
  2. Assessing Writing
  3. Assessing Writing
  4. Bejar, I. (2006). Automated scoring of complex tasks in computer-based testing. Psychology Press.
  5. Bejar, I. (2012). Rater cognition: Implications for validity. Educational Measurement: Issues and Practice, 3…
     ↗
  6. Broad, B. (2016). This is not only a test: Exploring structured ethical blindness in the testing industry. Jo…
  7. Brough, J. A., & Pool, J. E. (2005). Integrating learning and assessment: The development of an assessment cu…
  8. Bryant, L. H., & Chittum, J. R. (2013). ePortfolio effectiveness: A(n ill-fated) search for empirical evidenc…
  9. Chun, M. (2002). Looking where the light is better: A review of the literature on assessing higher education …
  10. Cushman, E. (2016). Decolonizing validity. Journal of Writing Assessment, 9(1). Retrieved from http://journal…
  11. Assessing Writing
  12. Assessing Writing
  13. Assessing Writing
  14. Elliot, N. (2016). A theory of ethics for writing assessment. Journal of Writing Assessment, 9(1). Retrieved …
  15. Elliot, N., Rudniy, A., Deess, P., Klobucar, A., Collins R., & Sava, S. (2016). ePortfolios: Foundational mea…
  16. Ewell, P. T. (1991). To capture the ineffable: New forms of assessment in higher education. Review of Researc…
     ↗
  17. Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. …
     ↗
  18. Fleiss, J. L. (1981). Statistical methods for rates and proportions. 2nd ed. New York: John Wiley.
  19. Freedman, S. W., & Calfee, R. C. (1983). Holistic assessment of writing: Experimental design and cognitive th…
  20. The WAC Journal
  21. Graham, M., Milanowski, A., & Miller, J. (2012). Measuring and promoting interrater agreement of teacher and …
  22. Hafner, J., & Hafner, P. (2003). Quantitative analysis of the rubric as an assessment tool: An empirical stud…
     ↗
  23. Hamp-Lyons, L. (1991). Assessing second language writing in academic contexts. Norwood, NJ: Ablex Publishing …
  24. Assessing Writing
  25. Assessing Writing
  26. Hartmann, D. P. (1977). Considerations in the choice of interobserver reliability measures. Journal of Applie…
     ↗
  27. Huber, M., & Hutchings, P. (2004). Integrative learning: Mapping the terrain. Washington, DC: American Associ…
  28. College Composition and Communication
  29. Hutchings, P. T. (1990). Learning over time: Portfolio assessment. American Association of Higher Education B…
  30. Assessing Writing
  31. Assessing Writing
  32. Inoue, A. B. (2015). Antiracist writing assessment ecologies: Teaching and assessing writing for a socially j…
     ↗
  33. Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational conseque…
     ↗
  34. Kelly-Riley, D., Elliot, N., & Rudniy, A. (2016). An empirical framework for ePortfolio assessment. Internati…
  35. Assessing Writing
  36. Kohn, A. (2006). Speaking my mind: The trouble with rubrics. English Journal, 95(4), 12-15.
     ↗
  37. Landis, J. R., & Koch, G. G. (1977). A one way components of variance model for categorical data. Biometrics,…
     ↗
  38. Looney, J. W. (2011). Integrating formative and summative assessment: Progress toward a seamless system, OECD…
  39. Mansilla, V. B., & Duraisingh, E. D. (2007). Targeted assessment of students' interdisciplinary work: An empi…
     ↗
  40. Mansilla, V. B., Duraisingh, E. D., Wolfe, C. R., & Haynes, C. (2009). Targeted assessment rubric: An empiric…
     ↗
  41. Meier, S. L., Rich, B. S., & Cady, J. (2006). Teachers' use of rubrics to score non- traditional tasks: Facto…
     ↗
  42. Myers, M. (1980). A procedure for writing assessment and holistic scoring. ERIC Clearinghouse on Reading and …
  43. National Science Foundation. (2015). Collaborative Research: The Role of Instructor and Peer Feedback in Impr…
  44. Newell, J. A., Dahm, K. D., & Newell, H. L. (2002). Rubric development and inter-rater reliability issues in …
  45. Penny, J., Johnson, R. L., & Gordon, B. (2000). Using rating augmentation to expand the scale of an analytic …
     ↗
  46. Poe, M., & Cogan, J.A. (2016). Civil rights and writing assessment: Using the disparate impact approach as a …
  47. Pula, J.J., & Huot, B.A. (1993). A model of background influences on holistic raters. In M.M. Williamson & B.…
  48. Assessing Writing
  49. Ross, V., Liberman, M., Ngo, L., & LeGrand, R. (2016a). Weighted log-odds- ratio, informative Dirichlet prior…
  50. Ross, V., Wehner, P., & LeGrand, R. (2016b). Tap Root: University of Pennsylvania's IWP and the financial cri…
  51. Schneider, C. G. (2002). Can value added assessment raise the level of student accomplishment? Peer Review, 4…
  52. Shepard, L. A. (2000). The role of assessment in a learning culture. Educational Researcher, 29(7), 4-14.
     ↗
  53. Shohamy, E., Gordon, C. M., & Kraemer, R. (1992). The effect of raters background and training on the reliabi…
     ↗
  54. Slomp, D. (2016). Ethical considerations and writing assessment. Journal of Writing Assessment, 9(1). Retriev…
  55. Stemler, S. E. (2004). A comparison of consensus, consistency, and measurement approaches to estimating inter…
  56. Stemler, S.E., & Tsai, J. (2016). Best practices in interrater reliability: Three common approaches. In J. Os…
     ↗
  57. Stock, P. L., & Robinson, J. L. (1987). Taking on testing: Teachers as tester- researchers. English Education…
     ↗
  58. Thaiss, C., & Zawacki, T. M. (2006). Engaged writers and dynamic disciplines: Research on the academic writin…
  59. Tinsley, H. E. A., & Weiss, D. J. (2000). Interrater reliability and agreement. In H. E. A. Tinsley & S. D. B…
     ↗
  60. Assessing Writing
  61. University of Pennsylvania. (2014). Assessment of Student Learning. Accreditation and 2014 self-study report.…
  62. University of Pennsylvania. (2016). Incoming Class Profile. Philadelphia, PA: Author. Retrieved from http://w…
  63. Vann, R. J., Meyer, D. E., & Lorenz, F. O. (1984). Error gravity: A study of faculty opinion of ESL errors. T…
     ↗
  64. Vaughan, C. (1992). Holistic assessment: what goes on in the rater's mind? In Hamp-Lyons, L., editor, Assessi…
  65. Viera, A. J., & Garrett, J. M. (2005). Understanding interobserver agreement: The kappa statistic. Fam Med, 3…
  66. Assessing Writing
  67. College Composition and Communication
  68. White, E. M. (1985). Teaching and assessing writing: Understanding, evaluating and improving student performa…
  69. Assessing Writing
  70. College Composition and Communication
  71. White, E. M., Elliot, N., & Peckham, I. (2015). Very like a whale. Boulder, Colorado: University Press of Colorado.
  72. Assessing Writing
  73. Wilson, M. (2006). Rethinking rubrics in writing assessment. Portsmouth, NH: Heinemann.
  74. Assessing Writing
  75. Assessing Writing
  76. Written Communication
  77. College Composition and Communication