-
Eubanks, D. A. (2008). Assessing the general education elephant. Assessment Update, 20(4), 4.
-
Eubanks, D. A. (2016a) Interrater facets. Github Repository. http://github.com/stanislavzza/Inter-Rater-Facets
-
Eubanks, D. A. (2016b) Rethinking interrater agreement. Assessment Update, 28(4), 8-14. DOI: 10.1002/au.30065
-
Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378.
-
Fleiss, J. L., Cohen, J., & Everitt, B. (1969). Large sample standard errors of kappa and weighted kappa. Psy…
-
Fleiss, J. L., Levin, B., & Paik, M. C. (2003). Statistical methods for rates and proportions (3rd ed.). Hobo…
-
Galton, F. (1892). Finger prints. London: Macmillan and Company.
-
Gamer, M., Lemon, J., & Singh, I. F. P. (2012). Irr: Various coefficients of interrater reliability and agree…
-
Gwet, K. (2015). Testing the difference of correlated agreement coefficients for statistical significance. Ed…
-
Gwet, K. L. (2014). Handbook of inter-rater reliability:Tthe definitive guide to measuring the extent of agre…
-
Haertel, E. H. (2006). Reliability. In R. L. Brennan (Ed.), Educational measurement (4th ed., pp. 65-110). We…
-
Hodgson, R. T. (2008). An examination of judge reliability at a major US wine competition. Journal of Wine Ec…
-
Kelly Riley, D., & Whithaus, C. (2016). A theory of ethics for writing assessment. [Special issue]. Journal o…
-
Landis, J. R. and Koch, G. G. (1977) The measurement of observer agreement for categorical data. Biometrics, …
-
Lane, S, & Stone, C. (2006), Performance Assessment. In R. L. Brennan (Ed.), Educational measurement (4th ed.…
-
Messick, S. (1986). The once and future issues of validity: Assessing the meaning and consequences of measure…
-
Moss, P. A. (2004). The meaning and consequences of "reliability". Journal of Educational and Behavioral Stat…
-
Moss, P. A., Pullin, D. C., Gee, J. P., Haertel, E. H., & Young, L. J. (Eds.). (2008). Assessment, equity, an…
-
National Research Council. (2012). Education for life and work: Developing transferable knowledge and skills …
-
Powers, D. M. (2012). The problem with kappa. In Proceedings of the 13th Conference of the European Chapter o…
-
R Core Team. (2015). R: A language and environment for statistical computing [Computer software manual]. Vien…
-
Rezaei et al. (2010)
Assessing Writing
-
Roberts, C. (2008). Modelling patterns of agreement for nominal scales. Statistics in Medicine, 27(6), 810-830.
-
Smeeton, N. C. (1985). Early history of the kappa statistic. International Biometric Society, 41(3).
-
Stanovich, K. E. (1986). Matthew effects in reading: Some consequences of individual differences in the acqui…
-
Stemler, S. E. (2004). A comparison of consensus, consistency, and measurement approaches to estimating inter…
-
Wang, J., Engelhard, G., Raczynski, K., Song, T., & Wolfe, E. W. (2017). Evaluating rater accuracy and percep…