Peer assessment is reliable and valid on average, but that average hides real variation across courses. Xiong, Schunn, and Wu identified what predicts when peer assessment works well and when it doesn't.
Xiong, Y., Schunn, C. D., Wu, Y. (2023). "What predicts variation in reliability and validity of online peer assessment?: A large-scale cross-context study." Journal of Computer Assisted Learning, 39(6), 2004–2024.
Decades of research, including Peerceptiv's own founding study, established that peer assessment is generally reliable and valid. But "generally" is doing a lot of work in that sentence. This study asked the more useful question for anyone actually running a peer-review program: what specific, controllable factors predict whether peer assessment is reliable and valid in a given course, versus not?
The researchers analyzed peer assessment data across a large number of courses and contexts, testing which design and implementation factors were associated with higher or lower reliability and validity of the resulting peer ratings.
Rubric design matters concretely: rubrics aligned to the actual performance range of students in a class — ones with real variation across the scale, rather than a scale where everyone clusters near the top or bottom — produce more reliable and valid peer ratings.
A rubric that can't distinguish between good and great work is a rubric that will produce noisy, unreliable peer ratings no matter how well-trained the reviewers are. Before troubleshooting a peer-review program's reliability by adding more reviewer training, check whether the rubric itself has room to discriminate at the actual performance levels employees are working at.
Chris Schunn co-authored this large-scale study identifying the concrete design factors that make peer assessment reliable. He is a Professor of Psychology, Learning Sciences and Policy, and Intelligent Systems at the University of Pittsburgh and a Senior Scientist at the University's Learning Research and Development Center (LRDC), where he has directed research projects backed by more than $80M in federal grants. As Peerceptiv's Chief Learning Scientist, his research directly shapes how the platform structures reviewer prompts, rubrics, and feedback workflows.