With modest training on a structured rubric, groups of student peer reviewers produced ratings of writing quality that matched individual expert graders on both reliability and validity.
Cho, K., Schunn, C. D., & Wilson, R. W. (2006). "Validity and Reliability of Scaffolded Peer Assessment of Writing from Instructor and Student Perspectives." Journal of Educational Psychology, 98(4), 891–901.
Peer review creates more opportunities for students to write and get feedback than an instructor could provide alone, but it raises an obvious concern: can grades assigned by peers, most of whom have no training as evaluators, be trusted? Cho, Schunn, and Wilson set out to test whether structured peer-generated grades hold up against the two standards that matter for any assessment — reliability (do independent raters agree with each other) and validity (do the ratings match what an expert would conclude).
This study was conducted on SWoRD (Scaffolded Writing and Rewriting in the Discipline), the web-based reciprocal peer review system built at the University of Pittsburgh's Learning Research and Development Center. SWoRD is the system that was later licensed, renamed, and commercialized as Peerceptiv — this paper is direct research on Peerceptiv's own predecessor platform.
College students, given modest training in how to apply a structured evaluation rubric, rated their peers' writing assignments. Researchers compared the reliability of these peer ratings against each other (using intraclass correlation) and their validity against ratings from subject-matter experts and instructors.
Individual peer ratings varied, the same way individual expert ratings vary. But aggregating ratings from multiple peer reviewers produced scores that were highly reliable and moderately-to-highly valid compared to individual expert ratings. The multiplicity of reviewers is what makes peer assessment trustworthy: no single peer needs to grade like an expert, because the collective judgment of several trained reviewers converges on an accurate score.
This is the founding evidence for a claim workforce L&D teams are often skeptical of: that employees without formal evaluator training can produce trustworthy assessments of each other's work. The mechanism is aggregation, not individual expertise. A single employee reviewing a colleague's work isn't a reliable substitute for a manager's judgment. Several employees reviewing the same piece of work, using a shared rubric, is. This is why Peerceptiv assigns multiple reviewers to every piece of submitted work rather than relying on a single peer's assessment — it's the same design principle this study established nearly two decades ago, running on the same underlying platform.
Chris Schunn co-authored this foundational study on SWoRD, the system he helped build and that later became Peerceptiv. He is a Professor of Psychology, Learning Sciences and Policy, and Intelligent Systems at the University of Pittsburgh and a Senior Scientist at the University's Learning Research and Development Center (LRDC), where he has directed research projects backed by more than $80M in federal grants. As Peerceptiv's Chief Learning Scientist, his research directly shapes how the platform structures reviewer prompts, rubrics, and feedback workflows.