How Swedish Teachers used RM Compare to support assessment calibration at scale

Across school systems and education ministries, assessment is increasingly focused on complex, open-ended outcomes.

But consistently evaluating skills, knowledge and understanding in this context can be difficult. Traditional rubric-based marking can lead to variation between evaluators, an emphasis on surface-level features rather than the quality of learning, and substantial moderation workload.

Dr Eva Hartell led a repeated-measures study in Sweden to explore whether a distributed evaluator group could develop a shared understanding of quality through Adaptive Comparative Judgement (ACJ). The study also examined whether that shared standard could help create reliable assessment scales with fewer judgements over time.

Dr Hartell will be presenting the full findings from her research at the ResearchEd National Conference on Saturday September 5th 2026.

You can read the full Case Study here right now.

The approach

The study used RM Compare across two linked sessions with the same evaluator panel and assessment context. The central question was simple: after working comparatively through the first session, would Assessors have developed a clearer, more shared view of quality - and could RM Compare produce a stable scale more efficiently as a result?

The potential for assessors, teachers and students - as a matter of normal practice - to have access to, and familiarity with, work from a national sample of schools and not just their own classroom has profound implications for the curriculum, pedagogy and assessment.

Summary of findings

A very short Comparative Judgement experience, even for Teachers new to the system, delivered highly impactful results

  • Assessors were making more consistent decisions after taking part in the comparative-judgement exercise.
  • Assessors had internalised a clearer standard of quality. With less cognitive friction, they could focus on making informed professional judgements rather than repeatedly interpreting criteria from first principles.
  • Assessors also used more precise, shared language connected to curriculum and subject expectations. Their comments became shorter and more decisive, indicating greater confidence in how they defined and recognised quality.
  • The findings suggest that, once a panel has developed a shared standard, RM Compare can support the construction of reliable scales with reduced judging volume.

The results support previous research concerning the power and impact of Learning by Evaluating. Participants found the experience engaging, enjoyable and rewarding. The time commitment was small and the impact very high.

The study shows that there is now a clear way froward for sharing and securing learners performance standards across schools. This can be done at scale across regions, countries or even the entire world.

The RM Compare 3 module Ecosystem provides a powerful way forward.