How Swedish Teachers used RM Compare to support assessment calibration at scale

The challenge

Across school systems and education ministries, assessment is increasingly focused on complex, open-ended outcomes and artefacts.

But consistently evaluating skills, knowledge and understanding in this context can be difficult. Traditional rubric-based marking can lead to variation between Assessors, an emphasis on surface-level features rather than the quality of learning, and substantial moderation workload.

Dr Eva Hartell led a repeated-measures study in Sweden to explore whether a distributed assessors group could develop a shared understanding of quality through Adaptive Comparative Judgement (ACJ). The study also examined whether that shared standard could help create reliable assessment scales with fewer judgements over time.

Summary of findings

A very short Comparative Judgement experience, even for Teachers new to the system, delivered highly impactful results

The results support previous research concerning the power and impact of Learning by Evaluating. Participants found the experience engaging, enjoyable and rewarding. The time commitment was small and the impact very high.

The study shows that there is now a clear way froward for sharing and securing learners performance standards across schools. This can be done at scale across regions, countries or even the entire world.

The RM Compare 3 module Ecosystem provides a powerful way forward.

The potential for assessors, teachers and students - as a matter of normal practice - to have access to, and familiarity with, work from a national sample of schools and not just their own classroom has profound implications for the curriculum, pedagogy and assessment.

The approach

The study used RM Compare across two linked sessions with the same assessor panel and assessment context.

  • Establishing a baseline: In the first session, Assessors judged complex student work using their existing assessment habits and internal standards. This provided a baseline rank order of work, while revealing the natural variation in evaluator alignment, decision times and approaches to judging quality.
  • Testing calibration: The same panel then completed a shorter second session in the same assessment context. The platform configuration, assessment domain and item pool remained stable, allowing the study to explore the effect of Assessors experience.

The central question was simple: after working comparatively through the first session, would Assessors have developed a clearer, more shared view of quality - and could RM Compare produce a stable scale more efficiently as a result?

The impact

More consistent evaluator judgement

RM Compare uses a judge-misfit measure to show how closely an Assessor’s decisions align with the emerging consensus of the group.

During the first session, Assessor alignment varied widely. Some judges’ decisions differed noticeably from the group’s emerging view of quality.

In the second session, those patterns moved closer to the cohort consensus. Misfit values clustered into a tighter range, suggesting that the same Assessors were making more consistent decisions after taking part in the initial comparative-judgement exercise.

Rather than relying on lengthy moderation meetings to achieve alignment, the comparative process helped Assessors calibrate through the act of judging.

Greater fluency in decision-making

Decision-time data showed a complementary pattern.

In the baseline session, decision times varied considerably, including longer periods of hesitation and occasional rapid decisions. In the second session, decision times became more consistent, with Assessors making judgements at a steadier professional pace.

This suggests that Assessors had begun to internalise a clearer standard of quality. With less cognitive friction, they could focus on making informed professional judgements rather than repeatedly interpreting criteria from first principles.

From checklists to holistic judgement

The study also identified a meaningful change in the way Assessors described their decisions.

In the first session, comments often reflected a checklist mindset, with attention given to layout, formatting and other surface-level features. By the second session, comments focused more often on clarity of reasoning, conceptual depth and the coherence of each response.

Assessors also used more precise, shared language connected to curriculum and subject expectations. Their comments became shorter and more decisive, indicating greater confidence in how they defined and recognised quality.

Stable scales with fewer judgements

Although the second session was intentionally shorter, it still produced a stable scale.

High-quality anchor pieces and lower-performing work occupied similar positions across both sessions, while the overall ordering of work showed strong agreement. Difficult boundary cases - work that sits between performance levels - also settled more consistently into intermediate positions.

The findings suggest that, once a panel has developed a shared standard, RM Compare can support the construction of reliable scales with reduced judging volume.

A framework for scalable assessment

Dr Hartell’s study points towards a practical approach for systems seeking to build more consistent, efficient and transparent assessment.

Using the RM Compare 3 Module Ecosystem

At the heart of the RM Compare 3 Module Ecosystem is the ability to turn Ranks into Rulers. This revolutionary technology is set to change the assessment landscape with the introduction of Comparative Judgement on-demand for the first time.

Want to know more? Get in touch.

What next?

The next phase of the project is underway with a larger sample of Teachers from across Sweden. The intention is to also test the approach across more curriculum areas.

Research notes

This case study presents anonymised, directional observations from platform telemetry. Full quantitative results, participant details and methodological analysis will be released separately following publication of the primary academic study