- Research
Fewer Comparisons, Higher Bar: What the Latest ACJ Research Means for RM Compare Users
For years, the rule of thumb in Adaptive Comparative Judgement (ACJ) was straightforward: aim for a Scale Separation Reliability (SSR) of around 0.70, and allow roughly 26–37 comparisons per item if you wanted to push reliability towards 0.90. That guidance, drawn from a widely cited 2018 meta‑analysis of comparative judgement studies, has shaped how assessment designers plan workload, set judging windows and justify ACJ to sceptical stakeholders.
Recent research invites us to rethink that picture, not because ACJ has become less trustworthy, but because we can aim higher on reliability while asking judges to do less work. For RM Compare users, that is good news: it supports designing sessions that are both more efficient and more defensible.
ACJ, SSR and the inflation story
Adaptive Comparative Judgement is an approach to assessing complex work such as essays, portfolios, projects where judges compare pairs of student work and decide which is better, instead of scoring each piece against a detailed rubric. Over many comparisons, an underlying scale of quality emerges and each script finds its place on that scale.
Scale Separation Reliability (SSR) is the statistic most commonly used to summarise how consistently judges’ decisions support that emerging scale. Higher SSR suggests a stable rank order and broad agreement among judges. However, earlier work raised a serious question: could ACJ systems be reporting reliability that looked better than reality?
In 2015, Bramley used simulated data to show that naïve adaptive algorithms could generate “spurious separation”: scales that appeared impressively reliable on paper even when the underlying judgements were mediocre. In those simulations, aggressive adaptivity combined with very few comparisons per item made SSR more optimistic than the true inter‑rater reliability. The system, in effect, seemed to know more than it really did.
That critique was taken seriously. Researchers such as Kimbell and colleagues, working closely with platform developers including the RM Compare team refined pairing algorithms and emphasised reading SSR in context rather than as a standalone badge of quality. The aim was to reduce the risk of inflation and to encourage practitioners to look beyond a single headline number.
More recently, studies using RM Compare in real high‑stakes contexts have added empirical reassurance. In spoken language proficiency work led by Wang and Zheng, high SSR values were backed up by strong split‑half reliability, good alignment with Rasch‑based models, and close agreement between ACJ ranks and expert judgements based on detailed rubrics. Where RM Compare reported high reliability, other methods agreed that the judgements were genuinely robust, rather than artefacts of the algorithm.
Against that backdrop that inflation concerns were real, prompted changes, and have since been checked, the latest meta‑analytic work provides updated guidance on two practical questions: how many comparisons you really need, and what SSR threshold you should now treat as “reliable enough”.
What the new meta‑analysis tells us
A major meta‑analysis published in 2025 re‑examined over a hundred comparative judgement datasets across education and other domains. Two findings stand out.
First, the study suggests that robust scale reliability can often be achieved with as few as 10 comparisons per item. This is a substantial shift from the earlier 26–37 comparisons rule of thumb. It does not mean that 10 will always suffice in every context, but it does show that, in realistic ACJ configurations, you can reach high reliability without imposing heavy, research‑style workloads on judges.
Second, the authors argue that practitioners should treat an SSR of about 0.8 (rather than the traditional 0.7) as the meaningful threshold for a trustworthy ranking in consequential contexts. By aiming for 0.8 rather than relaxing once you pass 0.7, you build in a buffer against residual inflation and ordinary measurement noise.
Put simply: the new evidence says you can often get to a robust level of reliability with far fewer comparisons than previously assumed, but you shouldn’t be content with “just past 0.7” when the decisions you are making carry serious consequences.
What this means for high‑stakes RM Compare sessions
Bringing together the inflation story, the validation work and the new workload guidance gives RM Compare users clearer design rules for summative and other high‑stakes scenarios.
When outcomes matter it makes sense to treat SSR 0.8 as your target rather than 0.7. For these contexts, RM Compare sessions should be planned so that the scale climbs to at least that level before you stop judging. The latest research indicates that you can do this with leaner workloads than older rules suggested. Instead of locking in 26–37 comparisons per item, you can design judging rounds around a baseline of around 10 comparisons per item, then monitor how the scale behaves and extend only if needed.
RM Compare’s live reporting make this approach practical. As a session runs, you can see SSR and how often each item has appeared. That allows you to start with a lean plan, watch reliability improve and decide whether to add further rounds only if the scale has not yet reached your desired threshold. Rather than committing in advance to heavy workloads, you treat SSR and exposure counts as dynamic signals that tell you when you have “enough” evidence.
High‑stakes use also demands attention to validity, not just reliability. A high SSR tells you that judges were consistent with each other; it does not guarantee that they focused on the right construct. Some studies have found that ACJ can favour clarity and structure over depth of content in certain tasks. RM Compare makes it straightforward to sample borderline scripts, compare ACJ ranks with rubric‑based grades, or run small cross‑check exercises. Doing this periodically in high‑stakes work helps ensure that the ranking reflects the kind of quality you intend to measure.
Taken together, this lets you tell a stronger story to regulators and stakeholders. You can say that RM Compare’s judgements are produced by an algorithm that has responded to inflation critiques, validated against multiple independent checks, and used within a framework that aims higher on reliability while managing judge workload responsibly.
Implications for everyday and non‑high‑stakes use
Not every RM Compare session is a high‑stakes exercise. Many are about learning, moderation, professional development or exploratory research. The same evidence base has implications here, and they are largely liberating rather than restrictive.
In everyday classroom projects, departmental moderation, peer assessment or professional learning tasks, it is often enough to see SSR in the 0.7 - 0.8 range. In these contexts, the goal is to surface a sensible rank order, stimulate discussion and generate feedback, not to certify or award. Knowing that robust scales can be achieved with as few as 10 comparisons per item means you can safely run lighter rounds, accepting slightly lower SSR in exchange for faster turnaround, lower teacher burden and the ability to run ACJ more frequently.
The emerging rule of thumb is simple: for high‑stakes decisions, design towards SSR of 0.8 or above and plan around ten or more comparisons per item, with occasional validity checks built in. For non‑high‑stakes use, SSR around 0.7+ with leaner workloads is typically sufficient, provided the outcomes are used for feedback, ranking or calibration rather than formal awards.
Because multiple studies now show that ACJ scales from systems like RM Compare correlate well with rubric scores and expert judgements, you can have confidence using the tool not just for big summative moments but for regular cycles of “learning by evaluating”, internal moderation and programme review. The same engine that can underpin high‑stakes decisions can also support formative practice.
The bottom line for RM Compare
The trajectory of research around ACJ and SSR has, at times, been challenging. Early inflation critiques forced difficult questions about how reliability should be measured and interpreted. For RM Compare, however, the direction of travel is strongly positive.
The platform now sits on an evidence base showing that inflation concerns were real but have been understood and addressed, that its high SSR values align with other reliability and validity checks, and that realistically achievable workloads - as few as 10 comparisons per item - can still deliver robust scales when you aim for an SSR of around 0.8 in consequential contexts.
For RM Compare users, the practical message is clear. You can design ACJ sessions that are less burdensome for judges and more robust for stakeholders by planning leaner workloads, aiming higher on SSR when stakes demand it, and keeping validity checks in view alongside reliability. That combination positions RM Compare not just as a convenient ACJ platform, but as a research‑aligned assessment engine you can trust in both everyday practice and high‑stakes decision making.