Blog
Posts for category: Research
-
Beyond the Black Box: What New 2026 Research Teaches Us About Validating AI Assessment
As artificial intelligence and Large Language Models (LLMs) become deeply embedded in global education, a central question looms over the industry: Can we truly trust AI to act as an autonomous judge of quality?
-
Fewer Comparisons, Higher Bar: What the Latest ACJ Research Means for RM Compare Users
Recent research indicates that we can aim higher on reliability while asking judges to do less work. For RM Compare users, that is good news: it supports designing sessions that are both more efficient and more defensible.
-
Setting and Maintaining Verifiable Standards in Complex Curriculums and Aligning Teacher Assessment at Scale
As educational ministries and professional bodies globally modernise their frameworks, learning objectives have fundamentally shifted. Traditional education models assessed knowledge recall through closed, easily quantifiable test formats.
-
The State of Learning by Evaluating
There's a familiar moment in teaching that almost every educator recognises. You hand out the rubric, explain the assignment, and watch thirty faces look back at you with the same expression — a kind of polite blankness that says I understand the words, but I don't yet understand what you mean. The criteria make sense in the abstract. What "good" looks like in practice is another matter entirely.
-
From Steady State to Rulers: how RM Compare is building the future of shared standards
Assessment systems often talk about standards, but too often those standards remain abstract. Teachers, examiners and assessors are expected to align to a wider benchmark, yet in day-to-day practice they usually see only the work directly in front of them: their own class, their own cohort, their own centre. That gap matters.
-
When AI Beats Economists – And Why That’s Good News For Assessment
Not so long ago, the idea that an AI system could out‑analyse a room full of economists would have sounded like science fiction. Yet that’s exactly what a recentFederal Reserve working paper set out to test.
-
Time for Tea? Can AI spot a nice cuppa?
Just like everyone else in the UK we are getting very excited about National Tea Day which takes place on the 21st April. The day is a great opportunity to share tea brewing preferences, and the strength of the perfect 'cuppa' is always hotly debated. So we though it would be interesting to get the view of AI.
-
Using Comparative Judgement to keep exam grades fair when tests change (OFQUAL Research report 2025)
Every year, exams change. New papers are written, formats evolve, and sometimes whole qualifications are refreshed. Yet everyone from students, parents, teachers and universities still expects one simple promise to hold: a grade this year should mean the same as a grade last year. So what do you do then?
-
The OECD Just Mapped the Certification Problem. Here's the Solution.
A response to the OECD's The Theory and Practice of Upper Secondary Certification (2026). The OECD report has mapped the territory of the problem with exceptional care. The standardisation challenge is real, it is persistent, and it has defeated every country that has tried to solve it from within the existing paradigm. The solution is not a better mark scheme. It is a better question.