HumanOS for assessment: putting AI in the human loop

Imagine being told that AI helped mark a piece of work you spent weeks producing. What would you want to know?

You might ask whether it recognised what was good about your work, rather than simply matching it to familiar patterns. You might want to know whether a qualified person would consider your work if the result seemed wrong. Above all, you would want to know that the organisation awarding your qualification remained accountable for the decision.

These are not objections to technology. They are reasonable expectations of an assessment process that treats candidates with dignity.

As AI becomes more capable, we need to look more closely at a familiar phrase: human in the loop. It can describe meaningful human control. But it can also describe a process in which AI leads and people review selected outputs. That leaves a more fundamental question unanswered: who decided what the AI should value in the first place?

What if we reversed the starting point and put AI in the human loop?

Human judgement sets the standard

Brighteye Ventures describes a HumanOS: systems that help people learn, work and adapt as software becomes more intelligent. Its premise is that technology can multiply the value of human skills, knowledge and talent.

In assessment, that idea matters. A candidate has not submitted work merely to receive a number. They have asked an awarding organisation to make a considered judgement about what they know or can do. AI may help with that task, but it should not quietly become the source of the standard against which the candidate is judged.

For an awarding organisation, a HumanOS approach starts with the assessment and the people who understand it. What does quality look like in this subject? Where do experienced assessors agree, and where is judgement more difficult? What kinds of work might an AI system misunderstand?

Only then do we ask what role AI can responsibly play. This does not mean a human must read every response in every scenario. It means AI’s role is defined within a human-governed assessment process, rather than human involvement being added around an AI-led one.

Making the principle testable

A commitment to human judgement needs more than reassuring language. Awarding organisations need practical ways to examine whether an AI approach reflects the standards their experts apply.

This is the thinking behind the proposed RM Compare AI Validation Layer. RM Compare can bring together judgements from qualified assessors to develop a structured human reference. An AI approach can then be tested against it: where do its judgements align with the assessors’, where do they differ, and which differences might affect a candidate’s outcome?

The most useful finding may not be a headline accuracy figure. It may be a piece of work that experienced assessors value highly but the AI scores poorly, or the reverse. Investigating those differences could help an awarding organisation understand the limits of the AI approach, improve the way it is used, and design more effective human review.

We are exploring this through a proof of concept. It does not certify AI marking, and calibrating AI scores against human judgements cannot guarantee the right result for every candidate. The aim is to test whether comparative judgement can provide a useful human reference for deeper, assessment-specific validation.

What candidates should be able to expect

If an awarding organisation uses AI in assessment, candidates deserve an honest account of what happened. Was their individual work reviewed by a person? Or was the AI process tested against human judgements made on other work? Those are different assurances, and we should not blur them.

Candidates should also know who is accountable for their result and have a meaningful route to question it if they believe it fails to reflect their work. A process cannot claim to respect human judgement while leaving the person most affected by it unable to challenge an outcome.

That is why HumanOS is more than a new label for oversight. It is a design test for the whole assessment experience: does AI strengthen an awarding organisation’s ability to make and stand behind good judgements, while preserving the candidate’s dignity and trust?

That question cannot be answered by a model provider alone. It needs to be explored with awarding organisations, their expert assessors and the people whose work is being assessed.

Putting AI in the human loop is not a promise that technology will never get things wrong. It is a commitment to keep human standards, human accountability and the candidate’s experience at the centre as assessment evolves.