From Marked Evidence to Reusable Rulers: A Practical Path for Awarding Organisations (Part 2/3)

In the first article in this series, we explored why assessment organisations need more than a repository of examples. They need a standards repository: a governed environment in which evidence of quality can be preserved, understood and reused.

For awarding organisations, that raises an immediate practical question.

They already hold a significant body of standards evidence. There are marked candidate responses, definitive scripts, examiner annotations, mark schemes, standardisation materials and awarding records. Much of it has been created through robust, rubric-based marking processes.

Can that evidence become a reusable ruler in RM Compare?

The answer is yes, but not automatically, and not by pretending that rubric-based marking and comparative judgement are the same thing.

The opportunity is to begin with the evidence an awarding organisation already trusts, preserve the context and authority that make it meaningful, and then apply Adaptive Comparative Judgement where it adds value: creating a stable, reusable relative scale for suitable forms of complex work.

Start with evidence you already trust

Awarding organisations do not start with an empty repository.

Every assessment cycle creates a valuable record of how standards have been interpreted and applied. Marks show how performance was judged against an approved scheme. Annotations record examiner reasoning. Definitive and seed scripts make standard-setting decisions visible. Standardisation and moderation processes capture the shared professional understanding that sits behind consistent assessment.

This is not simply operational material left over after an assessment series. It is an organisation’s accumulated evidence of quality.

Yet it is often difficult to use beyond its immediate purpose. An approved script may be embedded in a PDF. A key annotation may remain locked in a marking system. A senior examiner may know exactly which examples illuminate a difficult boundary, but that knowledge may not be structured, searchable or readily available to the next examiner cohort.

A standards repository creates a more durable home for that evidence. It can retain the original mark, grade, task, mark scheme, annotation, specification version and approval history alongside the candidate work itself.

This allows the organisation to begin with a simple but powerful proposition:

The standards evidence created through marking should remain useful after the marking window closes.

A rank order is useful - but it is not yet an ACJ ruler

Imagine an awarding organisation has selected a set of approved responses to an extended-writing assessment. Each response has been marked against the published rubric, reviewed by senior examiners and linked to a confirmed mark or grade.

Those responses can be sorted into an order based on their total marks. The organisation may have excellent work at the upper end, examples around key grade boundaries and a range of responses that reveal how the mark scheme has been interpreted in practice.

That ordered set is already useful.

It can support examiner induction. It can help standardisation discussions. It can make an organisation’s interpretation of performance levels more visible. It can give moderators and quality leads a reliable starting point when they need to retrieve examples of work across a mark range.

But it is not automatically an ACJ ruler.

A mark-derived order reflects the logic of the original mark scheme. The assessment may award credit separately for knowledge, communication, technical accuracy, method, task fulfilment or other criteria. It may deliberately weight one assessment objective more heavily than another. It may contain non-compensatory requirements, where a response must demonstrate a particular feature to receive credit.

An ACJ ruler is established differently. Experts make repeated pairwise comparisons of complete responses against an agreed comparative construct. Over time, those comparisons create a relative scale. The ruler is valuable when the relative positions of its reference items have become sufficiently stable for the defined purpose.

These are complementary forms of evidence, but they are not interchangeable.

A mark-derived reference set tells us:

“This work received a higher outcome than that work under the approved rubric and marking process.”

An ACJ-calibrated ruler tells us:

“Expert judges consistently placed this work above that work against the agreed comparison construct.”

The difference matters because RM Compare should preserve the authority of the original marks without making a false claim that mark order alone proves comparative stability.

The construct comes first

The key question is not, “Can we place marked scripts in order?”

Of course we can.

The more important question is:

Does the comparison that RM Compare users will make represent the same construct that the original rubric was designed to assess?

Consider two responses to a writing task.

One response is original, persuasive and well structured. It communicates powerfully with its intended audience, but contains frequent technical errors. Another is technically accurate and controlled, but less compelling, less distinctive and less effective as a whole.

If technical accuracy has substantial, separately awarded weight in the rubric, the technically controlled response may receive the higher total mark. If an assessor is asked only, “Which is the better piece of writing?”, they may choose the more persuasive response.

Neither judgement is necessarily wrong. They may simply represent different interpretations of what matters most.

This is why an organisation should not assume that a holistic ACJ process will reproduce a rank order created through analytic marking. A stable ACJ rank is meaningful only when it is stable for the construct the organisation intends to use.

Before calibrating a marked reference set, an organisation should be clear about the purpose of the proposed ruler.

Is the aim to reflect the official total mark for a particular task? To judge overall effectiveness? To support discussion of a single assessment objective? Or to provide a shared reference point for examiner training and moderation?

The answer determines whether ACJ is appropriate, how the comparison statement should be written and what claims can safely be made about the resulting ruler.

Where there is strong alignment between the intended comparative construct and the original assessment construct, ACJ may provide a powerful way to establish a reusable reference scale. Where alignment is partial, the reference set may still be highly valuable for training, moderation and professional discussion. Where alignment is weak, an organisation may be better served by retaining analytic marking or comparing distinct components separately.

From approved examples to comparative calibration

Where the construct and intended use are appropriate, an awarding organisation can select a subset of approved marked work for ACJ calibration.

This does not mean re-marking every historical script. It means curating a set of items that can function as credible reference points for a defined purpose.

The set should be deliberately bounded. It should relate to a known specification, task family or assessment component. It should represent the intended performance range. It should contain work that is suitable for comparison, sufficiently complete and appropriately anonymised. It should retain its original assessment evidence, including the mark, grade, rubric, annotations and approval record.

The next step is to define the comparative construct clearly.

A comparison statement should not ask judges to choose the “better” response in a vague or unconstrained sense. It should make visible the relevant assessment purpose and, where appropriate, the intended weight of published criteria.

For example:

“Considering the response as a whole, and giving appropriate weight to the published assessment objectives for argument, organisation and technical control, which response demonstrates stronger performance?”

This remains a holistic comparison, but it is not an unstructured one. The mark scheme continues to inform what quality means, while comparative judgement allows experts to focus on the relative strength of complete performances.

Judges then compare pairs of items against that shared statement. As comparative evidence accumulates, RM Compare develops a relative rank and estimates where each reference item sits on the scale.

The purpose is not simply to create a list from best to worst. It is to determine whether the relative item positions are sufficiently stable to support the ruler’s stated use.

RM Compare’s own reliability guidance makes clear that no single statistic should be treated as conclusive. Scale Separation Reliability is a useful indicator, but a credible stability decision should also consider the amount and pattern of comparison evidence, item-level uncertainty, the coherence of judge decisions and the appropriateness of the underlying construct.

What makes a ruler ready?

A reusable ruler is more than a stable rank.

It needs enough reference coverage to be useful. A rank made up entirely of mid-range work may be stable, but it will be of limited value if an assessor needs to distinguish work across a broad performance range. Equally, a ruler can contain a wide spread of work but still be unsuitable if individual items are ambiguous, unrepresentative or poorly aligned to the intended construct.

The quality of the item set and the quality of the judgement process therefore matter alongside rank stability.

A reference item must be suitable for its role. It should be relevant to the declared task and construct, complete enough to interpret, appropriately contextualised and approved for reuse. The judgement process must also be coherent: the panel should have the appropriate expertise, understand the comparative statement and produce sufficient evidence to support the intended use.

Finally, every ruler needs clear governance. It should have an owner, a defined scope, a version, permitted uses and a review cycle. It should be explicit about whether it is intended for examiner training, standardisation, moderation, on-demand comparison or another purpose.

This means a ruler is not declared ready simply because it has achieved a high reliability figure. It becomes ready when its evidence, coverage and governance support its claimed use.

A dual-evidence standard

When an awarding organisation brings existing marked evidence into RM Compare and calibrates an appropriate subset through ACJ, the result is not a replacement for the original assessment model.

It is a dual-evidence standard.

The original marking record remains intact. The rubric, mark, grade, task context and examiner rationale retain their authority. Comparative judgement adds a further layer of evidence: a stable relative scale, calibration history and a reusable framework for comparing future work.

This is valuable because it gives organisations more options.

They can use their mark-derived reference sets immediately for standardisation, training and moderated review. They can identify the collections most likely to benefit from comparative calibration. They can build reusable rulers selectively, beginning with complex assessment components where relative quality is difficult to establish through absolute marking alone.

Over time, this creates a more durable standards operating model.

Rather than repeatedly recreating materials and shared understanding for each assessment series, organisations can maintain a governed body of evidence. They can review and refresh reference sets. They can add new examples where coverage is weak. They can retire material when specifications change. And they can use stable comparative rulers to support consistent decisions wherever they are needed.

From marking evidence to shared standards

The proposition is not that traditional marking should be replaced.

Rubric-based marking provides important evidence of performance against defined criteria. It can support detailed feedback, criterion-level reporting and the specific rules required by many qualifications.

Comparative judgement provides something different: a way to establish and maintain a stable relative scale for suitable forms of complex work.

A standards repository brings those forms of evidence together.

It allows awarding organisations to start with the evidence they already trust, preserve its meaning and authority, and strengthen selected reference sets through ACJ where the construct and intended use justify it.

The result is not simply a better archive of scripts.

It is a governed, reusable standard: one that can support examiner standardisation, moderation, quality assurance and controlled live comparison long after the original assessment event has passed.

Start with the assessment evidence you already trust. Preserve the context that makes it meaningful. Apply comparative calibration where it is appropriate. Then turn the result into a reusable standard for consistent judgement at scale.

The final post in the series how the ruler can evolve over time.