Steady State, Not Static: Maintaining Assessment Standards When Questions Change (3/3)

In the first two articles in this series, we explored why assessment organisations need standards repositories rather than static item libraries, and how existing marked evidence can become governed reference sets and, where appropriate, reusable ACJ rulers.

For awarding organisations, the next question is unavoidable.

Exam boards do not set the same question every year.

They refresh topics, prompts, source materials and contexts. They need to protect assessment integrity, avoid predictability and ensure that qualifications remain current. Yet they also need to maintain a consistent interpretation of quality across assessment series.

So can an organisation really use a shared standard when every year produces different candidate work?

Yes, but only if a shared standard is understood as something living, not fixed.

A steady-state approach does not mean keeping the same historic examples forever. It means preserving a stable core of shared quality, then deliberately testing, extending and governing that standard as new assessment tasks are introduced.

Steady state is not a frozen ruler. It is a stable core of shared standards, continuously tested and deliberately maintained.

A living master ruler

The most useful model is a living master ruler.

This is not simply a collection of historic scripts held in a repository. It is a governed reference scale for a clearly defined construct and task family. It retains a core set of approved reference items, while allowing new evidence to be added, restricted or retired as assessment conditions change.

For example, an awarding organisation might define a master ruler around:

Persuasive extended writing for a specified audience and purpose, assessed against the published objectives for the current qualification.

That is a meaningful and bounded claim. It is much more defensible than trying to create a universal ruler for “writing quality.”

The master ruler contains selected anchor items. These are approved reference points that are well understood, spread across the intended performance range and retained because they provide continuity between assessment series. They remain connected to the information that explains their status: task context, specification version, original assessment outcome, approval history and permitted use.

When the organisation introduces a new annual task, it does not throw away the existing standard. Nor does it assume that new work will automatically fit the old scale.

It uses bridge calibration to investigate the relationship.

Same construct, different evidence

Two tasks can be designed to assess the same construct and still elicit different evidence from candidates.

Consider two persuasive-writing assessments. Both may assess effectiveness of communication for audience and purpose. Both may use the same assessment objectives. Both may require candidates to select ideas, organise a response and control language appropriately.

But the task itself can change what candidates are able to demonstrate.

One question may invite an article about an issue familiar to most candidates. Another may require a speech on an unfamiliar or more abstract topic. One may make it easy to adopt a clear position and structure an argument. Another may place greater demands on adapting tone to audience. One may favour candidates with strong topical knowledge, while another rewards imaginative engagement with a scenario.

The intended construct may be shared. The response evidence may not be.

This is why awarding organisations cannot assume that a ruler built from one year’s candidate work can be applied unchanged to a different task simply because both papers were designed to assess the same objective.

The assessment field has long recognised this issue. Where different forms of a test are intended to support comparable outcomes, organisations use linking and equating approaches to distinguish differences in candidate performance from unintended differences in task difficulty or form demand. Common or anchor items are a well-established way of providing a shared reference frame across different test forms.

For comparative judgement, the equivalent principle is straightforward:

Do not assume that a new task belongs on an established standard. Test the relationship through evidence.

Anchor items create continuity

A living master ruler depends on a carefully maintained set of anchor items.

In conventional test equating, common items provide a bridge between different assessment forms. They help organisations place results on a common scale by giving them shared evidence across otherwise different tests.onlinelibrary.wiley+1

Anchor items play a comparable role in a comparative-judgement model.

They are not simply old examples kept because they are familiar. They are deliberately selected and governed reference points. Each has a known place within the master ruler and a clear rationale for continued use.

A good anchor should be relevant to the defined construct, suitable for the task family and useful as a point of comparison. Collectively, the anchor set should span the range of performance the ruler is intended to represent. It should include items around important grade, pass/fail or category boundaries, not only clear examples at the extremes.

This does not mean that an anchor’s position can never change. It means that the organisation has enough evidence to understand its role in the shared standard and can monitor it as new data is introduced.

When a new task is introduced, selected responses can be compared with these anchors. The aim is to establish whether the new work can be located credibly against the existing reference scale.

Bridging a new assessment series

Bridge calibration is the practical process that connects a new assessment task to the existing master ruler.

The organisation begins by selecting a representative sample of new candidate responses. This sample should span the expected range of quality and include work around the points where important decisions are made. It should reflect legitimate differences in how candidates have approached the task, not simply the most familiar or convenient responses.

Experts then make structured comparative judgements.

New-task responses are compared with one another, so their relative positions can emerge. They are also compared with selected anchor items from the master ruler. The anchor items provide a stable frame of reference, while the new-task evidence reveals how the new assessment is behaving.

The comparison statement needs to focus on the construct the organisation intends to maintain, not on irrelevant differences in topic, format or surface style.

For a persuasive-writing task, it might ask:

“Considering the overall effectiveness with which each response fulfils its specified persuasive purpose for its intended audience, which response demonstrates stronger performance against the published assessment objectives?”

This helps judges compare responses across task contexts while remaining focused on the qualities the qualification is designed to assess.

The resulting calibration evidence can then be reviewed. Does the new task appear to elicit the intended construct? Do the new responses locate coherently against anchor items? Are key boundaries behaving as expected? Is the new task exposing gaps in the existing reference set? Are some differences caused by the task itself rather than differences in candidate quality?

These are not peripheral technical questions. They are the evidence base for deciding whether the shared standard can be maintained across the new assessment series.

Link, limit or separate

The purpose of bridge calibration is not to force every new task onto one common ruler.

Its purpose is to allow awarding organisations to make an explicit, evidence-led decision about the relationship between new task-specific evidence and the established standard.

Sometimes, the evidence will support a clear link. The new task has elicited the intended construct, judges can apply the comparison statement consistently and new work locates coherently against the anchors. Approved new responses can then be added to the master ruler, making the standard more current and representative without losing continuity.

Sometimes, the relationship will be credible but limited. The new task may behave differently at one part of the performance range, or it may reveal a type of response that the existing anchor set does not represent well. In this case, the organisation can retain a task-specific layer, gather more evidence or restrict use while the relationship is better understood.

And sometimes, the task will be too different. It may introduce a materially different construct, non-compensatory requirements or response conditions that make a single shared ruler misleading. In that situation, the responsible decision is to maintain a separate task-specific ruler or reference set.

A standards repository should support all three outcomes. Its role is not to conceal important differences in assessment evidence. It is to make those differences visible, governed and usable.

A ruler needs a lifecycle

A master ruler should not be treated as complete once it has been created.

Assessment tasks change. Specifications are revised. New candidate responses reveal different ways of demonstrating quality. Some historic examples become less representative. Others remain valuable as long-term anchors. The ruler needs to evolve in a controlled way.

This is why versioning matters.

Each ruler version should state the construct and task family it represents, the qualification and specification to which it relates, the anchor items it retains and the new items added through bridge calibration. It should make clear which evidence supported the decision to link a new task, where limits apply and which uses have been approved.

It should also have an owner, a review cycle and a clear process for retiring or recalibrating items when the evidence requires it.

This is not simply administrative housekeeping. It is what makes a shared standard defensible over time.

An organisation should be able to show not only the items it uses as references, but also the evidence that explains why those items remain appropriate for a particular purpose.

The place of RM Compare

RM Compare can provide the comparative layer within this operating model.

It can help awarding organisations retain and govern anchor items, curate new task-specific evidence and run focused bridge-calibration activity. It can make the relationship between new work and an established reference scale visible, while preserving the original assessment context and approval record.

Most importantly, it can help ensure that the learning created through one assessment series is not lost when the next begins.

AOs will always create new questions. Candidate work will always change. Standards will always require expert interpretation.

The purpose of a steady-state model is not to eliminate these realities.

It is to make them manageable.

Instead of rebuilding shared understanding from scratch, organisations can carry forward a stable core of evidence. Instead of assuming that annual tasks are equivalent, they can test how new evidence relates to the established standard. Instead of allowing standards knowledge to remain locked in documents and individual expertise, they can maintain it as a living, governed organisational asset.

Monitoring cohort performance over time

A living master ruler can do more than support consistent judgements within an assessment series. When new task-specific work is successfully linked back to the same reference scale, it can also provide evidence about changes in cohort performance over time.

This does not mean that every shift in the scale should be read as genuine improvement or decline. New tasks, different candidate populations and changing assessment conditions all need to be considered. But where construct continuity and bridge-calibration evidence are strong, a linked ruler can show whether a cohort’s work is locating differently against an established standard.

For awarding organisations, this creates an additional evidence stream for reviewing standards, understanding task effects and monitoring the health of an assessment over time. It complements established statistical and judgemental approaches to awarding; it does not replace them.

A linked master ruler can provide an additional longitudinal evidence stream for standards maintenance, task review and cohort-performance analysis. It should complement, not replace, formal awarding and comparability evidence.