Skip to the main content.

Hey Compono!

A coach that actually gets you.

Get 10 minutes free, then $15 a month. Cancel anytime.

Get Started ≫

2 min read

Part 4: Reliability: Why a Good Assessment Result Should Not Feel Like Luck

Part 4: Reliability: Why a Good Assessment Result Should Not Feel Like Luck
Part 4: Reliability: Why a Good Assessment Result Should Not Feel Like Luck
4:08

This is Part 4 of the 'Better Assessment Design Series:
Why Good Tests Are About More Than Questions and Scores.'
Click here for more in the series.

An assessment result should not feel like luck.

It should not depend on which version of a question someone received, which assessor happened to mark the response, or whether the wording happened to be slightly clearer for one person than another.

That is why reliability matters.

Reliability is about consistency.

If the same person took the assessment again under similar conditions, would they get a similar result?

If two assessors marked the same response, would they reach the same conclusion?

If different questions are meant to measure the same capability, do they produce consistent evidence?

Good assessment design is not just about measuring the right thing. It is about measuring it consistently enough that the result can be trusted.

Validity and reliability are different

Validity asks:

Are we measuring the right thing?

Reliability asks:

Are we measuring it consistently?

Both matter.

An assessment can be reliable but not valid. For example, it may consistently measure memory when the real goal is to measure judgement.

An assessment can also aim to measure the right thing but do so unreliably. For example, it may include realistic scenarios, but the scoring criteria may be so vague that different assessors mark responses differently.

Reliability gives consistency.

Validity gives meaning.

Good assessments need both.

Reliability + Validity = Good assessment

Why inconsistent results are dangerous

Imagine two people with the same level of understanding take the same assessment.

One passes. One fails.

Not because one understands the content better, but because one received clearer questions and the other received ambiguous questions.

That is a reliability problem.

Or imagine two assessors reviewing the same practical demonstration.

One says the person meets the standard. The other says they do not.

If there is no clear marking rubric, both assessors may believe they are right.

That is also a reliability problem.

In both cases, the assessment result becomes difficult to defend.

What weakens reliability

Reliability can be weakened by:

    • unclear questions
    • ambiguous answer options
    • too few items
    • inconsistent scoring rules
    • vague rubrics
    • assessor bias
    • inconsistent test conditions
    • trick wording
    • excessive guessing
    • questions that do not measure the same thing consistently

The higher the stakes, the more reliability matters.

A short knowledge check may not need the same level of rigour as a licensing, certification, selection or progression assessment. But any assessment used to make a meaningful decision should be stable enough to support that decision.

Practical advice

To improve reliability:

    • Use clear scoring rules.
    • Avoid vague questions.
    • Avoid questions with more than one defensible correct answer.
    • Use enough items to measure the outcome properly.
    • Train assessors where human judgement is involved.
    • Use marking rubrics for short answer, oral, portfolio, practical or scenario-based assessments.
    • Keep instructions and conditions as consistent as possible.
    • Review item performance after launch.

For multiple-choice assessments, item analysis can be very useful.

Metric

What it tells you

Item difficulty

How many people got the question right

Item discrimination

Whether stronger performers do better on the item

Distractor performance

Whether wrong answer options are working properly

Non-response rate

Whether people are skipping or misunderstanding the item

 

If almost everyone gets a question right, it may be too easy.

If almost everyone gets it wrong, it may be too hard, poorly taught, or badly written.

If stronger performers get it wrong, the item may be ambiguous.

If no one selects a distractor, that distractor is probably not doing its job.

Reliability is not about making assessments mechanical.

It is about making the evidence stable enough to trust.

A good assessment result should feel earned, not accidental.

 

Related

Part 2: Readability: If People Cannot Understand the Question, You Cannot Trust What the Answer Means

1 min read

Part 2: Readability: If People Cannot Understand the Question, You Cannot Trust What the Answer Means

This is Part 2 of the 'Better Assessment Design Series: Why Good Tests Are About More Than Questions and Scores.' Click here for more in the...

Read More
Part 3: Validity: The Most Important Question in Assessment Design

1 min read

Part 3: Validity: The Most Important Question in Assessment Design

This is Part 3 of the 'Better Assessment Design Series: Why Good Tests Are About More Than Questions and Scores.'Click here for more in the series.

Read More
Part 1 - Bloom’s Taxonomy: The Difference Between Knowing the Answer and Understanding What to Do

1 min read

Part 1 - Bloom’s Taxonomy: The Difference Between Knowing the Answer and Understanding What to Do

This is Part 1 of the 'Better Assessment Design Series: Why Good Tests Are About More Than Questions and Scores.'Click here for more in the series.

Read More