Skip to the main content.

Hey Compono!

A coach that actually gets you.

Get 10 minutes free, then $15 a month. Cancel anytime.

Get Started ≫

3 min read

Part 3: Validity: The Most Important Question in Assessment Design

Part 3: Validity: The Most Important Question in Assessment Design
Part 3: Validity: The Most Important Question in Assessment Design
5:19

This is Part 3 of the 'Better Assessment Design Series:
Why Good Tests Are About More Than Questions and Scores.'
Click here for more in the series.

 

 

Assessments can look credible long before they are valid.

They may have professional wording, structured answer options, scoring rules, pass marks and polished reports. But none of that proves the assessment is measuring what it claims to measure.

That is why validity sits at the centre of good assessment design.

Validity asks whether the result supports the decision or interpretation being made from it.

In plain language:

Are we measuring what we think we are measuring?

Good assessment design is not just about producing a score. It is about making sure the score means what we think it means.

Why validity matters

Assessments are rarely neutral.

They often lead to decisions.

Someone receiving a certification

Someone passes or fails. Someone progresses or does not. Someone receives a credential. Someone is judged ready, not ready, competent, not competent, suitable, unsuitable, proficient, or in need of further support.

That means the quality of the assessment affects the quality of the decision.

If the assessment is weak, the decision is weak.

This is why validity matters so much.

A test can look professional and still produce misleading evidence. It can have clear questions, polished reporting, a neat score and a pass mark, but still not measure what it claims to measure.

The stop sign example

Imagine the goal is to assess whether someone understands road safety.

This question is limited:

What does a stop sign mean?

Stop Sign

It measures recall.

That may be useful. But if the assessment claims to measure safe road judgement, it is not enough.

A stronger question would be:

You are approaching a stop sign near a school. It is raining, children are nearby, and visibility is poor. What should you do and why?

This gives better evidence.

It tests whether the person can connect the rule to risk, context, judgement and safe behaviour.

That is closer to the intended capability.

The issue is not whether the first question is wrong.

The issue is what we claim from the answer.

What assessments measure by accident

Many assessments claim to measure understanding, readiness or competence, but actually measure something else.

They may measure:

    • memory
    • reading ability
    • test-taking skill
    • confidence
    • guessing ability
    • familiarity with the question style
    • prior exposure to similar questions

This does not make the assessment useless. But it does mean we need to be careful.

If an assessment only measures recall, we should not use it as strong evidence of applied competence.

If an assessment uses unnecessarily complex wording, we should not assume the result only reflects subject knowledge.

If an assessment asks narrow questions, we should not claim broad capability.

Validity is partly about humility. It asks us not to overstate what the result can prove.

Different types of validity

There are several forms of validity that matter.

  1. Content validity asks whether the assessment covers the right material. If a syllabus, standard, framework or learning program includes ten important areas, but the assessment only tests three, the assessment is too narrow.

  2. Construct validity asks whether the assessment measures the underlying capability it claims to measure. If we say we are measuring judgement, are we really measuring judgement, or are we measuring memory and reading ability?

  3. Face validity asks whether the assessment appears relevant and credible to the people taking it and to stakeholders. This matters because people engage more seriously with assessments that feel relevant. But face validity alone is not enough. A scenario can look realistic and still be poorly designed.

  4. Criterion-related validity asks whether assessment results relate to meaningful outcomes. For example, do higher scores relate to better performance, fewer errors, safer behaviour, stronger learning outcomes or better decisions?

Practical advice: design validity in from the start

Validity should be considered before the questions are written.

A simple blueprint can help.

Intended outcome

Learning objective

Level of thinking

Question type

Evidence required

Understand road safety risk

Explain why stopping is required near schools

Understanding

Scenario question

Explains risk to pedestrians

Apply stopping rules

Choose the safest action at an intersection

Applying

Multiple-choice scenario

Selects safe action

Make safe decisions

Prioritise hazards in changing conditions

Analysing and evaluating

Judgement item

Chooses safest option and explains why

Every question should map back to:

    • the learning outcome, competency or standard
    • the learning objective
    • the level of thinking
    • the evidence required
    • the decision being made from the result

If a question cannot be mapped, it probably should not be in the assessment.

Validity is not created by good intentions.

It is created through careful design, clear mapping, strong item writing, appropriate scoring and ongoing review.

A valid assessment gives us confidence that the result means what we think it means.

And that confidence is the whole point. 

 

Related

We've tested whether new drivers can see danger for 20 years. Here's what the evidence says.

1 min read

We've tested whether new drivers can see danger for 20 years. Here's what the evidence says.

I have two boys. Five and two. Driving is a long way off for both of them, and yet here I am already worrying about it. Not the parallel parking or...

Read More
Part 2: Readability: If People Cannot Understand the Question, You Cannot Trust What the Answer Means

1 min read

Part 2: Readability: If People Cannot Understand the Question, You Cannot Trust What the Answer Means

This is Part 2 of the 'Better Assessment Design Series: Why Good Tests Are About More Than Questions and Scores.' Click here for more in the...

Read More
Part 1 - Bloom’s Taxonomy: The Difference Between Knowing the Answer and Understanding What to Do

1 min read

Part 1 - Bloom’s Taxonomy: The Difference Between Knowing the Answer and Understanding What to Do

This is Part 1 of the 'Better Assessment Design Series: Why Good Tests Are About More Than Questions and Scores.'Click here for more in the series.

Read More