19 min read
Assessment Design Matters: Choosing the Right Question Type for the Decision You Need to Make
Rudy Crous
September 21, 2026
A good assessment is not just a collection of questions. It is a structured measurement process.
Too often, organisations focus on what they want to assess, but not enough on how they are assessing it. The question format matters. A binary question, a multiple-choice question, a Likert scale, a forced-choice item, a situational judgement test, and a work sample do not measure people in the same way. Each format creates a different type of data, a different respondent experience, and a different level of interpretability.
That is why assessment design should never start with, “What kind of test looks good?” It should start with, “What decision are we trying to make, and what evidence do we need to make that decision confidently?”
This is also consistent with modern psychometric standards. The Standards for Educational and Psychological Testing emphasise that validity is not a property of the test in isolation. Validity is about the evidence supporting the interpretation and use of test scores for a particular purpose. In other words, an assessment can be appropriate for one decision and inappropriate for another.
The first principle: match the question type to the construct
Different question types are suited to different things.
Some questions are good for testing knowledge. Some are better for measuring attitudes. Some are useful for understanding preferences. Some are stronger for observing behaviour. Others are better suited to assessing judgement, application, or decision-making in context.
The mistake many organisations make is treating all assessment formats as interchangeable. They are not.
A multiple-choice question may work well when assessing factual knowledge. A Likert scale may work well when assessing attitudes or self-perceptions. A situational judgement test may work well when assessing judgement in a specific work context. A work sample may be stronger when assessing actual capability to perform a task.
No question type is inherently “best”. The right format depends on the construct, the context, the stakes, and the interpretation you want to make from the result. This is why test development texts typically separate the construct definition, item development, test design, scoring, and interpretation processes, rather than treating question writing as a stand-alone activity.
1. Binary questions
Binary questions usually ask the respondent to choose between two options: yes/no, true/false, agree/disagree, present/absent, competent/not yet competent.

Where they work well
Binary questions are useful when the issue being measured is genuinely binary.
For example:
“Has the person completed the required training?”
“Is this document compliant with the required standard?”
“Did the candidate provide evidence of the required licence?”
“Is this answer correct or incorrect?”
In these contexts, binary questions can be efficient, clear, and easy to analyse.
Strengths
The main advantage of binary questions is simplicity. They are easy to answer, easy to score, and easy to report. They work well in compliance, eligibility, safety, screening, and audit contexts where the required outcome is clear.
They can also reduce ambiguity. When a business needs to know whether someone has completed something, passed something, or met a minimum requirement, a binary format can be appropriate.
Limitations
The limitation is that binary questions compress complexity into only two options. They can force a response where the reality is more nuanced. For example, asking someone whether they “feel confident” may not be well suited to a yes/no format because confidence exists on a continuum.
Binary questions also provide limited discrimination. They do not tell you how much, how often, to what extent, or under what conditions. This can reduce their usefulness when you need to understand capability, attitude, behaviour, or development need in more detail.
Best use
Use binary questions when the construct is genuinely binary, when the decision is simple, or when you are collecting evidence for compliance, completion, eligibility, or minimum standard purposes. Avoid using them when the underlying construct is continuous, complex, or developmental.
2. Multiple-choice questions
Multiple-choice questions ask respondents to select one correct answer, or sometimes multiple correct answers, from a set of options.
Where they work well
Multiple-choice questions are commonly used to assess knowledge, comprehension, technical understanding, policy awareness, and applied reasoning. They are particularly useful when there is a clearly correct answer.
For example:
“What is the correct procedure when X occurs?”
“Which of the following best describes the purpose of Y?”
“What is the first step in managing this safety risk?”
Strengths
Multiple-choice questions are efficient, scalable, and relatively easy to score. They can assess a broad range of content in a short time. They can also be designed to test more than rote recall, especially when the options require the respondent to apply knowledge or distinguish between plausible alternatives.
Research on multiple-choice item writing shows that item quality matters enormously. Haladyna, Downing, and Rodriguez reviewed multiple-choice item-writing guidelines and identified a taxonomy of 31 guidelines, drawing on textbooks and empirical studies. Their work reinforces a basic point: multiple-choice questions can be useful, but only when stems, options, distractors, and scoring rules are carefully designed.
Limitations
The quality of multiple-choice questions depends heavily on the quality of the distractors. Poorly written distractors make questions too easy. Ambiguous distractors make questions unfair. If the wrong answer options are obviously wrong, the question may test test-taking skill more than knowledge.
Multiple-choice questions can also encourage recognition rather than recall. A person may recognise the right answer when they see it, but not be able to produce the answer independently in the real world.
Best use
Use multiple-choice questions for knowledge, technical content, policy understanding, compliance requirements, and applied scenarios where there is a defensible correct answer. Avoid using them to assess complex interpersonal behaviour, values, motivation, or performance without careful design.
3. Multiple-response questions
Multiple-response questions require the respondent to select more than one correct option.

Where they work well
These are useful when real-world decisions require identifying several correct actions, risks, causes, or requirements.
For example:
“Which of the following are required before operating this equipment?”
“Select all factors that may contribute to this safety incident.”
“Which behaviours are consistent with the organisation’s leadership expectations?”
Strengths
Multiple-response questions can test broader understanding than single-answer multiple-choice questions. They are useful when the task is not simply to find the “best” answer, but to recognise all relevant factors.
They can also reduce guessing because the respondent must evaluate each option.
Limitations
Scoring can become more complicated. If a respondent selects three correct options and one incorrect option, should they receive partial credit? Should incorrect selections cancel out correct selections? Should all correct answers be required for full marks?
This scoring decision matters because it changes what the assessment is measuring. A harsh scoring model may punish near-competence. A generous scoring model may overstate competence.
Best use
Use multiple-response questions when the real-world task requires identifying multiple correct elements. Be clear about scoring rules before the assessment is deployed.
4. Likert-type questions
Likert-type questions ask respondents to indicate their level of agreement, frequency, confidence, importance, or perceived competence on a scale.

For example:
Strongly disagree to strongly agree.
Never to always.
Not confident to very confident.
Not important to very important.
Where they work well
Likert-type questions are useful for measuring attitudes, perceptions, self-ratings, confidence, engagement, culture, behavioural tendencies, and preferences.
The original Likert approach was developed as a method for measuring attitudes. That matters because Likert-type questions are strongest when the construct is subjective, attitudinal, perceptual, or continuous, rather than a simple right-or-wrong knowledge point.
Strengths
Likert-type items are easy for respondents to understand and can produce useful quantitative data. They allow degrees of response rather than forcing a simplistic yes/no answer. They are also useful for tracking change over time, comparing groups, and identifying trends.
When multiple items are combined into a properly designed scale, Likert-type assessments can provide more reliable measurement of broader constructs. Modern scale development guidance consistently emphasises the need to define the construct, generate items, examine reliability and validity evidence, and avoid treating a single item as a complete measure.
Limitations
Likert-type questions rely on self-report. That means they are vulnerable to response biases such as social desirability, acquiescence, central tendency, extreme responding, and impression management. Research on common method bias and survey response behaviour has repeatedly shown that the way questions are asked, labelled, ordered, and framed can influence responses.
They also require careful wording. A vague item like “I am a good leader” is much less useful than a behavioural item like “I regularly clarify priorities and expectations with my team.”
Another limitation is that Likert-type data is often over-interpreted. One item by itself rarely tells you much. The strength comes from a well-designed set of items that collectively measure a defined construct.
Four-point versus five-point scales
This is where assessment design often gets sloppy.
A four-point scale can be useful when the organisation wants a directional judgement. For example:
Strongly disagree
Disagree
Agree
Strongly agree
This removes the midpoint and requires the respondent to lean one way or the other.
A five-point scale can be useful when neutrality, uncertainty, or mixed sentiment is meaningful. For example:
Strongly disagree
Disagree
Neither agree nor disagree
Agree
Strongly agree
This is often appropriate for attitudes, perceptions, sentiment, and experience measures, because a neutral response may be valid data.
The practical rule is not that organisational assessments must always use four points and attitudinal assessments must always use five. The better rule is this: include a midpoint when neutrality, uncertainty, or mixed views are legitimate. Remove the midpoint when the assessment purpose requires a directional judgement.
Research comparing five, seven, and ten-point scales has shown that scale format can influence mean scores and comparability, even when other distributional features are similar. This reinforces the practical point that response format is not a cosmetic choice. It affects the data.
Best use
Use Likert-type questions when measuring attitudes, perceptions, confidence, frequency, preferences, or self-reported behaviours. Use multiple items per construct and be careful not to treat one isolated item as a definitive measurement.
5. Rating scales
Rating scales ask someone to evaluate performance, behaviour, competence, or quality against a defined scale.
For example:
1 = Not demonstrated
2 = Developing
3 = Competent
4 = Highly competent
Where they work well
Rating scales are commonly used in performance reviews, capability frameworks, competency assessments, leadership assessments, and manager evaluations.
Strengths
They are useful when a person’s capability exists in levels rather than pass/fail categories. They allow assessors to make developmental distinctions, which is important when identifying training needs, readiness, or progression.
When aligned to behavioural indicators, rating scales can be very useful. For example, “competent” should not just be a label. It should describe observable behaviour.
Limitations
Rating scales are vulnerable to assessor bias. Different managers may interpret the scale differently. One manager’s “competent” may be another manager’s “developing”.
Common rating errors include leniency, severity, halo effect, recency bias, and central tendency bias.
The scale is only as good as the behavioural anchors behind it. A generic rating scale without clear behavioural descriptors can produce inconsistent and unreliable data.
Best use
Use rating scales when assessing performance, capability, behaviour, and competence over levels. Always define the scale clearly and provide behavioural anchors wherever possible.
6. Behaviourally anchored rating scales
Behaviourally anchored rating scales, often called BARS, combine rating scales with specific behavioural examples.
For example:
1 = Avoids giving feedback, even when performance issues are clear
2 = Gives feedback inconsistently or only when prompted
3 = Gives clear and timely feedback linked to expectations
4 = Coaches others proactively and adapts feedback to the person and situation
Where they work well
BARS are useful for leadership assessments, competency frameworks, behavioural capability reviews, and structured performance evaluations.
Strengths
The advantage of BARS is that they make ratings more observable and less subjective. They help assessors focus on evidence rather than impressions.
This approach is not new. Smith and Kendall’s early work on behaviourally anchored rating scales focused on constructing more unambiguous anchors by using examples of expected behaviour. Their work helped establish the idea that rating scales should be anchored in observable behaviour, not vague adjectives.
Limitations
They take more work to design. Each level must be carefully written, behaviourally specific, and aligned to the construct being assessed.
They can also become too rigid if written poorly. Real work is messy, and not every behaviour will fit neatly into a predefined descriptor.
Best use
Use BARS when assessment quality matters and when multiple assessors need to make consistent judgements about behaviour or competence.
7. Forced-choice questions
Forced-choice questions require respondents to choose between two or more options, often where both options may be socially desirable.
For example:
“Which statement is more like you?”
A. I enjoy generating new ideas.
B. I enjoy making sure plans are completed properly.
Where they work well
Forced-choice formats are often used in personality, preference, values, and behavioural style assessments.
Strengths
Forced-choice questions can reduce some forms of response bias because respondents cannot simply rate themselves highly on every positive trait. They must make trade-offs.
This can be useful when measuring relative preferences. For example, a person may be capable of both pioneering and maintaining work, but the forced-choice item helps identify which type of work they are more naturally drawn to.
Limitations
Forced-choice formats can be frustrating for respondents if both options feel true or neither option feels true.
There is also a psychometric issue. Classical scoring of multidimensional forced-choice questionnaires can produce ipsative data. That means scores are relative within the person and can make comparisons between people problematic. More modern item response theory approaches can help address this, but only when the assessment has been designed and scored appropriately.
Best use
Use forced-choice questions when measuring relative preferences, styles, or tendencies. Be cautious when using them for high-stakes selection unless the scoring model and validation evidence support that use.
8. Ranking questions
Ranking questions ask respondents to place options in order of preference, importance, priority, or effectiveness.
Where they work well
Ranking questions can be useful when you want to understand priorities or trade-offs.
For example:
“Rank the following leadership behaviours from most important to least important in your role.”
“Rank these benefits in order of value to you.”
“Rank these actions from most to least appropriate.”
Strengths
Ranking questions force prioritisation. This can be useful because people often rate everything as important when using standard rating scales.
Ranking can reveal what matters most when choices must be made.
Limitations
Ranking questions become difficult when there are too many items. Respondents can usually rank three to five options reasonably well, but longer lists become cognitively demanding and less reliable.
Ranking also tells you relative order, not absolute strength. If someone ranks “career development” above “pay”, it does not mean pay is unimportant. It only means career development was ranked higher in that specific set.
Best use
Use ranking questions when trade-offs matter. Keep the number of options small. Avoid over-interpreting the distance between ranked items.
9. Semantic differential questions
Semantic differential questions ask respondents to rate something between two opposing adjectives.
For example:
Clear 1 2 3 4 5 Confusing
Supportive 1 2 3 4 5 Unsupportive
Innovative 1 2 3 4 5 Traditional
Where they work well
These are useful for measuring perceptions, brand attributes, culture, user experience, and emotional reactions.
Strengths
Semantic differential items can capture nuanced perceptions quickly. They are often easy to complete and can be visually intuitive.
They can be useful when measuring how people experience an organisation, product, leader, or process.
Limitations
The quality depends on the choice of adjective pairs. Poorly chosen opposites can confuse respondents or introduce bias.
They may also feel simplistic if the construct is complex.
Best use
Use semantic differential questions when measuring perception, sentiment, or comparative impressions. They work best when the opposing descriptors are clear and meaningful.
10. Open-ended questions
Open-ended questions allow respondents to answer in their own words.
For example:
“What is the biggest barrier to performing your role effectively?”
“What would improve the onboarding experience?”
“Describe how you would handle this situation.”
Where they work well
Open-ended questions are useful for exploration, diagnosis, feedback, qualitative insight, and understanding context.
Strengths
They allow people to express things the assessment designer may not have anticipated. This is particularly valuable when exploring culture, engagement, barriers, risk, or improvement opportunities.
They can reveal nuance that closed questions miss.
Limitations
Open-ended responses are harder to score consistently. They require interpretation, coding, or qualitative analysis. This can introduce subjectivity.
They are also more time-consuming for respondents and assessors.
In high-stakes assessment, open-ended questions need clear marking criteria, rubrics, and ideally multiple markers or moderation processes.
Best use
Use open-ended questions when you need insight, explanation, examples, or context. Do not rely on them alone for scalable, objective scoring unless you have a structured rubric.
11. Short-answer and constructed-response questions
Short-answer questions require respondents to generate an answer rather than recognise one.
For example:
“List three risks associated with this process.”
“What is the purpose of this control?”
“Explain the first action you would take and why.”
Where they work well
These are useful for assessing recall, comprehension, explanation, reasoning, and applied knowledge.
Strengths
They reduce guessing compared with multiple-choice questions. They also provide better evidence that the person can produce the answer independently.
Constructed responses can be particularly useful when you need to assess reasoning, not just the final answer.
Limitations
They take longer to mark and require clear scoring guides. Responses may vary in wording even when they are conceptually correct.
Automated scoring can help, but it needs careful validation.
Best use
Use short-answer questions when recall, explanation, or reasoning matters. Use scoring rubrics to maintain consistency.
12. Scenario-based questions
Scenario-based questions present a workplace situation and ask the respondent what they would do, what risk they see, or what decision they would make.
Where they work well
They are useful for applied knowledge, judgement, policy interpretation, safety, leadership, customer service, and problem-solving.
Strengths
Scenario-based questions improve contextual relevance. They help move beyond abstract knowledge and into applied decision-making.
They often have stronger face validity because respondents can see the link to the job.
Limitations
The scenario must be realistic, fair, and aligned to the role. If the scenario is too specific, too obscure, or based on assumed experience, it may disadvantage some respondents.
There is also a risk that the question measures familiarity with the scenario rather than the underlying capability.
Best use
Use scenario-based questions when you want to assess whether someone can apply knowledge in context. Ensure the scenario reflects realistic work demands and is not overly dependent on niche experience.
13. Situational judgement tests
Situational judgement tests, or SJTs, present realistic workplace scenarios and ask respondents to choose, rate, or rank the most appropriate response.
For example:
“You notice a team member repeatedly ignoring a safety procedure. What would you do?”
The respondent may be asked to select the best response, rank several responses, or rate the effectiveness of each option.
Where they work well
SJTs are often used in recruitment, leadership assessment, graduate selection, safety roles, customer service, and professional judgement contexts.
Strengths
SJTs can have strong face validity because they look and feel relevant to the job. Candidates and employees can often see the connection between the assessment and the work.
They can assess judgement, prioritisation, interpersonal decision-making, values in action, and applied reasoning. They are also scalable compared with interviews or simulations. Whetzel and McDaniel define SJTs as assessments that present work-related situations and ask respondents to evaluate likely or effective courses of action.
Meta-analytic research has found that SJTs can predict job performance, particularly when they are designed around job-relevant constructs. Christian, Edwards, and Bradley found that SJTs commonly assess leadership and interpersonal skills, and that matching the construct being assessed to the performance criterion improves validity.
Limitations
The major limitation is interpretability. An SJT response is often only valid in relation to the specific situation presented, unless the assessment has been deliberately designed to measure a broader construct.
This is a common overreach in assessment design. A situational judgement test may tell you how someone responds to a particular scenario, under particular assumptions, with particular options. It does not automatically tell you how they will behave across all contexts.
This concern is echoed in the research literature. Lievens and Motowidlo argue that SJTs are often treated as contextualised simulations, but many SJTs lack clear construct definition. They note that despite SJT popularity and evidence of validity, there has often been insufficient theory explaining exactly why SJTs work and what they measure.
SJTs can also be vulnerable to coaching, cultural assumptions, and item-writing bias. The “best” answer must be defensible, job-relevant, and ideally based on expert judgement or criterion-related evidence.
Another challenge is that SJTs can appear sophisticated while being psychometrically weak. Good face validity does not automatically mean strong validity.
Best use
Use SJTs when judgement in context matters. They are best used as part of a broader assessment battery, not as a standalone measure of broad capability or personality. Be careful to interpret results at the right level: “judgement in these kinds of situations”, not “global competence”.
14. Work samples
Work samples ask respondents to perform a task that closely resembles the actual work.
For example:
A sales candidate prepares a discovery call plan.
A customer service candidate responds to a difficult customer email.
A manager reviews a team performance scenario and prepares a coaching response.
A technical candidate completes a job-relevant task.
Where they work well
Work samples are useful when you need evidence of actual capability. They are especially valuable when the role requires practical output, applied judgement, technical skill, communication, or problem-solving.
Strengths
Work samples often have strong job relevance. They assess what a person can actually do, not just what they say they can do.
They can also improve perceived fairness when designed well because candidates are assessed against the work itself, not just proxies such as CV history or interview performance. Research on applicant reactions has found that interviews and work samples tend to be perceived more favourably than many other selection methods, likely because applicants can see the relevance to the role.
Limitations
Work samples take more time to design, administer, and score. They must be realistic without being exploitative. They also need structured scoring rubrics to avoid subjective judgements.
A work sample may assess current capability but not necessarily potential, learning agility, or motivation.
Best use
Use work samples when actual job performance evidence matters. They are particularly useful in selection, promotion, technical roles, leadership development, and high-stakes capability assessment.
15. Simulations and assessment centres
Simulations place respondents in complex, realistic exercises such as role plays, group tasks, presentations, inbox exercises, case studies, or leadership scenarios.
Where they work well
They are useful for leadership assessment, executive selection, graduate programs, high-potential identification, and complex role readiness.
Strengths
Simulations can assess behaviour in action. They allow observation of communication, judgement, prioritisation, influencing, collaboration, and leadership under realistic conditions.
They can provide rich evidence, especially when multiple assessors and multiple exercises are used.
Assessment centre research generally supports their usefulness, but also shows that design matters. Lievens’ review of assessment centre construct validity concluded that careful design and high inter-rater reliability are necessary, but not always sufficient, to establish construct-related validity.
Limitations
They are expensive and resource-intensive. They require trained assessors, structured scoring, calibration, and careful design.
They can also create performance anxiety and may favour people who are comfortable in artificial assessment environments.
Best use
Use simulations when the decision is high-stakes and the role requires complex behavioural capability. They are strongest when combined with clear competency frameworks and trained assessors.
16. Portfolio and evidence-based assessment
Portfolio assessment involves reviewing collected evidence of competence over time.
For example:
Completed work products
Supervisor observations
Training records
Assessment results
Reflections
Credentials
Practical demonstrations
Customer or stakeholder evidence
Where they work well
Portfolio assessment is useful in competency-based environments, professional development, licensing, compliance, and role progression.
Strengths
The strength of portfolio assessment is that it can capture competence over time rather than relying on a single test event. It allows multiple forms of evidence to be considered.
This is particularly important when competence is not just about knowing something, but being able to apply it repeatedly and reliably in real contexts.
Limitations
Portfolio assessment can become inconsistent if the evidence requirements are unclear. It also requires governance. What counts as evidence? Who validates it? How current does it need to be? How is quality assured?
Without structure, portfolio assessment can become a collection of documents rather than a valid assessment process.
Best use
Use portfolio assessment when competence needs to be demonstrated across time, roles, levels, or real-world contexts. Define evidence standards clearly.
17. Adaptive assessments
Adaptive assessments adjust the difficulty or content of questions based on the respondent’s previous answers.
Where they work well
They are useful in knowledge testing, skills diagnostics, placement testing, and large-scale assessments where efficiency and precision matter.
Strengths
Adaptive testing can reduce assessment time and improve measurement precision. People are not forced through large numbers of questions that are too easy or too hard.
It can also provide more tailored diagnostic information.
Limitations
Adaptive assessments require strong item banks, item calibration, and a robust measurement model. They are more complex to build and maintain.
They can also be harder to explain to respondents because not everyone receives the same questions.
Best use
Use adaptive assessments when you have a mature item bank, enough data, and a clear need for efficient, precise measurement.
The practical design question: what are you trying to infer?
The real issue in assessment design is not just the question format. It is the inference you want to make from the result.
A binary question may support a simple compliance inference.
A multiple-choice question may support a knowledge inference.
A Likert scale may support a perception or attitude inference.
A rating scale may support a capability-level inference.
An SJT may support a contextual judgement inference.
A work sample may support a practical performance inference.
A simulation may support a behavioural readiness inference.
A portfolio may support an evidence-over-time inference.
Good assessment design keeps the inference honest.
The danger is when organisations overclaim. For example:
A person passed a knowledge test, so we assume they can perform the task.
A person selected the best answer in an SJT, so we assume they have strong leadership capability.
A person rated themselves highly, so we assume they are competent.
A person performed well in one work sample, so we assume they will perform well in every part of the role.
Each of these may be partly true, but each requires caution.
Recent personnel selection research reinforces this caution. Sackett, Zhang, Berry, and Lievens revisited meta-analytic validity estimates in personnel selection and concluded that many previous estimates had been overcorrected and overstated, although many selection methods still remain useful. The lesson for organisations is not to abandon assessment. The lesson is to stop overclaiming what any single assessment method can prove.
A useful way to choose assessment formats
When designing an assessment, ask five questions.
1. What is the construct?
Are you measuring knowledge, skill, attitude, behaviour, judgement, preference, performance, competence, or potential?
2. What type of evidence is required?
Do you need self-report, observed behaviour, correct answers, applied reasoning, work output, or evidence over time?
3. What decision will be made?
Is this for development, selection, compliance, certification, promotion, diagnosis, or workforce planning?
The higher the stakes, the stronger the evidence needs to be.
4. How will the result be interpreted?
Will you interpret the result at item level, scale level, competency level, role level, or organisational level?
This matters because not every question type supports broad interpretation.
5. What are the risks of getting it wrong?
Could the assessment unfairly exclude someone? Could it falsely certify competence? Could it create misleading development data? Could it give leaders confidence in results that are not valid?
Practical summary: when to use each question type
|
Question type |
Best for |
Be careful of |
|
Binary |
Compliance, completion, eligibility, true/false knowledge |
Oversimplifying complex constructs |
|
Multiple choice |
Knowledge, technical content, applied understanding |
Weak distractors, guessing, recognition bias |
|
Multiple response |
Identifying multiple correct risks, actions, or requirements |
Complex scoring rules |
|
Likert-type |
Attitudes, perceptions, confidence, frequency, preferences |
Self-report bias and overinterpreting single items |
|
Rating scales |
Capability, performance, competence levels |
Rater bias and vague scale definitions |
|
Behaviourally anchored scales |
Behavioural assessment and competency frameworks |
Design effort and rigidity |
|
Forced choice |
Preferences, style, relative tendencies |
Frustration, ipsative scoring, comparability limits |
|
Ranking |
Priorities and trade-offs |
Cognitive load and overinterpreting rank distance |
|
Semantic differential |
Perceptions, brand, culture, experience |
Poor adjective pairs |
|
Open-ended |
Qualitative insight, explanation, context |
Subjective scoring and analysis effort |
|
Short-answer |
Recall, explanation, applied reasoning |
Marking consistency |
|
Scenario-based |
Application of knowledge in context |
Scenario specificity and assumed experience |
|
SJTs |
Contextual judgement and decision-making |
Overclaiming broad capability from specific scenarios |
|
Work samples |
Actual job-relevant performance |
Scoring consistency and design effort |
|
Simulations |
Complex behavioural capability |
Cost, assessor training, artificiality |
|
Portfolio evidence |
Competence over time |
Evidence quality and governance |
|
Adaptive testing |
Efficient, precise knowledge or skills diagnostics |
Item bank and psychometric complexity |
The final point: assessment design is a validity decision
Assessment design is not a formatting exercise. It is a validity decision.
Every question type carries assumptions. Every scale shapes the data. Every response format influences how people answer. Every scoring model affects what the result means.
That is why the best assessment systems do not simply ask more questions. They ask better questions, in the right format, for the right purpose, with a clear understanding of what can and cannot be inferred from the result.
In organisational settings, this matters enormously. Assessments are often used to make decisions about hiring, promotion, capability, compliance, development, leadership, culture, and workforce risk. Poor assessment design creates poor evidence. Poor evidence creates poor decisions.
The goal is not to make assessments look sophisticated. The goal is to make them useful, fair, valid, reliable, and interpretable.
A well-designed assessment should help an organisation answer three questions with confidence:
What are we trying to measure?
What evidence do we have?
What decisions can we responsibly make from that evidence?
That is the discipline of good assessment design.
Rudy Crous is a corporate psychologist and the co-founder and CEO of Compono, a people and culture platform that brings hiring, engagement, learning and credentialling together in one place, adding behavioural science to the processes most HR tools only track.
Reference list
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. American Educational Research Association.
- Brown, A., & Maydeu-Olivares, A. (2011). Item response modeling of forced-choice questionnaires. Educational and Psychological Measurement, 71(3), 460–502.
- Christian, M. S., Edwards, B. D., & Bradley, J. C. (2010). Situational judgment tests: Constructs assessed and a meta-analysis of their criterion-related validities. Personnel Psychology, 63(1), 83–117. Verified via Wiley and Cambridge reference records.
- Crocker, L., & Algina, J. (1986). Introduction to Classical and Modern Test Theory. Holt, Rinehart and Winston.
- Dawes, J. (2008). Do data characteristics change according to the number of scale points used? An experiment using 5-point, 7-point and 10-point scales. International Journal of Market Research, 51(1), 1–20.
- DeVellis, R. F., & Thorpe, C. T. (2021). Scale Development: Theory and Applications (5th ed.). SAGE.
- Haladyna, T. M., Downing, S. M., & Rodriguez, M. C. (2002). A review of multiple-choice item-writing guidelines for classroom assessment. Applied Measurement in Education, 15(3), 309–333.
- Hausknecht, J. P., Day, D. V., & Thomas, S. C. (2004). Applicant reactions to selection procedures: An updated model and meta-analysis. Personnel Psychology, 57(3), 639–683.
- Lane, S., Raymond, M. R., & Haladyna, T. M. (Eds.). (2016). Handbook of Test Development (2nd ed.). Routledge.
- Lievens, F. (2009). Assessment centers: A tale about dimensions, exercises, and dancing bears. European Journal of Work and Organizational Psychology, 18(1), 102–121.
- Lievens, F., & Motowidlo, S. J. (2016). Situational judgment tests: From measures of situational judgment to measures of general domain knowledge. Industrial and Organizational Psychology, 9(1), 3–22.
- McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P. (2001). Use of situational judgment tests to predict job performance: A clarification of the literature. Journal of Applied Psychology, 86(4), 730–740.
- Podsakoff, P. M., MacKenzie, S. B., Lee, J.-Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research: A critical review of the literature and recommended remedies. Journal of Applied Psychology, 88(5), 879–903.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149–155.
- Whetzel, D. L., & McDaniel, M. A. (2009). Situational judgment tests: An overview of current research. Human Resource Management Review, 19(3), 188–202.

Better hiring decisions, before the interview
Compono Hire helps you predict job-fit and team-fit using behavioural science, so you can shortlist with confidence.
Request a demoBuilt for mid-market hiring teams.

Practise your next tough conversation
Voice-first coaching that adapts to your personality. Get actionable steps you can take this week.
Start freeBuilt by Compono. Not therapy — practical behaviour change.

.webp)
