← Knowledge base
πŸ”˜

Forced choice and Likert scales

Whether a test asks you to pick A or B, or to rate a statement from one to five, decides what kind of answer it can give you β€” and one of the two formats creates the categories it later reports as findings.

How we measure

Two ways of asking the same question.

At a party, do you (a) talk to many people including strangers, or (b) talk to a few people you know?

"I start conversations with people I don't know." Strongly disagree β€” disagree β€” neutral β€” agree β€” strongly agree.

The first is forced choice. The second is a Likert scale, after Rensis Likert, who proposed the format in 1932. They look like stylistic alternatives. They are not. The format decides what kind of result the test is capable of producing, before a single question has been written.

The short version: forced choice throws away the degree, and a system that has thrown away the degree can only report a side. If you have ever wondered why type tests produce types and trait tests produce percentages, this is most of the answer. The categories were not discovered in the data. They were built into the answer sheet.

What each format keeps

Likert keeps distance. Someone who strongly agrees and someone who mildly agrees have said different things, and the scoring preserves the difference. Twenty items summed give a score with a lot of possible values, which can then be placed against a comparison group β€” the percentiles and norms machinery. It also makes the middle sayable: neutral is a real answer, and for a genuinely middling respondent it is the accurate one.

Forced choice keeps direction. You picked a. The test records a. Whether you picked it at 51 per cent or 99 per cent is not recorded, because it was not asked. Twenty items produce a count out of twenty, and every respondent's tally lands on one side or the other of ten.

That has one genuine advantage, which is why the format persists. Forced choice is harder to fake in a direction, because both options are usually made equally attractive. If you are answering to look good to an employer, a five-point scale lets you agree strongly with everything flattering; an A-or-B item between two desirable statements does not. For high-stakes selection that matters, and it is why well-designed occupational instruments still use the format.

The cost is that a person with no real preference is required to invent one. Twenty invented preferences make a type.

The ipsative problem

Forced choice carries a second and less obvious consequence, named by Cattell in 1944 and worked out in detail by Hicks.

When each item makes you choose between two traits, the traits are competing for a fixed budget. Score high on one and you have necessarily scored lower on another, because the same answer did both jobs. This is ipsative scoring: the result describes the ordering of traits within you, not your standing among other people.

The consequence is blunt and routinely ignored. Ipsative scores cannot be compared between people. "I'm high on dominance and you're low" is not a statement the data supports, because both of you spent the same fixed budget. You can say dominance is your highest; you cannot say it is higher than someone else's. Two candidates with identical profiles may behave nothing alike, and two with opposite profiles may be identical in absolute terms.

This is why a DISC profile is a shape rather than a height, and why the sensible use of one is a conversation about your own ordering rather than a comparison across a team. Brown and Maydeu-Olivares showed the problem is technically solvable β€” item-response models can recover normative scores from forced-choice data β€” but the solution requires a specific design and a calibration sample, and almost nothing you can take for free on the internet has either.

Where Likert goes wrong

The format that keeps more information also collects more noise, and it is not the obviously better choice.

The middle becomes a hiding place. Neutral is used for genuinely middling, for don't-know, for depends, for don't-want-to-say and for not-reading-carefully. Five distinct states collapse into one response, which is why some instruments remove the midpoint and force a lean β€” reintroducing the forced-choice problem in miniature.

Style contaminates content. Some people avoid the ends of any scale and some live at them. That tendency is stable within a person and it shifts their scores on every trait at once, which is the subject of response styles.

The anchor is your own reference group. Agreeing that you are talkative means talkative compared to whom. People answer against the people around them, so the same behaviour produces different ratings in different environments β€” the reason cross-cultural comparisons of raw Likert means are so often wrong.

Item wording does more work than it looks like. I start conversations with strangers and I find it easy to start conversations with strangers measure different things. One is behaviour, the other is perceived difficulty, and a lot of people answer differently to the two.

How to read a result once you know the format

Look at the answer format before you look at the result. It takes one screen and it tells you what the output can mean.

If it was A or B, the output is a shape and a direction. Do not read it as a quantity, do not compare your profile with someone else's as though both were measured on a common ruler, and expect the labels to be unstable β€” a fifty-one per cent preference recorded as a side is exactly the situation test–retest reliability fails in.

If it was one to five, the output is a score and it means something only against a comparison group. Ask what the group was. A percentile computed against everyone who happened to take a test on a website is not the same object as one computed against a stratified sample.

If the result is a category and the format was a scale, someone drew a line, and the line is the least reliable part of what you were told. Find the underlying number if the page shows one.

Our own tests use both formats, and the mapping is not accidental: the Big Five is Likert and reports percentages, the DISC profile is closer to ipsative and reports a shape, and the MBTI-style test reports percentages beside the four letters specifically so that the discarded degree is visible somewhere. Where a result of ours gives you a label and no number, the label is carrying more weight than the format underneath it can support, and you should read it as the weakest claim on the page.

sources

  • Β· Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 140, 1–55.
  • Β· Cattell, R. B. (1944). Psychological measurement: normative, ipsative, interactive. Psychological Review, 51(5), 292–303.
  • Β· Hicks, L. E. (1970). Some properties of ipsative, normative, and forced-choice normative measures. Psychological Bulletin, 74(3), 167–184.
  • Β· Brown, A., Maydeu-Olivares, A. (2013). How IRT can solve problems of ipsative data in forced-choice questionnaires. Psychological Methods, 18(1), 36–52.

More from the knowledge base