← Knowledge base
βœ‚οΈ

The cutoff problem

Personality scores pile up in the middle, and type systems put their dividing line exactly there. Why most people are borderline, why that is the majority case rather than an edge case, and what a borderline result is good for.

How we measure

Draw the distribution of any personality trait across a large sample and you get one hump. A single peak in the middle, tapering on both sides β€” the shape of height, or reading speed, or how long people can hold their breath.

You do not get two humps. There is no valley between introverts and extraverts where nobody lives, no gap separating thinkers from feelers. McCrae and Costa established this for the MBTI dimensions specifically in 1989, and Bess and Harvey went looking again with better methods in 2002 and found the same thing: continuous, unimodal, no natural seam.

A type system has to cut that hump somewhere, and it cuts it in the middle. Which means the dividing line runs straight through the densest part of the data, where more people sit than anywhere else.

That is the cutoff problem, and its consequence is not subtle. Being borderline is the normal case, not the exception. If your result has ever felt half-right, the most likely explanation is that it is, and that most people reading the same page are in the same position.

How many people this is

Take one dimension with a cut at the midpoint. The scores nearest the line are the most common ones, so a band of moderate width around the middle catches a large share of everyone β€” on a normal distribution, the middle third of the range holds far more than a third of the people.

Now do it four times. A four-letter code requires four independent readings to each land clearly on one side. Even if each dimension separately gives a clear lean for most respondents, the chance of all four being clear falls quickly, because the probabilities multiply. The result is that a minority of people have a code where every letter is confidently theirs, and a large majority have at least one letter balanced on the line.

This is the arithmetic underneath the retest figures in test–retest reliability. It is not that the questionnaire is badly built. It is that four coin-flips near the edge are four opportunities to come back different.

The most familiar case is the one the Introvert or Extravert test has to handle on every result screen: the largest group on that dimension is neither, and a two-way question has nowhere to put them.

It is also why "what percentage of people are INFJ" questions have unstable answers across sites. The rarity of a type is partly a fact about people and partly a fact about where a particular publisher put its lines.

When a category is legitimate

Cutting a continuum is not automatically wrong. Psychology has a method for deciding when a category is really there, and it is worth knowing that the method exists, because it settles the argument in a way that opinion does not.

Paul Meehl developed taxometrics for exactly this: a set of techniques that ask whether a latent category β€” a taxon β€” underlies a set of continuous measures, or whether the measures are simply dimensional. The techniques can come back either way. Applied to biological sex, they find a taxon. Applied to most things, they do not.

Haslam, Holland and Kuppens reviewed 177 taxometric studies across personality and psychopathology and found that the great majority of constructs came out dimensional. Some candidates survived, including a small number in the clinical literature. Normal personality traits generally did not.

So the defensible statement is narrow: personality differences are real, measurable, consequential and continuous. The categories laid over them are conveniences. A convenience can still be useful β€” clinicians draw a diagnostic threshold on continuous distress because treatment decisions are binary, and the threshold earns its place by being followed by a decision. The question to ask of any personality category is the same one: what decision does this line support that the underlying number would not support better? For a clinical cutoff there is usually an answer. For a four-letter code there is usually not.

What a cut costs

Three specific losses, each of which shows up in ordinary use.

Two people on opposite sides can be nearly identical. A score of 49 and a score of 51 differ by less than the measurement error, and get different letters, different descriptions and different advice. Meanwhile a 51 and a 95 get the same letter and the same description, though they have far less in common.

The description is written for the extremes. Type portraits are drawn from people who score clearly, because those are the people the pattern is visible in. A reader at 52 per cent reads a portrait built from people at 90 and concludes the system does not work, or worse, concludes they are a bad example of their own type.

Error gets amplified rather than absorbed. On a scale, a few misread items move your score slightly. On a category, the same few items can move you to a different type with a different name. Discretising takes small measurement noise and converts it into a large change in output, which is the opposite of what a measurement should do.

There is one thing a cut genuinely buys, and it is not nothing: a category is memorable, communicable and socially useful. Four letters can be said in a sentence, printed on a profile, used to find people. A percentile cannot. That is a real benefit and it is the actual reason type systems dominate β€” not that anyone believes the categories are natural, but that a number is not shareable and a name is. The entry on type nicknames is about what happens once that name starts doing work.

If your result is borderline

Most likely it is. Here is what it is actually telling you.

A middling score is a finding, not a failure. It does not mean the test failed to read you or that you answered badly. It means you are near the average on that dimension, which is where most people are, and which is frequently an advantage β€” moderate extraversion is more adaptable than either extreme, and the same is true of most traits.

Read the dimension, not the letter. If our MBTI-style test puts you at 53 per cent toward one side, the useful sentence is "roughly balanced on this", not the letter. The percentages are printed beside the code for this reason, and the code is the part to distrust.

Do not chase the clear result. Retaking a test until a letter firms up is selecting for a lucky sitting, not discovering a truth. If the number sits near the middle across two sittings, that is the answer.

Prefer the scale when a decision is at stake. For anything consequential β€” a job, a course, an argument with a partner about how much socialising is enough β€” the continuous score in the Big Five carries more information than any four-letter code, precisely because nothing was thrown away at a boundary. And if the question underneath the decision is how much company you can take before you need the evening back, social battery is closer to the thing you are asking about than any letter is.

sources

  • Β· McCrae, R. R., Costa, P. T. (1989). Reinterpreting the Myers-Briggs Type Indicator from the perspective of the five-factor model of personality. Journal of Personality, 57(1), 17–40.
  • Β· Bess, T. L., Harvey, R. J. (2002). Bimodal score distributions and the Myers-Briggs Type Indicator: Fact or artifact? Journal of Personality Assessment, 78(1), 176–186.
  • Β· Meehl, P. E. (1995). Bootstraps taxometrics: Solving the classification problem in psychopathology. American Psychologist, 50(4), 266–275.
  • Β· Haslam, N., Holland, E., Kuppens, P. (2012). Categories versus dimensions in personality and psychopathology: A quantitative review of taxometric research. Psychological Medicine, 42(5), 903–920.

More from the knowledge base