Response styles
Two people with the same trait can produce different scores because of how they use a rating scale. Acquiescence, extreme responding, social desirability and the reference-group effect, and why none of them are visible from the inside.
How we measure
Two people are equally talkative. One rates I enjoy meeting new people at five out of five; the other, who lives at the same level of talkativeness, rates it four, because they reserve fives for things they are certain about.
The scores differ. The trait does not. What differs is response style β the habitual way a person uses a rating scale, independent of what is being rated.
This is not a rounding error. Weijters and colleagues showed these habits are consistent within a person across questionnaires and across time, which means they behave like traits in their own right: stable, individual, and measurable. And because a style applies to every item, it shifts every scale at once in the same direction, which makes it invisible in exactly the way that matters. A profile skewed by response style still looks like a coherent profile.
The uncomfortable part is the last: you cannot see your own. Nobody experiences themselves as an acquiescent responder. They experience themselves as answering honestly, which they are.
The four that do most of the damage
Acquiescence β the tendency to agree. Some people say yes to a statement and then also say yes to its opposite, not from carelessness but because agreement is the path of least resistance and most statements have something true in them. A test that phrases every item in one direction cannot distinguish an acquiescent responder from someone who genuinely has the trait. This is why well-built scales reverse-score some items, and why a questionnaire with no reversed items is a small warning sign.
Extreme versus midpoint responding. Some people live at one and five; others compress everything into three and four. The compressor's scores gravitate toward the middle on every trait, which on a type test means landing near every boundary at once β the situation the cutoff problem describes, produced by a style rather than by an actual balance of traits.
Socially desirable responding. Not simply lying. Paulhus separated it into two components, and the distinction is the useful part: impression management, the deliberate presentation of a better version to an audience, and self-deceptive enhancement, the honest but inflated self-view. The first responds to anonymity. The second does not, because the person answering believes their answers. Crowne and Marlowe's scale, built to catch the general tendency, has been in use since 1960 and still detects a lot of it.
The reference-group effect. The most counterintuitive and, in some settings, the largest. You rate yourself against the people you are surrounded by. Heine and colleagues showed this wrecks cross-cultural comparisons of raw Likert means: members of a culture with high average conscientiousness rate themselves against high-conscientiousness peers, so national averages can come out backwards from behavioural indicators. The same mechanism runs at smaller scale. Change teams and your self-rated organisation shifts without your habits changing at all.
What moves an answer in the moment
Style is stable. Context is not, and it is stacked on top.
Who you think is reading. The same questionnaire answered for yourself, for a therapist, and for an employer yields three profiles. Not dishonesty β a different audience makes a different aspect of the truth relevant.
Which "you" you are answering as. Asked whether you are organised, are you answering about work or home? Most people silently pick one, and most questionnaires never specify. Instruments used in occupational settings add a frame of reference for exactly this reason, and consumer tests almost never do.
Recency. "How often do you lose your temper" is answered from the last vivid instance, not from a tally. A bad week produces a different profile from a good one, which is a large part of why results move between sittings.
Question order and fatigue. Answers late in a long questionnaire are flatter and closer to the midpoint than answers early in it. On a 40-item test this is small. It is not zero.
Common method variance. Podsakoff and colleagues catalogued what happens when everything is measured the same way at the same moment by the same person: the correlations between scales are inflated by the shared method. Applied to a personality profile, it means the relationships between your scores β the shape of the profile, which is what gets interpreted β are the part most contaminated.
What can be done about it
Research instruments have partial fixes. Consumer tests use almost none of them, and the reasons are commercial rather than technical.
Balanced keying. Reverse half the items so acquiescence cancels. Cheap, effective, and it makes a questionnaire feel repetitive and slightly adversarial, which costs completion.
Ipsatising within a person. Standardise each person's answers against their own mean and spread, removing the style and keeping the pattern. This works, and it destroys between-person comparability β the trade-off described in forced choice and Likert scales.
Forced choice. Pitting two equally desirable statements against each other defeats impression management well. It also throws away the degree.
Validity scales. Clinical instruments include scales that detect inconsistent answering, over-claiming and defensiveness, and discard a protocol that fails them. Almost no free test does this, because telling a user their result has been discarded is a bad product experience.
Longer scales. More items per trait average out more noise. Also more dropout.
The screening-style banks are the place this bites hardest, because a band boundary behaves like any other cutoff. A midpoint responder can land a whole point lower on the Anxiety Check than an extreme responder with the same experience, which is part of why that page says "not a diagnosis" wherever a score appears.
Notice the shape of every trade: every correction costs items, time or comparability. A free web test optimises for completion, so it takes the other side of each one. That is not a conspiracy, it is an incentive, and knowing which incentive built the instrument tells you most of what you need about what its scores can bear.
Reading your own result with this in mind
Three practical adjustments, and one piece of honesty about our tests.
If your whole profile is middling, suspect your scale use before you conclude you are unremarkable. A flat profile is produced as often by midpoint responding as by actual moderation. The check is easy: look at whether you used one and five at all anywhere in the questionnaire.
If your whole profile is flattering, look at it again. Self-deceptive enhancement is not detectable from the inside β that is its definition β and the correction is not introspection, it is asking someone. The self-report entry covers what other people's ratings add and where they beat yours.
Read the shape, cautiously, and the relative ordering more than the absolute heights. Style inflates or deflates everything together, so your highest scale is probably still your highest scale even if the whole profile is shifted. The ordering survives what the levels do not.
As for us: none of our tests use balanced keying throughout, none carry validity scales, and none ask you to specify a frame of reference before you begin. Our Big Five is 30 items across five traits β six items per trait, which is a short measurement by any standard and leaves plenty of room for style to matter. We report percentages rather than raw scores, which helps with interpretation and does nothing about this problem. If you want a result less exposed to it, take the same test twice a fortnight apart in different moods and read the overlap rather than either sitting β and if the two sittings differ a lot, the Moody Test is measuring the thing that moved.
sources
- Β· Paulhus, D. L. (1984). Two-component models of socially desirable responding. Journal of Personality and Social Psychology, 46(3), 598β609.
- Β· Crowne, D. P., Marlowe, D. (1960). A new scale of social desirability independent of psychopathology. Journal of Consulting Psychology, 24(4), 349β354.
- Β· Weijters, B., Geuens, M., Schillewaert, N. (2010). The individual consistency of acquiescence and extreme response style in self-report questionnaires. Applied Psychological Measurement, 34(2), 105β121.
- Β· Heine, S. J., Lehman, D. R., Peng, K., Greenholtz, J. (2002). What's wrong with cross-cultural comparisons of subjective Likert scales? The reference-group effect. Journal of Personality and Social Psychology, 82(6), 903β918.
- Β· Podsakoff, P. M., MacKenzie, S. B., Lee, J.-Y., Podsakoff, N. P. (2003). Common method biases in behavioral research: A critical review of the literature and recommended remedies. Journal of Applied Psychology, 88(5), 879β903.