Three results, three different answers, and the uneasy feeling that the problem is the person taking the test rather than the test.
It is worth naming what that feeling assumes: that there is a single true label sitting underneath, and that a sufficiently honest sitting would recover it. That premise is the part that does not survive contact with the measurement literature.
.31
rank-order consistency of personality traits measured in childhood
.74
the same figure between the ages of 50 and 70
152
longitudinal studies behind those numbers
6.7 yr
the interval they are held constant at
How stable is personality, actually
Roberts and DelVecchio pooled 152 longitudinal studies and 3,217 test-retest correlations to answer this properly, holding the interval constant at 6.7 years so that the numbers are comparable to each other.
- Childhood: 0.31r
- College years: 0.54r
- Age 30: 0.64r
- Ages 50 to 70: 0.74r
Two readings of that chart are both correct and they point in opposite directions. Personality is meaningfully stable, and the stability increases with age, and by middle life the ordering of people on a trait holds up impressively across seven years. Also, none of those numbers is close to 1, and the childhood figure is low enough that it barely constrains anything.
What the chart measures is rank order: whether the people who were high stay high relative to everyone else. That is the thing that holds. It is not the same as an individual score staying put, and it is nowhere near the same as a label staying put.
What moves between two sittings
Something real, something rounded, and something that belongs to the instrument rather than to anyone.
The real part is that a trait is not a constant quantity a person carries around. Fleeson's experience sampling work found that across two or three weeks of ordinary life the typical person expresses nearly every level of every trait, while the average across all those moments stays almost perfectly stable. The average is what a good instrument is trying to estimate. A single sitting is one draw from a wide distribution, and asking someone to summarise a distribution in twenty questions on a Tuesday is a lossy operation by design.
The rounded part is where most of the drama lives. Any instrument that reports a category has a cut point somewhere, and a score is on one side of it or the other. A position at 52 against 48 produces a capital letter with exactly the same confidence as a position at 89 against 11. Two answers change, four points move, the letter flips, and in a four-letter system the entire profile flips with it.
And the instrument's own part is that every version uses different items, a different number of them and a different comparison group. A score is a position against a norm sample, so changing the sample changes the number without anything about the person changing at all. How percentiles and norms work is the ten minutes that makes every result on every site readable.
| What changed | How far it can move a score | Is it about the person |
|---|---|---|
| Mood, sleep, the week | Several points on energy and outlook items | It is about the week, which is a real thing and not the trait |
| Which context was answered for | Often more than the week does | Yes, and it is worth recording which one |
| Different items, different norm sample | Several points, invisibly | No |
| Position near a cut point | Nothing at all, and the label still flips | No |
| Six years of adult life | A modest amount, in a consistent direction | Yes, and this is the only row that means what people hope it means |
Why type systems make it worse
They are not unlucky. They are built in a way that guarantees this.
If the sixteen types were genuine categories, scores on each dimension would pile up at two ends with a gap in the middle, the way a measurement does when it is picking out two distinct kinds of thing. McCrae and Costa checked this in 1989 and found the ordinary shape of a continuous trait instead: a single hill with most people near the centre. Fraley and colleagues later ran taxometric analyses on adult attachment and reached the same verdict there.
That distribution is the direct cause of the instability. Most people sit near the middle of at least one dimension, so most people are near at least one cut point, so most people have at least one letter that is being generated by rounding. Pittenger's review of the four-letter instrument found retest agreement poor enough to undercut the classification claim while leaving the underlying dimensions intact, which is a more awkward verdict than either side of that argument usually quotes.
What a moved result is worth reading as
A moved result is usually treated as a failure of the test, or of the person. It is more informative read as a measurement, which is what it is.
The dimension that keeps flipping is telling you where you sit. Being genuinely balanced between two poles is a way to be, and it is a way that a two-box system has no notation for, so it alternates between two descriptions that are both slightly wrong. That is information about a person, delivered badly.
The direction of a change carries something too. Results taken after a hard stretch often move toward a less flattering profile, and the honest reading is not always noise. People answer first sittings somewhat aspirationally and later ones somewhat more plainly, and a shift in that direction may be a gain in accuracy rather than a loss of stability.
And a result that moves after a real structural change is doing its job. A new city, a new job, the end of something long. Those events move scores, and an instrument that reported the same number through all of them would be measuring less than it claims to.
The part that does not move
It would be easy to finish on the note that nothing is measurable, and that conclusion is not supported either.
The rank-order figures in the chart are, by the standards of psychology, strong. Someone at the far end of a dimension at thirty is very likely to be near that end at forty. Mean levels drift in a consistent direction across adulthood, with most people becoming somewhat more conscientious and somewhat more emotionally stable, and that drift takes decades and is visible in population data rather than in one person's two sittings. Whether believing that traits are changeable does anything to that drift is a separate claim with a far thinner evidence base than its popularity suggests, and what a growth mindset test actually measures goes through what happened when it was pooled across several hundred studies.
So the instrument is doing something. What it is not doing is issuing a verdict, and the gap between those two claims is where almost all of the disappointment with personality testing lives. How accurate the four-letter model is covers the specific case, and why a result differs between sittings covers the version most people arrive with.
The question people ask is which of the results is the real one. The better question is what the distance between them is measuring, because that distance is usually the most specific thing either sitting produced.
Sources
- Roberts, B. W., DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age: a quantitative review of longitudinal studies. Psychological Bulletin, 126(1).
- Fleeson, W. (2001). Toward a structure- and process-integrated view of personality: traits as density distributions of states. Journal of Personality and Social Psychology, 80(6).
- McCrae, R. R., Costa, P. T. (1989). Reinterpreting the Myers-Briggs Type Indicator from the perspective of the five-factor model of personality. Journal of Personality, 57(1).
- Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3).
- Fraley, R. C., Hudson, N. W., Heffernan, M. E., Segal, N. (2015). Are adult attachment styles categorical or dimensional? A taxometric analysis. Journal of Personality and Social Psychology, 109(2).
- Roberts, B. W., Walton, K. E., Viechtbauer, W. (2006). Patterns of mean-level change in personality traits across the life course. Psychological Bulletin, 132(1).