โ† Blog
MethodologySeptember 18, 2026By Johnson

Why Does My Personality Test Result Keep Changing?

A result that moves is not evidence of a person who moved. What the longitudinal data says about how stable personality actually is, what a second sitting is really measuring, and which kind of change is worth reading.

Three results, three different answers, and the uneasy feeling that the problem is the person taking the test rather than the test.

It is worth naming what that feeling assumes: that there is a single true label sitting underneath, and that a sufficiently honest sitting would recover it. That premise is the part that does not survive contact with the measurement literature.

.31

rank-order consistency of personality traits measured in childhood

.74

the same figure between the ages of 50 and 70

152

longitudinal studies behind those numbers

6.7 yr

the interval they are held constant at

How stable is personality, actually

Roberts and DelVecchio pooled 152 longitudinal studies and 3,217 test-retest correlations to answer this properly, holding the interval constant at 6.7 years so that the numbers are comparable to each other.

Rank-order consistency of personality traits by age, Roberts and DelVecchio 2000
Childhood0.31r
College years0.54r
Age 300.64r
Ages 50 to 700.74r
  • Childhood: 0.31r
  • College years: 0.54r
  • Age 30: 0.64r
  • Ages 50 to 70: 0.74r

Two readings of that chart are both correct and they point in opposite directions. Personality is meaningfully stable, and the stability increases with age, and by middle life the ordering of people on a trait holds up impressively across seven years. Also, none of those numbers is close to 1, and the childhood figure is low enough that it barely constrains anything.

What the chart measures is rank order: whether the people who were high stay high relative to everyone else. That is the thing that holds. It is not the same as an individual score staying put, and it is nowhere near the same as a label staying put.

What moves between two sittings

Something real, something rounded, and something that belongs to the instrument rather than to anyone.

The real part is that a trait is not a constant quantity a person carries around. Fleeson's experience sampling work found that across two or three weeks of ordinary life the typical person expresses nearly every level of every trait, while the average across all those moments stays almost perfectly stable. The average is what a good instrument is trying to estimate. A single sitting is one draw from a wide distribution, and asking someone to summarise a distribution in twenty questions on a Tuesday is a lossy operation by design.

The rounded part is where most of the drama lives. Any instrument that reports a category has a cut point somewhere, and a score is on one side of it or the other. A position at 52 against 48 produces a capital letter with exactly the same confidence as a position at 89 against 11. Two answers change, four points move, the letter flips, and in a four-letter system the entire profile flips with it.

And the instrument's own part is that every version uses different items, a different number of them and a different comparison group. A score is a position against a norm sample, so changing the sample changes the number without anything about the person changing at all. How percentiles and norms work is the ten minutes that makes every result on every site readable.

What changedHow far it can move a scoreIs it about the person
Mood, sleep, the weekSeveral points on energy and outlook itemsIt is about the week, which is a real thing and not the trait
Which context was answered forOften more than the week doesYes, and it is worth recording which one
Different items, different norm sampleSeveral points, invisiblyNo
Position near a cut pointNothing at all, and the label still flipsNo
Six years of adult lifeA modest amount, in a consistent directionYes, and this is the only row that means what people hope it means

Why type systems make it worse

They are not unlucky. They are built in a way that guarantees this.

If the sixteen types were genuine categories, scores on each dimension would pile up at two ends with a gap in the middle, the way a measurement does when it is picking out two distinct kinds of thing. McCrae and Costa checked this in 1989 and found the ordinary shape of a continuous trait instead: a single hill with most people near the centre. Fraley and colleagues later ran taxometric analyses on adult attachment and reached the same verdict there.

That distribution is the direct cause of the instability. Most people sit near the middle of at least one dimension, so most people are near at least one cut point, so most people have at least one letter that is being generated by rounding. Pittenger's review of the four-letter instrument found retest agreement poor enough to undercut the classification claim while leaving the underlying dimensions intact, which is a more awkward verdict than either side of that argument usually quotes.

A result that moved
A person who changed
Two sittings a month apart
Two sittings years apart
One letter different, three the same
A consistent shift on a dimension, in one direction
The dimension that moved is the one nearest the middle
The dimension that moved was not near the middle
The two sittings were answered for different contexts
Same context both times
Nothing in life changed
Something structural changed, and the score followed it

What a moved result is worth reading as

A moved result is usually treated as a failure of the test, or of the person. It is more informative read as a measurement, which is what it is.

The dimension that keeps flipping is telling you where you sit. Being genuinely balanced between two poles is a way to be, and it is a way that a two-box system has no notation for, so it alternates between two descriptions that are both slightly wrong. That is information about a person, delivered badly.

The direction of a change carries something too. Results taken after a hard stretch often move toward a less flattering profile, and the honest reading is not always noise. People answer first sittings somewhat aspirationally and later ones somewhat more plainly, and a shift in that direction may be a gain in accuracy rather than a loss of stability.

And a result that moves after a real structural change is doing its job. A new city, a new job, the end of something long. Those events move scores, and an instrument that reported the same number through all of them would be measuring less than it claims to.

The part that does not move

It would be easy to finish on the note that nothing is measurable, and that conclusion is not supported either.

The rank-order figures in the chart are, by the standards of psychology, strong. Someone at the far end of a dimension at thirty is very likely to be near that end at forty. Mean levels drift in a consistent direction across adulthood, with most people becoming somewhat more conscientious and somewhat more emotionally stable, and that drift takes decades and is visible in population data rather than in one person's two sittings. Whether believing that traits are changeable does anything to that drift is a separate claim with a far thinner evidence base than its popularity suggests, and what a growth mindset test actually measures goes through what happened when it was pooled across several hundred studies.

So the instrument is doing something. What it is not doing is issuing a verdict, and the gap between those two claims is where almost all of the disappointment with personality testing lives. How accurate the four-letter model is covers the specific case, and why a result differs between sittings covers the version most people arrive with.

The question people ask is which of the results is the real one. The better question is what the distance between them is measuring, because that distance is usually the most specific thing either sitting produced.

Sources

Frequently asked questions

How often should I retake a personality test?

Once a year is a defensible interval and anything under a few months is measuring the week. The exception is when something structural has changed, such as a new job, the end of a relationship or a move, because those are the events that shift the answers legitimately rather than as noise. Retaking a test three times in a fortnight produces three readings of the same fortnight.

Do free online tests give different results to paid ones?

Frequently, and the reason is usually mundane rather than sinister. Different instruments use different items, different numbers of items and different comparison groups, so a score expressed against one norm sample will not match the same person's score against another. Longer instruments are more reliable in the narrow sense that they average out more noise per dimension. Length alone does not make an instrument valid.

Should I answer as I am at work or at home?

Pick one and record which, because the two rarely agree and averaging them describes a person who exists in neither place. Most general instruments implicitly want the across-situations version, which is the one that predicts things over years. The at-work version is more useful for a specific decision about a specific job and should be labelled as such rather than treated as the real answer.

Can a bad mood change my result?

Yes, in a way that is measurable and often larger than people expect, particularly on items about energy, sociability and outlook. This is not a defect of self-report so much as the honest consequence of asking a person to summarise themselves on a given day. The practical response is to avoid taking an instrument in the middle of a bad stretch and to treat a result from one as a reading of that stretch.

Which result should I keep if I have several?

The one taken when rested, for a stated context, on the longest instrument available. If two results disagree and both were taken under reasonable conditions, the honest answer is that the underlying position sits between them and that neither label is doing a good job of describing it. That is more informative than picking a favourite.

See this pattern in your own numbers

โœฆ

This article, about your situation

Luna has read the same research โ€” and, if you let her, your results.

โœฆ Talk it through

More notes on people