Test–retest reliability
Whether a test gives you the same answer twice is a measurable number, and it is the first one worth asking for. What the figures look like for traits and for types, and why a category is so much less stable than the score underneath it.
How we measure
Test–retest reliability is the answer to one question: if the same person takes the same test twice, with nothing important happening in between, how similar are the two results?
It is reported as a correlation between the two occasions, and it is the least glamorous number in psychometrics. It is also the first one you should look for, because everything else a test claims depends on it. A measure that disagrees with itself across a fortnight cannot be measuring a stable property of you. It is measuring the fortnight.
The finding that matters most here is not that personality tests are unreliable. Trait scores are considerably more stable than people expect. It is that a category built on top of a stable score is much less stable than the score — and almost every popular test hands you the category and hides the score.
The numbers
Rough figures, from published reliability studies rather than marketing pages.
Trait scales over weeks. Well-built Big Five instruments retest at around r = 0.80 over a few weeks and stay near r = 0.70 across a year. That is high. It means the ordering of people barely changes, and that your score today is a good bet for your score next month.
Trait scales over decades. Roberts and DelVecchio meta-analysed 152 longitudinal studies and found rank-order consistency rises with age: around r = 0.31 in childhood, 0.54 in the twenties, and about 0.74 between 50 and 70. Personality is stable, increasingly so, and never completely fixed. That last point is the one both camps get wrong.
Four-letter types over weeks. Here it falls apart. Howes and Carskadon retested over a five-week interval and found that fewer than half of respondents got the same code on all four dichotomies — the individual dimensions held for roughly 60 to 80 per cent of people, and four of those in a row is a demanding requirement. Later figures vary with interval and sample, and the manual's own numbers do too, but the shape is consistent and Pittenger's review treats it as the instrument's central problem rather than a footnote.
Notice that these two facts are not in tension. The underlying preferences moved a little. The category moved a lot, because it was sitting on a boundary.
Why a category is so much less stable
This is the mechanism, and it is simple enough to see in one example.
Suppose your thinking–feeling score is 52 per cent toward feeling. A fortnight later, after a good week instead of a bad one, it is 48. The measurement barely moved — well inside the error of any short questionnaire — and by the standards of the trait it is the same reading twice.
But the letter flipped. F became T. The label changed, the type description changed, the four-letter code you tell people changed, and nothing about you did.
Now compound it. A type code has four of these boundaries, and a person only has to be near one of them for the code to change. That is why category instability is so much higher than scale instability: you are not asking whether one measurement is stable, you are asking whether four independent measurements all stay on the same side of four lines. The cutoff problem entry covers where those lines fall relative to where people actually score, which is what makes this so common rather than rare.
Two implications follow. The middle is where results are least reliable, and most people are in the middle. And a changed letter is usually not a changed person — it is the most likely thing to happen to a borderline score, and treating it as personal growth or as evidence of fraud are both overreadings.
What moves a score, legitimately
Not all instability is noise. Some of it is the test correctly picking up something real that you did not intend to measure.
Mood. Howes and Carskadon looked directly at this: retest agreement was lower for people whose mood had shifted between sittings. A questionnaire asking how you usually behave is answered from how you currently feel, because that is the evidence at hand.
Context of administration. Answering at work, for work, produces different responses from answering at home for yourself. The frame of reference changes what "usually" means.
Reference group. You rate yourself against the people around you. Move to a louder workplace and your self-rated extraversion falls without your behaviour changing at all, an effect documented across cultures and covered in response styles.
A state measure moving is not a failure. Some instruments are supposed to be unstable, because the thing they read is. A mood-pattern measure like the Moody Test reading differently in a hard week is the instrument working, not failing — the error is only in expecting a state reading to behave like a trait reading. Check which kind you took before you worry about a change.
Actual change. Personality does move, mostly slowly and mostly in one direction — people tend to become more conscientious and emotionally stable through their twenties and thirties. Over five years, some of a shift is real. Over five weeks, essentially none of it is.
The practical rule: a difference between two sittings a month apart is measurement. A difference across several years, in the same direction, on a scale rather than a letter, may be you.
What to ask of a test, including ours
Three things, in order of how much they tell you.
Does it publish a retest figure at all? Most free tests do not, and the absence is informative — not because the test is necessarily bad, but because nobody has checked. A test that reports no reliability statistics is asking to be trusted rather than evaluated.
Does it report a score as well as a label? A label alone cannot be checked for stability by the reader. Four percentages can: take it twice and see how far each moved. This is one reason the Big Five is the more useful instrument to retake — it never converts your score into a side, so there is nothing to flip.
Does it show a range or a point? A single number implies a precision no short questionnaire has.
By those standards our own tests are mixed, and it is worth being specific. The MBTI-style test reports the four dimensions as percentages beside the code, which is exactly so that a reader near 50 can see it — and 36 items across four dimensions is nine items per dimension, which is a short measurement and should be read as a band. We do not publish a retest coefficient for it, because we have not run the study. That is a real gap, it is the same gap almost every free test on the internet has, and it is not made smaller by being common.
What we can tell you without a study is the thing this page exists to say: if your result sits near the middle on any dimension, expect it to change, do not read the change as meaningful, and use the percentage rather than the letter.
sources
- · Myers, I. B., McCaulley, M. H., Quenk, N. L., Hammer, A. L. (1998). MBTI Manual: A Guide to the Development and Use of the Myers-Briggs Type Indicator (3rd ed.). Consulting Psychologists Press.
- · Howes, R. J., Carskadon, T. G. (1979). Test-retest reliabilities of the Myers-Briggs Type Indicator as a function of mood changes. Research in Psychological Type, 2(1), 67–72.
- · Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210–221.
- · Roberts, B. W., DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies. Psychological Bulletin, 126(1), 3–25.