Two tabs, two free tests, one evening. The first comes back with a 2. The second comes back with a 6. Both descriptions are uncomfortably good in places and plainly wrong in others.
The usual conclusion is that one of the tests is badly made. The more interesting possibility is that both worked as designed and the number was never fully reachable by the method.
104
independent samples in the 2021 systematic review
9
types the theory posits
fewer than 9
factors that factor analyses of Enneagram measures typically recover
0
studies that have derived the nine types by clustering rather than assuming them
The system sorts by something a questionnaire cannot observe
This is the part that makes the Enneagram different from trait questionnaires, and it is also the part that makes it hard to measure.
The defining claim is about motive. Two people can perform the same behaviour, arriving early, taking on the extra task, staying calm in the argument, and belong to different types because the reason underneath differs. Helping to be needed is a different type from helping to keep things safe, and the visible act is identical.
A questionnaire item can ask what a person does. It can ask what they agree with about themselves. It cannot watch why.
McClelland, Koestner and Weinberger set out the problem in 1989 in a paper that has shaped motivation research since. They distinguished self-attributed motives, which people report about themselves when asked directly, from implicit motives, measured indirectly through what people spontaneously produce. Measures of the same motive taken the two ways seldom correlate with each other, and they predict different classes of behaviour. Self-attributed motives track immediate responses to structured situations. Implicit motives track what somebody does over time when nothing is prompting them.
An online Enneagram test reaches the first kind only. It arrives at a motive by asking somebody to endorse a description of their own motive, which is a self-attribution, and the research says self-attribution and the motive that actually drives sustained behaviour are separate measurements that do not reliably agree.
What the review of the whole literature found
Hook and colleagues reviewed 104 independent samples in 2021 and reported mixed evidence on reliability and validity, which is a fair summary and worth unpacking, because the mix is not evenly distributed.
| The claim | What the review found |
|---|---|
| Enneagram scores are internally consistent and stable enough to use | Broadly supported, with reliability the strongest part of the evidence |
| Subscales relate to established traits in the predicted directions | Supported, with theory-consistent correlations to the Big Five |
| Items form nine distinct factors | Not supported, since factor analyses typically recover fewer than nine |
| The nine types are natural groupings of people | Untested, since no study has used clustering to derive them |
| Wings and movement under stress describe real dynamics | Little supporting research either way |
| People find it useful for personal growth | Supported by several studies |
The pattern in that table is specific rather than dismissive. The measurement behaves like a set of correlated scales that work reasonably well, sitting under a theory of nine discrete kinds of person that the measurement has never confirmed.
That is exactly the condition under which two tests disagree. If the underlying structure has fewer than nine dimensions in it, then the nine labels are cuts made across a smaller space, and two instruments that cut at slightly different places will hand the same person different numbers without either of them malfunctioning.
The near tie nobody shows you
Most free tests report a winner. Underneath the winner is a set of nine scores, and the gap between first and second is often a point or two.
Somebody whose top three land at 31, 30 and 29 has a result that will not survive a bad week, a differently worded item set, or a mood. Somebody whose top score leads by twelve points has a result that will. Both leave with a single number and a description, and nothing on the page distinguishes the two situations.
The reason a result moves between sittings covers the general version of this across every framework. The Enneagram version is sharper, because the categories are defined by motive rather than by degree, so there is no honest way to report a 4 who is nearly a 5. A trait score can say 61st percentile. A type has to pick.
Both descriptions felt true, and that is not evidence
The second thing that happens on a two-test evening is that both profiles land. That feeling is not a sign that the person is somehow both types.
Enneagram descriptions are built around core fears, and core fears are close to universal. Fear of being worthless, fear of being abandoned, fear of being controlled, fear of being ordinary. Most adults can find recent evidence for all of them, which is precisely the property that makes a personality description feel accurate regardless of whether it was assigned correctly. The Barnum effect is the name for the mechanism, and it is stronger for motive language than for behaviour language, because behaviour can be checked against last Tuesday and a fear cannot.
The way to break the tie is to stop reading the descriptions of the types and start reading the items that produced them.
What to do with two numbers
The pair is more informative than either one alone, and there is a specific way to read it.
Find the items where the two tests disagreed hardest. Two instruments that both returned a 9 tell you little, since agreement is cheap when the items are similar. Two that split between a 2 and a 6 are usually splitting on a single axis, in that case whether the underlying drive is toward being needed or toward being secure, and the honest answer is often that both are present and one of them shows up under pressure.
Then check the avoidance rather than the aspiration. The system claims types are organised around what a person is trying not to be, and self-report is least distorted on the question of what somebody avoids, because there is less social reward in the answer. Somebody who cannot say which of two descriptions they want to be true has usually already noticed which one they work hardest to keep away from.
None of this makes the framework worthless. It makes it a vocabulary with a measurement problem attached, which is a different complaint from the one usually levelled at it. The same gap between a system and the names it travels under turns up across every typology on the internet, and the Enneagram's version is the least visible because the numbers look so much like data.
Sources
- Hook, J. N., Hall, T. W., Davis, D. E., Van Tongeren, D. R., Conner, M. (2021). The Enneagram: a systematic review of the literature and directions for future research. Journal of Clinical Psychology, 77(4).
- McClelland, D. C., Koestner, R., Weinberger, J. (1989). How do self-attributed and implicit motives differ? Psychological Review, 96(4).
- Forer, B. R. (1949). The fallacy of personal validation: a classroom demonstration of gullibility. Journal of Abnormal and Social Psychology, 44(1).