โ† Blog
PersonalitySeptember 20, 2026By Johnson

Why Two Free Enneagram Tests Give You Two Different Numbers

A 2 in one tab and a 6 in the other, and both descriptions land. The disagreement is not sloppiness in one of the tests. It follows from what the system sorts by and what a questionnaire is able to see.

Two tabs, two free tests, one evening. The first comes back with a 2. The second comes back with a 6. Both descriptions are uncomfortably good in places and plainly wrong in others.

The usual conclusion is that one of the tests is badly made. The more interesting possibility is that both worked as designed and the number was never fully reachable by the method.

104

independent samples in the 2021 systematic review

9

types the theory posits

fewer than 9

factors that factor analyses of Enneagram measures typically recover

0

studies that have derived the nine types by clustering rather than assuming them

The system sorts by something a questionnaire cannot observe

This is the part that makes the Enneagram different from trait questionnaires, and it is also the part that makes it hard to measure.

The defining claim is about motive. Two people can perform the same behaviour, arriving early, taking on the extra task, staying calm in the argument, and belong to different types because the reason underneath differs. Helping to be needed is a different type from helping to keep things safe, and the visible act is identical.

A questionnaire item can ask what a person does. It can ask what they agree with about themselves. It cannot watch why.

McClelland, Koestner and Weinberger set out the problem in 1989 in a paper that has shaped motivation research since. They distinguished self-attributed motives, which people report about themselves when asked directly, from implicit motives, measured indirectly through what people spontaneously produce. Measures of the same motive taken the two ways seldom correlate with each other, and they predict different classes of behaviour. Self-attributed motives track immediate responses to structured situations. Implicit motives track what somebody does over time when nothing is prompting them.

An online Enneagram test reaches the first kind only. It arrives at a motive by asking somebody to endorse a description of their own motive, which is a self-attribution, and the research says self-attribution and the motive that actually drives sustained behaviour are separate measurements that do not reliably agree.

What the review of the whole literature found

Hook and colleagues reviewed 104 independent samples in 2021 and reported mixed evidence on reliability and validity, which is a fair summary and worth unpacking, because the mix is not evenly distributed.

The claimWhat the review found
Enneagram scores are internally consistent and stable enough to useBroadly supported, with reliability the strongest part of the evidence
Subscales relate to established traits in the predicted directionsSupported, with theory-consistent correlations to the Big Five
Items form nine distinct factorsNot supported, since factor analyses typically recover fewer than nine
The nine types are natural groupings of peopleUntested, since no study has used clustering to derive them
Wings and movement under stress describe real dynamicsLittle supporting research either way
People find it useful for personal growthSupported by several studies

The pattern in that table is specific rather than dismissive. The measurement behaves like a set of correlated scales that work reasonably well, sitting under a theory of nine discrete kinds of person that the measurement has never confirmed.

That is exactly the condition under which two tests disagree. If the underlying structure has fewer than nine dimensions in it, then the nine labels are cuts made across a smaller space, and two instruments that cut at slightly different places will hand the same person different numbers without either of them malfunctioning.

The near tie nobody shows you

Most free tests report a winner. Underneath the winner is a set of nine scores, and the gap between first and second is often a point or two.

Somebody whose top three land at 31, 30 and 29 has a result that will not survive a bad week, a differently worded item set, or a mood. Somebody whose top score leads by twelve points has a result that will. Both leave with a single number and a description, and nothing on the page distinguishes the two situations.

The reason a result moves between sittings covers the general version of this across every framework. The Enneagram version is sharper, because the categories are defined by motive rather than by degree, so there is no honest way to report a 4 who is nearly a 5. A trait score can say 61st percentile. A type has to pick.

Both descriptions felt true, and that is not evidence

The second thing that happens on a two-test evening is that both profiles land. That feeling is not a sign that the person is somehow both types.

Enneagram descriptions are built around core fears, and core fears are close to universal. Fear of being worthless, fear of being abandoned, fear of being controlled, fear of being ordinary. Most adults can find recent evidence for all of them, which is precisely the property that makes a personality description feel accurate regardless of whether it was assigned correctly. The Barnum effect is the name for the mechanism, and it is stronger for motive language than for behaviour language, because behaviour can be checked against last Tuesday and a fear cannot.

The way to break the tie is to stop reading the descriptions of the types and start reading the items that produced them.

What a questionnaire samples
What the type is defined by
What a person says they usually do
Why they did it the specific time it mattered
A motive the person can name about themselves
A motive that shows up in what they do unprompted
Agreement with a written description
A pattern visible to someone watching over years
A score that moves with the week
A structure the theory says is fixed from childhood
Nine ranked totals
One core type with the other eight excluded

What to do with two numbers

The pair is more informative than either one alone, and there is a specific way to read it.

Find the items where the two tests disagreed hardest. Two instruments that both returned a 9 tell you little, since agreement is cheap when the items are similar. Two that split between a 2 and a 6 are usually splitting on a single axis, in that case whether the underlying drive is toward being needed or toward being secure, and the honest answer is often that both are present and one of them shows up under pressure.

Then check the avoidance rather than the aspiration. The system claims types are organised around what a person is trying not to be, and self-report is least distorted on the question of what somebody avoids, because there is less social reward in the answer. Somebody who cannot say which of two descriptions they want to be true has usually already noticed which one they work hardest to keep away from.

None of this makes the framework worthless. It makes it a vocabulary with a measurement problem attached, which is a different complaint from the one usually levelled at it. The same gap between a system and the names it travels under turns up across every typology on the internet, and the Enneagram's version is the least visible because the numbers look so much like data.

Sources

Frequently asked questions

Should I just pick the type that feels most like me?

Self-typing is what most experienced practitioners recommend and it carries its own bias, so it is worth knowing which one. People tend to choose the type whose description flatters the story they already tell about themselves, and the type that fits is often the one that stings slightly. A practical check is to ask someone who has watched you under pressure which of two candidate descriptions sounds more like you at your worst, since the system claims to describe behaviour in the stressed direction as well as the settled one.

What is a wing, and does the research support it?

A wing is the neighbouring number said to colour the core type, written as 4w5 or 9w1. The 2021 systematic review found little empirical support for wings or for the movement between types under stress and growth, which are the two parts of the theory that make it a dynamic model rather than a list. That does not make the vocabulary useless in a conversation, but a wing is currently a description rather than a finding.

Is the Enneagram related to the Big Five?

Enneagram subscales do correlate with five-factor traits in ways the theory predicts, which is one of the more encouraging results in this literature. Types associated with anxiety and vigilance tend to score higher on neuroticism, and types associated with achievement tend to score higher on extraversion and conscientiousness. The correlations are real and moderate, which means the two systems are describing overlapping ground rather than the Enneagram measuring something entirely separate.

Where did the nine types come from originally?

Not from data. The nine-pointed figure reaches back through Gurdjieff in the early twentieth century, and the assignment of nine personality types to its points was made by Oscar Ichazo and developed by Claudio Naranjo in the 1960s and 1970s, largely through teaching and clinical observation rather than through measurement. Questionnaires came afterwards and were built to detect the nine types rather than to discover how many there are.

Can an employer use an Enneagram result?

It should not be used for selection, and the reason is narrower than a general objection to personality testing. The instruments were built for self-understanding, they are transparent enough that a motivated candidate can steer the result, and the validity evidence for the nine categories is not at the level any selection procedure requires. Team development conversations are the use the material was written for.

Does a number change over a lifetime?

Enneagram theory says the core type is fixed and that what changes is how healthily it is expressed, which is a strong claim with little longitudinal evidence behind it either way. What has been observed is that test results move, and the theory absorbs this by treating a changed result as a mistyping rather than as a change. That is a difficult position to test, since it makes the claim compatible with every possible observation.

See this pattern in your own numbers

โœฆ

This article, about your situation

Luna has read the same research โ€” and, if you let her, your results.

โœฆ Talk it through

More notes on people