← Knowledge base
📝

What a self-report can see

Every personality test you have taken asked you. What that method is genuinely good at, what it is structurally blind to, and why the people who know you disagree with your results in a predictable direction.

How we measure

There is one fact that applies to every personality test on this site and nearly every one anywhere else: the only instrument involved is you.

Nobody watched you at a party. Nothing was timed, counted or recorded. A set of statements was put in front of you and you reported how well each one fits. What comes back is a summary of those reports.

This is not a criticism. Self-report is cheap, fast, and — the part that surprises people — it works. Self-reported conscientiousness predicts job performance, academic outcomes, health behaviour and longevity across large samples. Self-reported neuroticism predicts relationship dissolution and mental health outcomes. These are among the more replicable findings in personality psychology, and they were all obtained by asking people.

But it is worth being exact about what was asked. A questionnaire measures your theory about yourself. That theory is informed, it is usually the best single theory available, and it is a different object from your behaviour. Most misreadings of a test result come from confusing the two.

What it is good at

Three things, and they are the things the method is uniquely suited to.

The inside. Nobody else can report what you feel in a meeting, what you were rehearsing before you spoke, or how much the small talk cost. For internal states — anxiety, motivation, the thoughts that loop at midnight — your report is not merely the best evidence, it is close to the only evidence. This is why screening-style measures for rumination or anxiety are self-report and should be.

Aggregation across situations. A researcher observing you for an afternoon sees one situation. You have watched yourself in thousands. When you answer "I usually", you are summarising a dataset nobody else has access to, including the private version and the version from years ago.

Cheapness, which is not trivial. Behavioural measurement is slow and expensive enough that most of it is never done. Everything the Big Five rests on — the five-factor structure, the outcome predictions, the age trends — was established by asking very large numbers of people rather than by watching any of them. Self-report is why personality psychology has samples in the hundreds of thousands, and sample size is why the field can distinguish a real small effect from noise at all — the difference that decided the growth mindset question.

What it is blind to

The blind spots are structural. They are not fixed by writing better questions.

You cannot report what you do not see. The traits people are least accurate about are the ones with the most at stake for how they see themselves. Vazire's self–other knowledge asymmetry model makes a specific prediction here and it holds up: people are more accurate than observers about internal, low-observability traits like anxiety, and less accurate than observers about evaluative, highly visible ones like how dominant, charming or intelligent they come across. Your friends know things about your effect on a room that you do not.

Self and informant reports agree only moderately. Connelly and Ones' meta-analysis puts the correlation between self-ratings and ratings by people who know you well around the 0.4 to 0.5 range, depending on trait and acquaintance. That is substantial agreement and substantial disagreement. Notably, a single well-acquainted informant predicts some real-world outcomes about as well as the self-report does, and combining a few informants can beat it.

Reports of behaviour are not behaviour. Baumeister and colleagues made the general complaint about the field, and Mehl's work with ambient recordings made it concrete: what people say they do and what a recorder captures them doing correlate, but not nearly as tightly as the questionnaire literature implicitly assumes.

The answer is shaped by how you answer. Acquiescence, extreme responding, the reference group you compare yourself to, and how you want to be seen all leave fingerprints on a score. That is enough of a topic to have its own entry: response styles.

Retrospection reconstructs. "How often do you lose your temper" is answered from the most available instance, not from a count. Recent, vivid and emotionally loaded episodes dominate, which is why the same person answers differently in a good week.

The gap is information

The useful move is not to distrust self-report. It is to treat the difference between your account and other accounts as a reading in its own right.

If you score yourself low on warmth and three friends independently describe you as warm, one of those is wrong, and the interesting question is which. Often the answer is that you are reporting your intention and they are reporting your effect — a real and common split, and one that matters in every place where the effect is what counts.

If you score high on a socially desirable trait and nobody who lives with you would agree, that is not necessarily dishonesty. Self-ratings on evaluative traits are the ones most contaminated by how we would like to be, and the contamination is mostly invisible from the inside.

If you score high on something unflattering, take it seriously, because the pressure runs the other way. A self-reported high score on an undesirable trait has crossed a headwind to get there.

The practical version of all this is simple and slightly uncomfortable: ask two people who know you to guess your scores before you show them. The places they miss are where your self-knowledge is good. The places they agree with each other and disagree with you are where it is not.

What this means for our tests

Every test on this site is a self-report. There is no behavioural measurement anywhere in the product, no informant version, and no observation of any kind. That is the ceiling on all of it and it is worth stating on the page rather than in a footnote.

Three consequences follow, and we would rather you knew them.

Our results describe your self-concept as much as your behaviour. On the internal scales — anxiety, overthinking, exhaustion — that distinction matters least, because the self-concept is much of the phenomenon. On the evaluative ones — leadership, workplace emotional intelligence, anything about how you come across — it matters most, and the Leadership Potential and Workplace EQ results should be read as how you see yourself operating, not as a measurement of your effect on other people. The research says those two things differ, and it says the gap is largest exactly here.

A test cannot catch a blind spot by definition. The Narcissism Spectrum test is the clearest case: the trait it measures is partly a distortion in self-perception, so the instrument is asking the distortion to report itself. Grandiose traits are measurable by self-report — people high on them often endorse the items readily — but a low score is much weaker evidence than a high one, and no screening instrument on any website resolves this.

Where our results are useful, it is as a prompt and a comparison, not a verdict. The number is worth about as much as the conversation it starts with somebody who has watched you. That is a smaller claim than most test pages make, and it is the one the method supports.

sources

  • · Paulhus, D. L., Vazire, S. (2007). The self-report method. In R. W. Robins, R. C. Fraley & R. F. Krueger (Eds.), Handbook of Research Methods in Personality Psychology. Guilford Press.
  • · Vazire, S. (2010). Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.
  • · Connelly, B. S., Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.
  • · Baumeister, R. F., Vohs, K. D., Funder, D. C. (2007). Psychology as the science of self-reports and finger movements: Whatever happened to actual behavior? Perspectives on Psychological Science, 2(4), 396–403.
  • · Mehl, M. R., Gosling, S. D., Pennebaker, J. W. (2006). Personality in its natural habitat: Manifestations and implicit folk theories of personality in daily life. Journal of Personality and Social Psychology, 90(5), 862–877.

More from the knowledge base