Validity
A test is not accurate or inaccurate β a claim made from a test is. The four things that get called accuracy, which of them a personality test usually has, and the one readers feel most strongly that matters least.
How we measure
"Is this test accurate?" is the most common question anyone asks about a personality test, and the question does not have an answer, because it is asking about the wrong object.
Validity is not a property of a test. It is a property of a specific claim made from a test score. The same questionnaire can support one inference well and another not at all: a conscientiousness scale may be a decent basis for predicting whether someone files their expenses on time and a terrible basis for deciding whether they should be promoted. Same instrument, same score, two claims, two verdicts.
This is not a technicality invented to avoid the question. It is the settled position of the field β Messick argued it at length and the Standards for Educational and Psychological Testing, the discipline's shared rulebook, opens on it. The useful reformulation is: what is this score being used to conclude, and has anyone checked that conclusion?
Do that and the question becomes answerable. It also becomes uncomfortable, because for most free tests the answer to the second half is no.
Four things called accuracy
They are routinely conflated, including by people selling tests, and they are ranked here roughly from least to most informative.
Face validity β does it look like it measures what it claims? An item asking whether you enjoy parties obviously concerns sociability. This is the only kind a reader can assess unaided, it is what "this test felt accurate" usually means, and it is worth close to nothing on its own. A set of items can look perfect and predict nothing. It can also run the other way: an item that predicts well may look absurd, which is why some clinical scales contain questions that seem irrelevant.
Content validity β do the items cover the whole construct? A conscientiousness scale that asks only about tidiness has missed persistence, deliberation and reliability. Short web tests fail here constantly, because coverage costs items and items cost completion rate.
Criterion validity β does the score predict something outside the test? This is where a measure earns its keep. Barrick and Mount's meta-analysis established that conscientiousness predicts job performance across occupations, with a correlation that is modest in absolute terms and among the sturdiest findings in applied psychology. Nothing about how the items look matters once you have a number like this.
Construct validity β does the score behave the way the underlying concept should? Cronbach and Meehl's contribution, and the broadest of the four. It requires a network of evidence: the score should correlate with things the concept should relate to (convergent), and not correlate with things it should be distinct from (discriminant). Campbell and Fiske's insight was that the second half is as important as the first, and it is the half most often skipped. A new scale that correlates 0.8 with an existing anxiety measure has not demonstrated a new construct. It has renamed one.
Reliable is not valid
The two get confused, and the relationship between them is worth stating exactly because it runs one way only.
A bathroom scale that reads five kilos heavy is perfectly reliable and completely invalid: same answer every time, wrong every time. A scale that jumps ten kilos between weighings is unreliable, and therefore cannot be valid either, because a measurement that disagrees with itself cannot agree with anything else.
Reliability caps validity; it does not supply it. This is why testβretest reliability is the first thing to ask about and never the last. A test that reports a high reliability figure and no validity evidence has told you it is consistent, which leaves entirely open what it is consistently measuring.
There is a related trap on the reliability side. Internal consistency β usually reported as Cronbach's alpha β rises when items are near-duplicates of each other. A scale of ten paraphrases of the same sentence will have superb alpha and terrible content validity. High alpha is sometimes a sign of a well-built scale and sometimes a sign of a narrow one.
How instruments actually compare
Painting with a broad brush, because the details are in the manuals, but the ordering is not controversial.
Big Five instruments have the strongest evidence in personality. Five-factor structure recovers across languages and methods, the scales predict occupational, health and relationship outcomes at small-to-moderate effect sizes, and the research base is large enough that the failures are documented too. The prediction sizes are genuinely modest β correlations in the range that explains single-digit percentages of variance β and anyone quoting them as large is overselling.
Clinical screening instruments are validated against diagnoses made another way, which is a demanding criterion and the reason they come with published sensitivity and specificity figures. They are also validated for a purpose β flagging who should be assessed properly β and are invalid for the purpose most people put them to, which is self-diagnosis.
The MBTI has reasonable reliability on the underlying scales, real convergence with Big Five dimensions, and weak evidence for the claims that make it distinctive: that the categories are categories, that the code predicts job performance or team fit, that the function stack exists. The convergence finding is the interesting one, because it means the instrument is measuring something real and then reporting it in a form that discards information. See type frameworks.
Entertainment formats β colours, animals, aesthetic quadrants β have face validity and nothing else, by design. That is not a failing as long as nobody pretends otherwise.
What to ask, and what we can answer
Four questions, in order. They take a minute and they will disqualify most of what you find.
What is the claim? Not "is it accurate" but "accurate for concluding what". Write the sentence you want to act on.
Against what was it checked? A validity claim needs an external criterion. If a test page cites only its own internal statistics, it has not made one.
How big is the effect? "Predicts job performance" is compatible with a correlation that shifts the odds slightly. Almost all real personality effects are small, and small is not nothing β it is just not what a sales page implies.
Who was in the sample? A validity coefficient established on undergraduates in one country is evidence about undergraduates in one country, and the percentiles and norms entry covers what happens when the comparison group does not resemble you.
Turned on ourselves, the answers are uneven and worth putting in writing. Our Big Five is built on the public-domain item-pool tradition, so the construct it targets has decades of validity evidence behind it β but that evidence attaches to the established instruments, not automatically to our 30-item version of them, and we have not published a validation study of our own. The career and workplace tests report patterns that are useful for reflection; we make no criterion-validity claim for them, and you should not treat a Leadership Potential score as a prediction about how you would lead or a Workplace EQ score as a measurement of how you land on colleagues β both are evaluative, highly visible traits, which is exactly where self-ratings are least accurate. The screening-style banks say "not a diagnosis" on every screen because that is the only claim they can support.
The general rule holds for us as much as for anyone: a test page that tells you what a result means, and never what it was checked against, is asking for trust it has not shown the working for.
sources
- Β· Cronbach, L. J., Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281β302.
- Β· Campbell, D. T., Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81β105.
- Β· Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741β749.
- Β· American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. AERA.
- Β· Barrick, M. R., Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1β26.