Free tests and formal assessments
What a few hundred pounds buys when the free version asks similar questions. Not accuracy β a norm sample, a manual, a qualified interpreter and a record β and when none of that is worth paying for.
How we measure
A formal personality assessment through an occupational psychologist costs somewhere in the hundreds. A free one on a website asks recognisably similar questions and takes six minutes.
The obvious explanation is that the paid one is accurate and the free one is not. That is mostly wrong, and it is worth saying so on a site that gives tests away.
What you pay for is not accuracy. It is accountability. A published manual, a documented norm sample, a person qualified to interpret the output and answerable for it, and a record that exists afterwards. Those are real goods with real prices attached. None of them is the same thing as the questionnaire being better at reading you.
Meanwhile a free test can be built on public-domain items with decades of research behind them and still leave you with nothing you can act on β because the items were never the scarce part.
What the money actually buys
Five things, and it is worth knowing which of them you care about before deciding.
A norm sample you can name. A formal instrument reports your score against a documented comparison group: how many people, recruited how, from where, when. A free test typically compares you against whoever has taken it, which is a self-selected sample of people who clicked a link. The percentile arithmetic is identical; the meaning is not, for reasons the percentiles and norms entry sets out.
A manual. Reliability coefficients, validity studies, the intended use, the populations it was validated on, and the uses the publisher explicitly warns against. Boring, and it is the document that lets anyone check the claims. A test with no manual is not necessarily bad; it is unaudited.
A qualified interpreter. The largest single difference, and the one people underrate. A trained practitioner reads your profile in the context of your history, tells you which scores to ignore, notices the contradiction between two scales, and is professionally accountable for what they say. A results page cannot do any of that β it has one script for everyone who lands on that score.
Restricted distribution. Serious instruments are sold only to qualified purchasers, which keeps items out of circulation and keeps the norms from decaying. A test whose items are on the internet is a test people can prepare for.
A record. Something that exists afterwards, can be revisited, and can be used in a process β a development plan, a selection decision, a clinical formulation. A free result is a screen you close.
What the money does not buy
Three things people assume are on the other side of the paywall and are not.
Better items. The public-domain item pools are genuinely good. Goldberg and colleagues built the IPIP precisely so that scales measuring the same constructs as proprietary instruments would be freely available, and the IPIP versions correlate strongly with their commercial counterparts. A well-built free Big Five scale is not measuring something cruder; it is measuring much the same thing without a licence fee.
Online administration as a defect. Buchanan and Smith looked at this early and found web-administered personality measures behaved much like paper ones, with comparable structure and reliability. The internet is not the problem. Twenty-five years of research on internet-delivered testing has broadly confirmed it, and the ITC guidelines are about supervision, security and fairness rather than about the medium being inherently worse.
Immunity from the general limits. A paid assessment is still a self-report, still shaped by response styles, still reporting effects that are real and small. Paying does not change the ceiling on the method. A confident consultant with an expensive report is capable of overclaiming considerably more than a free results page does, and the price tag makes it harder to argue with.
When to pay, and when not to
Pay when a decision has consequences for someone else. Selection, promotion, anything where a person is affected by the result and could reasonably ask on what basis. The Standards for Educational and Psychological Testing exist for this case, and an unvalidated web test used to make an employment decision is indefensible and in several jurisdictions unlawful.
Pay when the question is clinical. Distress, functioning, whether something is a disorder. Screening tools indicate; they do not diagnose, and the gap between the two is a qualified assessment.
Pay when you want the interpreter more than the score. Often the value is entirely the conversation. In that case you are buying a structured couple of hours with someone trained, and the questionnaire is the agenda rather than the product.
Do not pay for curiosity. If the question is which end of a trait you are nearer, a good free scale answers it about as well and you can retake it.
Do not pay for a certificate. A paid MBTI administration is a real service and it does not make the four-letter code more stable than the free version β ours included, and the MBTI-style test here has the same boundary instability the paid one does. The instability is in the cutoff, not in the vendor.
Do not pay to settle an argument. No instrument is authoritative enough to end a disagreement about what someone is like, and treating one as though it were is the failure described in psychological essentialism.
Where we sit
We give tests away and sell a subscription around what comes after them, so we have an obvious interest in you believing free tests are enough. Here is the version we would want a friend to have.
Our tests are not formal assessments and we do not present them as equivalent. There is no published manual for any of them. We have not run validation studies of our own. Our norms come from people who took the test on this site, which is a self-selected sample, and our percentiles should be read as a rough placement rather than a population statistic.
They are short. Our Career Interest Match is built on the RIASEC interest model, which is the same framework the established vocational inventories use β and 36 items is not a vocational assessment, and we are not going to claim it is. Our Big Five is 30 items across five traits; the research standard runs from 60 items up to 240, and the difference buys facet-level resolution that we cannot offer. Our HSP test is 15 items where Aron's scale has 27. Shorter tests have wider error bands, and that is a straightforward cost of the format, not a detail.
The screening-style banks are screens. They say "not a diagnosis" wherever a score appears, and that line is doing real work.
What we think we are good at is the layer after the number: connecting two results, putting the finding in language you can use on a Tuesday, and being explicit about what a score cannot bear β which is what this entire section of the knowledge base is for. That is a legitimate product and it is not a substitute for an assessment when you need an assessment.
If you are deciding between us and a formal assessment, the question is the one from validity: what are you going to conclude from the result? If the conclusion affects a job, a diagnosis or another person, pay someone. If it is a better vocabulary for something you have been failing to describe, a free test and an honest page about its limits will get you most of the way.
sources
- Β· American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. AERA.
- Β· International Test Commission (2005). International Guidelines on Computer-Based and Internet-Delivered Testing. ITC.
- Β· Goldberg, L. R., Johnson, J. A., Eber, H. W., Hogan, R., Ashton, M. C., Cloninger, C. R., Gough, H. G. (2006). The international personality item pool and the future of public-domain personality measures. Journal of Research in Personality, 40(1), 84β96.
- Β· Buchanan, T., Smith, J. L. (1999). Using the Internet for psychological research: Personality testing on the World Wide Web. British Journal of Psychology, 90(1), 125β144.