This site runs a typology test, so the unflattering half goes first.
Part of what makes the question hard is that the MBTI names two different things. One is an instrument with a publisher, a manual and a research literature, and the claims made for it can be checked. The other is a vocabulary that escaped into general use decades ago β the sixteen four-letter labels, the profiles, the memes, the thing people put in bios. The second has almost nothing to do with the first, and most of what people mean when they say the test was accurate is a response to the second.
What the research actually says
The type you get is not stable across retests. This is the finding that does the most damage and it is not disputed by anyone, including the instrument's own manual. Agreement on all four letters between two sittings a few weeks or months apart is commonly reported somewhere in the region of half, with figures ranging widely depending on the interval and the sample. Whichever end of that range you take, a large proportion of people get a different label from the same instrument with nothing about them having changed. An instrument with that property can describe; it cannot classify.
The dimensions are not actually two-humped. If sixteen types were real categories, scores on each dimension would pile up at two ends with a gap in the middle, the way height does for a species with two distinct sizes. They do not. On every dimension the distribution is a single hill with most people clustered around the middle β the ordinary shape of a continuous trait. McCrae and Costa reported this in 1989 and it has been replicated repeatedly since. It is also the direct cause of the retest problem: when most people are near the cut point, small changes flip the letter.
Four of the letters map onto existing trait dimensions; the fifth trait is missing. Extraversion and introversion track the extraversion factor closely. Sensing and intuition track openness. Thinking and feeling track agreeableness. Judging and perceiving track conscientiousness. Nothing in the four-letter model corresponds to emotional stability, which is the dimension with the strongest relationship to wellbeing, relationship outcomes and mental health. A model of personality that omits it is not neutral about what it leaves out. The type frameworks entry goes through where the four-letter model holds and where it gives way.
It does not predict the things people use it to predict. A 1991 review by a United States National Research Council committee looked at the evidence and concluded it did not support the instrument's use in career counselling, which was and remains one of its largest applications. Type is a weak predictor of job performance and of most outcomes people care about. It is used in hiring and team assignment anyway, in places where it should not be, and the publisher has said as much.
A result feeling right is weak evidence. The Forer effect β people rating vague, generally flattering descriptions as highly accurate portraits of themselves β was demonstrated in 1948 and has not stopped being true. Type profiles are written warmly, describe traits nearly everyone has some of, and contain no unflattering content that would allow you to reject one. The sense of recognition when you read your profile is real and it is not information about the measurement.
It is close to absent from the research literature. Working personality researchers use five-factor models and their descendants, and the four-letter model appears mostly in applied and organisational settings rather than in peer-reviewed personality journals. This is not a conspiracy; it is what happens to a model that measures real variance in a lossy way when a less lossy option exists.
What it is still good for
Having said all that: the popularity is not purely a marketing accident, and dismissing it as astrology misses what it is doing.
It is a shared vocabulary, and shared vocabularies have real value. I need to think before I answer and I am not being evasive is a sentence a lot of people cannot construct about themselves until a framework hands them the words. Couples, teams and families who have a common language for a difference tend to argue about the difference rather than about each other's character, and that is a meaningful improvement over what they were doing before.
It is also a reasonable scaffold for reflection, precisely because the stakes are low and no result is a bad one. The failure mode is not people reading their type; it is people using it as a reason. A type is a description of tendencies, not an explanation of a decision and not a permission slip, and the sentence I'm an introvert so I can't is the point where a useful vocabulary has become an excuse with four letters on it.
What it is not good for is anything with a gate: hiring, promotion, team selection, deciding who is suited to a job. There the unreliability stops being a curiosity and starts costing somebody a career.
What it depends on
The honest answer to how accurate it is for you personally is that it depends on where you sit on each of the four dimensions, and that is a number rather than a letter. Somebody at the 90th percentile on extraversion has a stable E that will still be an E next year. Somebody at the 52nd has a letter that will flip on a bad Tuesday, and the sixteen-type description built on it will be wrong in a way that is not the reader's fault.
This is why our version shows the dimensional percentile under each letter instead of only the label, and why the Big Five sits next to it on the same site β the five-factor version is duller, has better retest properties, and includes the dimension the four-letter model leaves out. How percentiles and norms work is worth ten minutes if you want to read any personality result properly, including ours. If the question underneath yours is how to tell a measured instrument from a well-designed quiz at all, reliability and validity in plain language is the general version of this page. And if the practical complaint is that the result moved rather than that the model is wrong, why a personality test result keeps changing separates the part that is rounding from the part that is a person on a different week.