Two results come back. One ranks physical touch first, the other ranks acts of service. Both people read their partner's list, feel briefly understood, and agree to try.
Three weeks later the same argument arrives in the same shape, and one of them says the sentence that sends people back to a search box. We did the test and it did not help.
5
languages in the model
.54 to .75
correlations between how much people say they value each of them
3
different factor counts published for the same five-item model
100
couples in the study most often cited as proof that matching works
The framework makes three claims and they fail differently
Gary Chapman's book appeared in 1992 and has outsold most of academic psychology combined. Its argument has three separable parts, which matters, because they do not stand or fall together.
The first is that each person has one primary language through which affection actually registers. The second is that there are five of them. The third is that couples do better when each partner expresses love in the other's preferred channel.
Impett, Park and Muise reviewed the framework against relationship science in 2024 and found the first two claims poorly supported and the third only partly so. Their summary is worth stating plainly rather than softening. People report liking and appreciating all five ways of receiving love, the five overlap heavily enough to question whether they are distinct, and the matching hypothesis has not produced the consistent effect the popular material implies.
Why a ranking appears even when the preferences are level
The review reports correlations between the five ratings running from .54 to .75. A person who rates words of affirmation highly tends to rate quality time and physical touch highly too, because what is being measured is partly a general appetite for affection.
That is the mechanical reason a primary language shows up anyway. Most free versions of the test use paired items that force a choice between two good things, and a forced choice produces a ranking whether or not there is a real gap between the options. Someone whose underlying ratings sit at 4.1, 4.0, 3.9, 3.8 and 2.2 leaves with a first and a second and a third that are noise, and a fifth that is real. How forced-choice scoring differs from a rating scale covers what those numbers can and cannot be compared against, and the short version is that they rank things inside one person and say nothing about how that person compares to anyone else.
The practical consequence is narrow and useful. The bottom of a love languages ranking carries more information than the top.
The number five was never fixed
Egbert and Polk ran a validity test on Chapman's categories in 2006 and found the five-factor structure only partly supported, with items loading across categories that were supposed to be separate. Surijah and Septiarly ran a factor analysis on an Indonesian sample in 2016 and reported reasonable internal consistency alongside a structure that did not reproduce the model cleanly.
Taken together the published solutions have come out at three factors, at four, and at five, depending on the sample and the items. That is not the signature of a construct with five natural joints in it. It is the signature of a reasonable vocabulary that has been cut at a memorable number.
The study that supports matching, read carefully
There is one piece of evidence that gets cited on the other side, and it deserves better than the summary it usually gets.
Mostova, Stolarski and Matthews recruited 100 heterosexual couples in 2022, measured each person's preferred way of receiving affection and each partner's expressed way of giving it, and computed a mismatch score from the discrepancies across all five languages. Mismatch was associated with lower satisfaction on both measures they took.
- Relationship satisfaction in women: 0.4absolute r
- Sexual satisfaction in men: 0.37absolute r
- Relationship satisfaction in men: 0.36absolute r
- Sexual satisfaction in women: 0.21absolute r
Those are respectable correlations for this literature. They are also cross-sectional, taken at one sitting from a Polish sample of mixed-sex couples, and they cannot say which way the arrow points. A couple whose relationship has been going badly for a year will express less affection of every kind, which registers as a high mismatch score without any language having been mismatched.
What a mismatch score might really be counting
Look at how the number is built. It sums discrepancies across all five languages, so the person who scores as badly mismatched is not the one aiming at the wrong category. It is the one who has stopped aiming.
Someone who is attending to their partner will land on the right channel some of the time by accident, because attention produces variety. Someone who has checked out expresses little of anything, and the sum of five discrepancies goes up together. A high mismatch score is therefore consistent with two people speaking different languages and equally consistent with one person no longer speaking.
This is the reading nobody offers, and it changes what to do next. Two rankings that differ are a translation problem and translation is learnable. A mismatch that covers all five categories at once is usually not about categories.
Why it helps some couples anyway
None of the above means the exercise is empty, and the honest position is more awkward than either camp usually admits.
Joel and colleagues pooled 43 longitudinal datasets covering 11,196 couples in 2020 to find out which self-reported variables actually predict relationship quality. The strongest ones were perceptions of this particular relationship, with perceived partner commitment and appreciation at the top. Nothing resembling preference matching appears anywhere on that list, and appreciation, which does, is a thing a person notices rather than a channel a partner selects.
Couples who take the test still report that it helped, and the most likely reason has nothing to do with whether the five categories are real. The test produces a structured conversation about what care looks like, held at a neutral moment rather than during an argument, with a vocabulary that lets someone say I need more of this without saying you have been failing me. That is a genuine intervention, and it would probably work with four categories or seven.
The site's entry on emotional needs covers what the research actually models underneath this, and the day after a fight covers the conversation that decides more than any ranking does. The love languages test here reports all five as scores rather than crowning one, which is the format the correlation data argues for.
The couple who searched the phrase did not want a taxonomy. They wanted to know why effort keeps failing to arrive, and the ranking answers that question only in the case where effort is still being made.
Sources
- Impett, E. A., Park, H. G., Muise, A. (2024). Popular psychology through a scientific lens: evaluating love languages from a relationship science perspective. Current Directions in Psychological Science, 33(1).
- Mostova, O., Stolarski, M., Matthews, G. (2022). I love the way you love me: responding to partner's love language preferences boosts satisfaction in romantic heterosexual couples. PLOS ONE, 17(6).
- Egbert, N., Polk, D. (2006). Speaking the language of relational maintenance: a validity test of Chapman's five love languages. Communication Research Reports, 23(1).
- Surijah, E. A., Septiarly, Y. L. (2016). Five love languages scale factor analysis. Makara Human Behavior Studies in Asia, 20(1).
- Joel, S., Eastwick, P. W., et al. (2020). Machine learning uncovers the most robust self-report predictors of relationship quality across 43 longitudinal couples studies. Proceedings of the National Academy of Sciences, 117(32).