The 90% divorce prediction was scored on the couples it was built from
A lab watches a couple argue for fifteen minutes and calls the outcome with over ninety percent accuracy. The number is real and correctly calculated. It is also not a prediction, and the difference is the whole story.
You have probably met this claim in one of two forms. Either a laboratory can watch a couple argue for fifteen minutes and tell you, with over ninety percent accuracy, whether that marriage will end. Or, in the compressed version that travels furthest, one researcher can predict divorce from a single conversation.
The numbers behind it are real. They were calculated correctly, from patient and expensive work, by researchers who were not trying to mislead anybody. The problem is sitting inside the word accuracy.
Those figures come from models that were built and then scored on the same couples. A model tested on the couples it was made from is telling you how well it fits them. It is a description of people whose outcomes were already known when the description was written, and a description is not a forecast about anyone.
The work behind the number
Start with what actually happened, because it deserves more respect than the myth around it usually gets.
John Gottman and Robert Levenson brought married couples into a laboratory built to look like an apartment, asked them to discuss a real ongoing disagreement for about fifteen minutes, and recorded it. Facial expressions and speech were coded second by second against a fixed scheme. Heart rate and skin conductance were measured throughout. Then the couples were followed for years to see what became of them[1].
The accuracy figures came out of that pipeline. A 1998 study followed 130 newlywed couples for six years and reported that divorce and stability were classified with 83 percent accuracy[2]. An earlier study took a different route, interviewing 52 couples about the history of their own relationship, how they met, how they decided to marry, what the hard years were like, and reported over 94 percent accuracy about who had separated three years later[5]. That 94 percent, from 52 couples, is where most of the ninety-something figures in circulation originally come from.
Almost none of the criticism that follows is about whether this work was done carefully. The coding schemes were painstaking, the follow-up was long, and observing what couples do rather than asking them how things are going was a genuine methodological step forward. The argument is entirely about what the resulting percentage means.
Describing and predicting are not the same thing
This is the distinction the accuracy figure hides. It is worth slowing down for, because most readers have never had it explained and it is not difficult once someone does.
Suppose you have 52 couples and you already know which of them separated. You also have a great many measurements on each one: how often each partner rolled their eyes, how quickly their pulse rose, which words they used about their wedding day. You go looking for some combination of those measurements that sorts the couples into the two groups whose answers you already hold.
You will find one. With enough measurements and few enough couples, you are close to guaranteed to find one, because nothing stops you from adjusting the rule until the sorting comes out right. The rule you end up holding fits those 52 couples extremely well. That is what the 94 percent is measuring: how well a rule, selected precisely because it fit, fits.
Whether that rule says anything about the 53rd couple is a separate question, and the percentage does not answer it. It cannot. Those 52 couples were spent on building the rule, so they can no longer be used to test it.
The everyday version is a practice exam. Memorise the answer key and you will score full marks on that paper. The score is real, you did not cheat to get it, and it tells you nothing about how you will do on a paper you have not seen, because what it measured was recall of that particular exam rather than knowledge of the subject.
The only way to find out whether a rule predicts is to point it at people who took no part in making it. Statisticians call this cross-validation. It is not an exotic technique reserved for difficult cases. It is the ordinary standard, and it is the step these headline figures skipped.
What happened when someone did cross-validate
In 2001, Richard Heyman and Amy Smith Slep published the test that had been missing[3].
They did not rebuild Gottman's laboratory. They took archival data from the 1985 National Family Violence Survey, a large nationally representative study, and assembled a sample of 528 people: 176 who had divorced, 176 with the highest marital disagreement scores, and 176 with the lowest. Then they built a divorce-prediction model the standard way, on half the sample, and held the other half back.
On the half the model was built from, it looked excellent. Overall accuracy of 90 percent. It correctly flagged 92 percent of the people who had actually divorced.
On the half it had never seen, overall accuracy fell to 69 percent, and the share of divorced people it flagged fell from 92 percent to 46 percent. Pointed at couples it did not already know, the model missed more of the divorces than it caught.
The more revealing number is the one nobody quotes. Of the people the model flagged as headed for divorce, the proportion who actually divorced fell from 65 percent in the building sample to 29 percent in the held-back one. And there is a second problem stacked on top of the first: a third of their sample was divorced by construction, far above the real world. Adjusting to a realistic population divorce rate of around 16 percent, the figure drops to 21 percent. Roughly four out of every five couples the model condemned would have stayed together.
Heyman and Slep's conclusion was narrow and, in hindsight, mild. Prediction results reported without cross-validation should be treated with great caution however impressive they look, because impressiveness is exactly what fitting a model to its own data produces.
The findings that survive all of this
None of the above touches the observations themselves, and the observations are the part worth keeping.
Coded conflict conversations really do separate couples in distress from couples who are not. Contempt, meaning sarcasm, mockery, and the eye-roll delivered from a position of superiority, shows up more in struggling relationships than any of the other patterns catalogued, and carries more signal than the sheer volume of arguing does. Couples who are doing well really do run a far higher proportion of warmth to hostility inside a disagreement than couples who are not. These are descriptive claims about how two groups differ, they are what the coding scheme was designed to detect, and it detects them.
That is not a minor result. Before this line of work, a great deal of couples research rested on asking people how their marriage was going, which is a question people answer badly for reasons ranging from self-flattery to genuinely not knowing. Watching what couples do and coding it consistently was a real advance, and the categories it produced have proved useful enough to escape the academy entirely.
What has to stay attached to them is that a difference between groups is not a probability attached to a person. Contempt being more common in relationships that end is a fact about relationships that end. It is not a percentage chance that yours will.
It is worth noting that the Gottman Institute's own current wording is more careful than the claim in circulation. Their research FAQ frames the finding as identifying that a particular couple is behaving like the couples in the group that went on to divorce[6]. That is a statement about resemblance, and it is defensible on this evidence. It is also considerably weaker than what routinely gets said on its behalf by people who read it in a book.
The scientific argument here was conducted in the open and remains unresolved rather than settled. Scott Stanley, Thomas Bradbury and Howard Markman raised a related set of concerns in 2000 about how quickly findings from this research programme were being converted into interventions for couples, and Gottman replied in the same issue of the same journal[4]. That is what a live disagreement between serious researchers looks like.
What a real prediction would have to look like
The useful thing to take from this is a test you can apply to the next confident number you meet, in this field or any other.
A claim to predict has to do three things. State the rule before knowing the outcomes rather than after. Apply it to people who took no part in constructing it. And report the hit rate against a realistic base rate, rather than one engineered by choosing a sample that is one-third divorced.
The largest recent attempt to forecast relationship outcomes did precisely this. Samantha Joel, Paul Eastwick and 84 colleagues pooled 43 longitudinal datasets covering more than 11,000 couples, trained their models on some of those datasets, and tested them on datasets the models had never encountered[7]. Any predictor that only worked inside its own sample failed the crossing and was dropped. That is why the surviving list is short, modest, and much less quotable than a single arresting percentage. We covered what it found separately; what matters here is the procedure rather than the results.
The pattern generalises. A short honest list of predictors will always lose the attention contest to one flattering number, and the flattering number is very often flattering because the test it passed was one it could not have failed.
What this means for you
If you have been carrying the ninety percent figure around as a private verdict, put it down. Nobody has ever demonstrated the ability to look at your conflict conversation and tell you where your relationship ends. The demonstration that would establish that has been attempted, and what it produced was a model that got roughly one in five of its condemnations right.
That cuts in both directions, which is why this is not simply reassuring. If the four horsemen are not a sentence, their absence is not an acquittal either. Plenty of marriages end without anybody rolling their eyes, and a relationship can be quietly failing in ways no fifteen-minute recording would capture.
What the evidence does support is the descriptive claim, and descriptive claims are more useful to an individual than forecasts anyway. Contempt is worth taking seriously, not because it foretells anything, but because of what it is: a way of speaking to someone you are supposed to be on the same side as, from above. If it has become the register your arguments run in, that is information about the relationship you are in now, and now is the only thing you can act on.
A forecast leaves you nothing to do but wait and see whether it comes true. A description of how Tuesday's argument went is something two people can change on Wednesday.
The divorce-prediction accuracy figures were produced by fitting models to small samples and then scoring those models on the same samples, which measures fit rather than foresight. The one serious published attempt to cross-validate the approach on data it had not been built from saw accuracy fall from 90 percent to 69 percent, and the share of flagged couples who actually divorced fall to around 21 percent once a realistic divorce rate was applied.
The behavioural observations underneath are a different matter and largely stand. Contempt really is the most corrosive thing in the catalogue, and observational coding really does distinguish couples in trouble from couples who are not. What was oversold was never the observation. It was the leap from describing a group to forecasting a person.
Sources
- [1]Marital processes predictive of later dissolution: Behavior, physiology, and health(opens in a new tab)
Gottman, J. M., & Levenson, R. W. · Journal of Personality and Social Psychology · 1992
- [2]Predicting marital happiness and stability from newlywed interactions(opens in a new tab)
Gottman, J. M., Coan, J., Carrère, S., & Swanson, C. · Journal of Marriage and the Family · 1998
- [3]The hazards of predicting divorce without crossvalidation(opens in a new tab)
Heyman, R. E., & Slep, A. M. S. · Journal of Marriage and Family · 2001
- [4]Structural flaws in the bridge from basic research on marriage to interventions for couples(opens in a new tab)
Stanley, S. M., Bradbury, T. N., & Markman, H. J. · Journal of Marriage and Family · 2000
- [5]How a couple views their past predicts their future: Predicting divorce from an oral history interview(opens in a new tab)
Buehlman, K. T., Gottman, J. M., & Katz, L. F. · Journal of Family Psychology · 1992
- [6]Research FAQ(opens in a new tab)
The Gottman Institute
- [7]Machine learning uncovers the most robust self-report predictors of relationship quality across 43 longitudinal couples studies(opens in a new tab)
Joel, S., Eastwick, P. W., et al. · Proceedings of the National Academy of Sciences · 2020
More from the blog
- Myth9 min read
The gap between men and women is smaller than the gap between two men
The premise that men and women need translating to each other has sold tens of millions of books. It does not survive contact with the meta-analyses. What does survive is more useful, and it has nothing to do with which sex you are.
- Explainer9 min read
Twenty million copies sold before anyone tested the five love languages
The framework is useful and it is largely untested, and those two facts have coexisted for three decades. The studies eventually ran. Here is what survived them and what did not.