← Blog

The Fuzzy Middle: Why Most Compatibility Scores Should Come With an Error Bar

A compatibility score is an estimate, not a measurement. Here's how measurement error, test-retest reliability, and model error make 78 and 82 the same match.

Two profiles land in front of you. One says 82% compatible. The other says 78%. Which do you message first?

Almost everyone picks the 82. And almost everyone is making a decision on noise.

That four-point gap feels like information because it's rendered as a number, and numbers carry an implicit promise: that the thing being counted is stable enough to count. Your height is 178 cm. Your bank balance is $1,240.17. A compatibility score looks like it belongs in that family. It doesn't. It belongs to a messier family — estimates derived from other estimates, each of which wobbles.

This isn't an argument that compatibility scores are worthless. It's an argument that a score without an error bar is a half-finished sentence. Here's how to finish it.

Where the number actually comes from

Every compatibility score, ours included, is the end of a chain:

  1. You answer items. Maybe 50, maybe 300. Each is a noisy proxy for something underlying.
  2. Items become trait scores. Agreeableness, attachment anxiety, conflict avoidance, novelty-seeking — each a weighted average of items, each carrying the error of its inputs.
  3. Two people's trait profiles get compared. Similarity on some dimensions, complementarity on others, deal-breaker filters on top.
  4. That comparison collapses into one scalar. Usually scaled to look like a percentage, because percentages feel trustworthy.

Four transformations. Error enters at every one and never leaves. By step four, the system has made roughly a dozen modelling decisions — which traits matter, how much, whether similarity or difference is good on each — and the output format hides all of them behind two digits.

Reliability: the number that should be printed next to every score

The foundational idea in psychometrics is almost embarrassingly simple. Any observed score is the true score plus error. The question is how much of the variation you see between people is signal and how much is churn.

Test-retest reliability answers that directly: give the same people the same instrument twice and see how well the two sets of scores correlate. For well-built personality instruments this is usually respectable but never perfect — and crucially, it depends on who you test. One analysis of the Big Five Inventory found a median test–retest reliability of .66 across a sociodemographically diverse adult sample, notably lower than the median of .78 typically reported from college-student samples. The same questionnaire. Different people. Meaningfully different stability.

That gap matters because almost all consumer matching runs on diverse adults, while much of the validation literature was built on undergraduates.

From reliability you get the standard error of measurement — the expected spread of someone's observed scores around their true score, calculated as the standard deviation times the square root of (1 minus reliability).

Let's run that arithmetic on a plausible compatibility score. This is illustration, not a published finding: suppose scores across your candidate pool have a standard deviation of 12 points, and the composite score has a reliability of .80 — generous, for a construct built on top of several trait estimates.

  • Standard error of measurement ≈ 12 × √0.20 ≈ 5.4 points
  • A 95% interval around a single score ≈ ±10.5 points

So your 82 is really "somewhere around 71 to 93." Your 78 is "somewhere around 67 to 89." Those intervals overlap across nearly their entire length. The 82 and the 78 are, for every practical purpose, the same match.

It gets stricter when you compare two scores. The error on a difference between two independently measured scores is larger than the error on either one — roughly √2 times the individual standard error, about 7.6 points in our example. For a gap to clear the 95% bar, it would need to be around 15 points. Under these assumptions, a 70 versus an 86 is a real difference worth acting on. A 78 versus an 82 is a coin flip wearing a lab coat.

Difference scores are the shakiest part of the machine

Here's the part that rarely makes it into product copy. Compatibility scoring frequently depends on difference or similarity terms — how far apart two people sit on conscientiousness, how aligned their conflict styles are. It is a long-standing point in psychometrics that difference scores inherit error from both components while often stripping out the shared variance that made each component reliable in the first place. A gap between two moderately reliable measurements is less trustworthy than either measurement alone.

Which means the parts of a matching model that feel most intuitive — "you're both high in openness!" — are frequently built from the least stable quantities in the system. We've written before about what the evidence actually supports on similarity versus complementarity; the measurement problem sits underneath that debate and constrains how confidently any answer can be delivered.

Add the fact that people don't answer honestly in a uniform way — socially desirable responding tilts scores rather than merely scattering them — and you have bias stacked on top of variance. Noise you can average out. Bias you cannot.

Thresholds turn small noise into big flips

Everything above concerns continuous scores. Categorical outputs are worse, because a cutoff converts a tiny measurement wobble into a total change of label.

The type-based traditions illustrate this vividly. Reviewing the evidence on the Myers-Briggs, psychologist Stephen Benning has noted that "more than one-third of people receive different four-letter types after a four-week period." Their personalities did not transform in a month. They were sitting near a dichotomy's boundary, and random measurement error pushed them across it.

Any matching system that reports tiers — "Excellent Match," "Good Match" — reintroduces exactly this cliff. Someone at 79.4 and someone at 80.1 get categorically different treatment because of a gap smaller than the measurement error around either of them. Bands are fine; pretending the band edges are real is not.

Even perfect measurement wouldn't save the score

Suppose you eliminated measurement error entirely. You'd still face a ceiling: how much of romantic outcome is predictable from pre-meeting self-report at all.

The evidence here is humbling and worth internalising. In a well-known speed-dating study applying machine learning to initial attraction, researchers found that a model could predict who tends to be desirable and who tends to be desiring — but not the specific, unique desire of one particular person for another.[1] The general tendencies were learnable. The chemistry between a specific pair was not.

A later effort pooled 43 dyadic longitudinal datasets from 29 laboratories. Relationship-specific variables — things like perceived partner commitment, appreciation, and sexual satisfaction — predicted up to 45% of variance in relationship quality at baseline and up to 18% at study end, while individual-difference variables predicted 21% and 12% respectively. Notably, a person's own reports carried far more predictive weight than their partner's.

Read that carefully, because it's the single most important constraint on this entire industry: the strongest predictors are properties of a relationship that doesn't exist yet. Before two people meet, you cannot measure appreciation or perceived commitment. You can only measure the weaker individual-level stuff.

This is why, in their comprehensive review of online dating, Finkel and colleagues concluded that no compelling evidence supported matching sites' claims that their mathematical algorithms produce romantic outcomes superior to other ways of pairing people.[2] Sociologist Michael Rosenfeld's assessment of such algorithms was blunter: they are, he said, "mostly smoke and mirrors."[3]

A number that explains a modest slice of variance can still be useful — it beats proximity and a photo. But it should be presented as what it is: a prior, not a prophecy.

What an honest compatibility system looks like

If you accept all of the above, several design conclusions follow, and we hold ourselves to them:

  • Show a range, not a point. "Likely in the 70s" is more truthful than "78."
  • Report confidence separately from score. Twenty answered items should not produce the same visual authority as two hundred. Low-information matches should look low-information.
  • Prefer rank over magnitude. "Top decile of your pool" survives measurement error better than a two-digit percentage, because ordering is more robust than spacing.
  • Refuse to differentiate inside the noise. If two candidates fall within the error band, present them as tied. Forcing a ranking you can't justify is a small lie repeated thousands of times.
  • Expose the drivers. Which three inputs moved this score, and in which direction? A score you can audit is a score you can argue with.
  • Say what would change it. "Answer the conflict module and this estimate tightens by roughly half" is more honest than any decimal place.

We'd rather tell you a match is uncertain than manufacture precision we can't defend. That's also why the avatars do the talking: a conversation surfaces the relationship-specific signal that no pre-meeting questionnaire can reach. The score narrows the field; the dialogue tests it.

A field guide: six questions for any score you're shown

Use these on us, too.

  1. What's the reliability of the underlying instrument, and in whom was it established?
  2. How many items produced this, and how recently were they answered?
  3. Is this score comparable across people, or is it rank-within-your-pool dressed as a percentage?
  4. What's the smallest difference the system claims is meaningful? If it can't answer, it doesn't know.
  5. Which inputs moved it most — and does that match my own sense of what matters?
  6. Has anything in this system been validated against actual relationship outcomes, or only against self-reported satisfaction with the match screen?

If a product can't answer most of these, the number is decoration.

The fuzzy middle is where you actually live

Here's the thing nobody says out loud: for most users, most of the time, the great majority of candidate matches sit in an indistinguishable band. A handful are clearly poor fits — deal-breakers collide, values diverge hard. A handful look genuinely promising. Everyone else is in the fuzzy middle, and the fuzzy middle is not a failure of the algorithm. It's the honest shape of the data.

What the score is good for is triage: pruning the obviously wrong, surfacing the plausibly right, saving you from the brute-force grind of finding out in person. What it is not good for is deciding between a 78 and an 82.

So treat the number like a weather forecast. Seventy percent chance of rain is useful; it tells you to bring a jacket. It does not tell you that today is wetter than yesterday's sixty-eight percent. Take the tests, read the range, and then go have the conversation that the number can only gesture at.

The most trustworthy thing a matching system can tell you isn't how compatible you are. It's how sure it is.

Sources

  1. Dating? A magic formula to predict attraction is more elusive than ever — ScienceDaily
  2. Online Dating: A Critical Analysis From the Perspective of Psychological Science — Psychological Science in the Public Interest
  3. Critics challenge the 'science' behind online dating — San Francisco Chronicle

Share this article

← All posts

Comments

No comments yet — be the first.