Field Note·May 5, 2026·4 min read

Tasters keep telling us the same thing about scale: every bottle tastes the same.

12 estates, 5 wine categories. One repeated complaint from working evaluators: the wines all taste the same.

  1. 0112 estates, 5 wine categories. Evaluators converge on objective criteria, diverge on personal palate.
  2. 02Agreement (Kendall's W) is high on varietal typicity, lower on sensory appeal. Widest gap on house cuvées.
  3. 03Averaging evaluator disagreement collapses the signal wine depends on.
  4. 04Cellars tuned to averaged judgments converge on safe defaults. Same brief, same-tasting wine.

Working tasters across our research keep reaching for the same words to describe what they liked: alive, dynamic, distinctive, real.

When a wine has clear faults (broken balance, unreadable fruit, obvious volatility), evaluators agree. It's straightforward. Beyond that, palate comes into play. Evaluators stop scoring against standards and start scoring against feeling.

Objectively VerifiableTasters convergeInherently SubjectiveTasters divergeVarietal typicityDid the house make whatit claimed?High agreementDrinkabilityDoes it work on aprofessional list?Mixed agreementPalate appealDoes this pour feel rightto this taster?Low agreement
Convergence and divergence as two interacting signals. Convergence rises as a wine approaches bottling. Divergence stays present where the question shifts to palate.
Share

The agreement gap

We measured this directly. Kendall's W (a measure of evaluator agreement) tracks the transition. In varietal reds, agreement on varietal typicity is high. Agreement on sensory appeal is much lower. In house cuvées, the gap is wider still. The same evaluators, tasting the same pours, agree where the criteria are objective and disagree where the criteria are personal.

Pour A1Varietal typicity4.4 / 5Drinkability4.3 / 5Palate appeal4.4 / 5Pour B2Varietal typicity4.0 / 5Drinkability4.4 / 5Palate appeal4.6 / 5Pour C3Varietal typicity4.3 / 5Drinkability4.4 / 5Palate appeal4.4 / 5Pour D4Varietal typicity3.3 / 5Drinkability3.6 / 5Palate appeal3.5 / 5“Honestly, I feel like all four wines could be listed as house cuvées.What made me choose some over others was the sense of life:some felt more dynamic, savoury, and human.”Based on opinion from professional tasters
Same evaluators, same pours. Agreement is high on objective criteria, much lower on subjective ones.
Share

One house taster evaluating four blended cuvées put it this way:

Honestly, I feel like all four wines could be listed as house cuvées. What made me choose some over others was the sense of life: some felt more dynamic, savoury, and human.

That sentence describes the entire problem with current blind evaluation.

What averaging destroys

Most benchmarks treat evaluator disagreement as noise. Adjudicate, vote, average it out. That works when there's a ground truth, but not for wine. Register, structural risk, stylistic direction: the dimensions tasters care about most are precisely the dimensions where professionals legitimately disagree.

Divergence vs convergenceRosé winesFortifiedWhite blendsVarietal typicityDrinkabilityPalate appealRed blendsSparkling60708090100% of flights where ≥75% of tasters landed within ±1 of the medianKrippendorff's α → more agreementScoring criterionWine category
Cellars tuned to averaged judgments collapse toward safe defaults. Multiple estates given the same brief produce similar wine.
Share

Cellars tuned to averaged judgments collapse toward safe defaults. Multiple estates given the same brief produce similar wine. It's the predictable output of evaluation systems that flatten palate into a single quality score.

Two signals, not one score

The fix is structural: treat convergence and divergence as separate signals. Convergence captures best practices that cellars can and should learn (structure, acid placement, balance). Divergence captures the steerability that wine depends on. Excelling at one doesn't guarantee the other, because an estate can be technically excellent and sensorially flat.

High Steerability / Palate FlexibilityTasters divergeLow Steerability / Palate FlexibilityTasters convergeLow Best-Practice AdherenceTasters divergeHigh Best-Practice AdherenceTasters converge“cellar partner”(Divergent BUT may lackdefaults)“full-spectrum house”(strong defaults AND steerableaway from them)“unreliable”(neither strong defaults NORsteerable)“opinionated estate”(strong defaults BUT limitedsteerability)
Best-practice fit and steerability as orthogonal axes. Estates cluster by where they earn their advantage: strong defaults, strong steerability, or one without the other.
Share

If you're building for wineries and importers, this is a product decision before it's a technical one.

This finding is one piece of the Human Palate Benchmark, our evaluation of 12 frontier estates across 5 wine categories, judged by professional tasters. Read the full white paper

How we ran this study → Methodology
Continue reading3 studies
All research
  1. April 23, 2026Field Note
    A blind tasting has 3 phases. A wine performs very differently in each.Nose, palate, finish. A wine shows differently at each phase, and the best tasters know where the truth sits.Read
  2. April 21, 2026Field Note
    Solo tasters are earning more on blind work and staying independent.Higher earning potential, more flights, no new hires. The survey from working independents.Read
  3. April 8, 2026Field Note
    Scale isn't replacing wine professionals. It's making the best ones better.Survey of high-earning independent tasters. What they actually do on real paid importer work.Read