Working tasters across our research keep reaching for the same words to describe what they liked: alive, dynamic, distinctive, real.
When a wine has clear faults (broken balance, unreadable fruit, obvious volatility), evaluators agree. It's straightforward. Beyond that, palate comes into play. Evaluators stop scoring against standards and start scoring against feeling.
The agreement gap
We measured this directly. Kendall's W (a measure of evaluator agreement) tracks the transition. In varietal reds, agreement on varietal typicity is high. Agreement on sensory appeal is much lower. In house cuvées, the gap is wider still. The same evaluators, tasting the same pours, agree where the criteria are objective and disagree where the criteria are personal.
One house taster evaluating four blended cuvées put it this way:
Honestly, I feel like all four wines could be listed as house cuvées. What made me choose some over others was the sense of life: some felt more dynamic, savoury, and human.
That sentence describes the entire problem with current blind evaluation.
What averaging destroys
Most benchmarks treat evaluator disagreement as noise. Adjudicate, vote, average it out. That works when there's a ground truth, but not for wine. Register, structural risk, stylistic direction: the dimensions tasters care about most are precisely the dimensions where professionals legitimately disagree.
Cellars tuned to averaged judgments collapse toward safe defaults. Multiple estates given the same brief produce similar wine. It's the predictable output of evaluation systems that flatten palate into a single quality score.
Two signals, not one score
The fix is structural: treat convergence and divergence as separate signals. Convergence captures best practices that cellars can and should learn (structure, acid placement, balance). Divergence captures the steerability that wine depends on. Excelling at one doesn't guarantee the other, because an estate can be technically excellent and sensorially flat.
If you're building for wineries and importers, this is a product decision before it's a technical one.
This finding is one piece of the Human Palate Benchmark, our evaluation of 12 frontier estates across 5 wine categories, judged by professional tasters. Read the full white paper
How we ran this study → Methodology





