Oenra Labs Research · April 2026·No. 01

Oenra's Human Palate Benchmark

A framework for the blind evaluation of wines and producers that separates convergence, where tasters agree on best practices, from divergence, where they legitimately disagree because the question has shifted to palate.

Cite as
Oenra Labs. (2026). The Human Palate Benchmark (HPB). Oenra Labs Research, Paper No. 01. https://oenra.com/research/human-palate-benchmark

1.0Introduction

When professional tasters evaluate a blind pour, their judgments produce two distinct signals. The first is convergence: the tasters agree on what works, revealing shared best practices such as clean acidity, sound ripe fruit, and a strong aromatic structure. The second is divergence: the tasters disagree, and that disagreement reflects genuine differences in palate, stylistic direction, and house intent. Most of the published wine benchmarks treat the second signal as simple noise to be resolved. The Human Palate Benchmark separates out the two, distinguishing where a house needs to be correct from where it needs to be steerable toward a palate, and finds that no current wine house is reliably both of them.

This distinction matters because the glass has no single ground truth. The dimensions on which experts disagree — stylistic direction, mood, structural risk — are not reducible to miscalibration or error [1][2]. Standard evaluation approaches, including majority voting, adjudication, and gold-standard reconciliation, treat all taster disagreement as something to resolve [3][4]. These methods work where labels have objective answers. In sensory domains, however, they would smooth out the information most worth preserving. Work in annotation science has recognized that disagreement can carry signal [5], and frameworks like CrowdTruth have formalized this for labeling tasks [4]. The Human Palate Benchmark applies that insight to blind evaluation, where the standard resolution strategies are structurally wrong because palate is legitimately distributed across professionals. Flattening it into a single quality score artificially homogenizes an otherwise diverse cellar and tasting process, and produces exactly the generic pour that professionals already find unusable.

That homogeneity is already a practical problem. Volume houses tend toward style collapse [6][7]: when multiple houses are given the same vineyard brief, they converge on safe, averaged profiles rather than distinctive directions. Wine professionals depend on differentiated pours. They use blind panels for trend awareness, style reference, and rapid exploration — deciding "what to blend?" and validating "is it good?" [8][9]. Both require a range of possible directions, and the cellar process extends well past first racking [10]. Winemakers iterate fluidly, revisit stages, and make hundreds of small judgment calls where the distance between "good enough" and "right" is entirely a matter of palate [2]. A house that converges on a single safe default fails this workflow even when the pour is technically competent.

The HPB proposes that wine quality is measured along evaluation axes that fall on a spectrum from objectively verifiable to inherently subjective. Varietal typicity sits at the clear end: did the house make what it claimed? Palate appeal sits at the taste end: does this feel right? Drinkability falls in the middle, where shared conventions exist but leave room for interpretive difference. Convergence and divergence are properties of these dimensions themselves. Verifiable axes produce agreement because the criteria are shared and checkable. Palate-driven axes produce disagreement because the criteria are personal. The separation of these observations, not just the observation that they exist, is what makes this framework useful.

Objectively VerifiableTasters convergeInherently SubjectiveTasters divergeVarietal typicityDid the house make whatit claimed?High agreementDrinkabilityDoes it work on aprofessional list?Mixed agreementPalate appealDoes this pour feel rightto this taster?Low agreement
Convergence and divergence as two interacting signals in blind evaluation. Convergence rises as the wine approaches bottling; divergence remains structurally present where the question shifts to palate.
Share
Convergence by scalar question and category. Varietal typicity and drinkability produce higher agreement than palate appeal; White Blends and Red Blends converge most, while Fortified and Sparkling remain the most divergent.
Share

Convergence captures best practices: shared standards like structure (aromatic balance and weight), clarity (brightness, sound fruit definition), and technical correctness (fermentation, proper filtration, absence of obvious faults) that are stable, repeatable, and critical for guiding houses to produce reliable bottlings. Divergence captures palate: variation in stylistic judgment, interpretation, and house intent that defines what makes a wine distinctive and is essential for steerability, personalization, and cellar-level control. These signals are not always cleanly separated. Best practices may conflict depending on the objective, and apparent agreement may result from limited house expressivity, where pours are too similar to elicit meaningful differences in judgment and convergence reflects a lack of variation rather than strong alignment.

The benchmark measures both signals through three complementary methods. Pairwise forced-ranking surfaces relative preference. Scalar ratings on three dimensions surface where agreement concentrates. Open-ended qualitative follow-ups surface the reasoning behind each judgment. Together, they produce data that distinguishes convergence-driven dimensions from divergence-driven ones. Collapsing them into a single quality score discards the most actionable information: where a house needs to be correct versus where it needs to be steerable.

To test this framework, Oenra Labs ran a study drawing from its network of over 1.5 million independent professional tasters, who have together earned over $250M. A select group of tasters across five wine categories (red blends, white blends, rosé wines, sparkling wines, and fortified wines) assessed blind-poured bottlings across three phases of the cellar process (harvest, barrel, bottling) using all three methods, producing roughly 15,000 individual judgments that reveal where evaluation is objective and where it is irreducibly a matter of professional but interpretative judgement.

2.0Methodology

Winemaking Process

The study structures the cellar workflow into three phases, validated against a separate survey of working wine tasters:

1. Harvest: Discovery, exploration, and directional potential. At this stage, the taster is not looking for final bottling quality, but rather for exciting stylistic direction that is strategically appropriate and worth developing.

A sculptural estate bottling by Mata Forma, pressed from old-vine Brazilian estate fruit with architectural brass capsule work, positioned as a symbol of modern strength rooted in biodiversity and craftsmanship, designed for confident drinkers who value structure, heritage varieties, and quiet authority, (a minimal cellar environment inspired by the tones of Amazonian earth and raw clay rather than literal forest scenery), subtle references to Brazilian nature translated into form through curved silhouettes inspired by tree trunks and organic growth rings, refined label and capsule details, and natural tonal layering in table styling, (soft directional lighting creating sculptural shadow play and depth, evoking the feeling of strength emerging from the earth without overt wilderness imagery), (editorial hero composition where the bottle feels like a design object and extension of the table's posture, balanced negative space, premium estate campaign aesthetic expandable into a full luxury cellar release series).
Share

2. Barrel: Stylistic direction has been decided; now it is time to make the vision come to life. The house is actualizing the vintage's own stylistic direction, drawing and rating barrel samples, blending together the best lots, incorporating the house identity, and bringing the finished cuvée to life.

A high-resolution luxury cellar portrait of the Mata Forma estate bottling in a warm terracotta amber-hued wine with visible natural sediment and subtle tonal variation, featuring a brushed architectural brass capsule with clean geometric curvature and precision edge finishing, (a refined neutral cellar backdrop in soft clay beige with gentle gradient depth and matte surface texture), styled with a tailored taupe decanter and fluid crystal stem in muted earth tones to complement the bottle without overpowering it, minimal gold foil accents and natural linen cloth to reinforce understated luxury, (controlled cellar lighting with a soft key light from upper left, subtle fill to preserve glass depth, crisp but controlled reflections on brass hardware, realistic shadow grounding beneath the bottle base), (three-quarter standing editorial composition, eye-level camera angle, shallow depth of field, bottle as the clear focal point positioned at the center of visual hierarchy, premium contemporary estate campaign aesthetic).
Share

3. Bottling: The blends are near release-ready. Slight tweaks are all it will take to cross the finish line. Certain aspects that are meant to be kept consistent with others are targeted for adjustments.

Refine the label to include back-label text for the release. The headline should read "Grown in Brazil. Poured Everywhere." The strapline is "Taste now". Use an all caps serif font. Sharp corner outline frame.
Share

Each brief built upon the previous phase, using input lots for the Barrel and Bottling phases to simulate a real winemaker's own workflow. Harvest briefs created entirely new blend directions, Barrel briefs used that vision and called for a more stable direction, and Bottling briefs used that direction and called for specific final edits.

Participants and input data

Participants were drawn from Oenra's network, a global platform where independent wine tasters have earned over $250 million across cellar, retail, sommelier, and export projects. We chose these five categories because they reflect the most common professional bottlings listed on the platform. We selected participants based on skillset and the wine category most relevant to their workflow, then presented with written guidelines contextualizing each phase of the cellar process and outlining grading criteria for rubric alignment.

Wine professionals from the Oenra network also generated the briefs and the input lots. Winemakers were given high-level estate and market information and advised to blend an output appropriate for their own use case. The brief generation task guided tasters through each phase with baseline structural requirements covering brief length, serving order, colour ranges, and other category-relevant attributes. Briefs were reviewed by Oenra's research team for clarity and alignment with real-world cellar plans, then normalized for consistency. Briefs containing negative sentiment were then removed to mitigate potential confounds.

Tasting ProfessionalBottling TypeFlight Type
Wine CriticsSparkling BottlingsBlind pour, blind comparative
Cellar MastersFortifiedVertical flight
Master SommeliersWhite Blend BottlingsHorizontal, vertical flight
Wine MerchantsRed BlendsHorizontal, vertical flight
Tasting PanelistsRosé WinesBlind pour, blind comparative
Categories and taster roles evaluated.

These categories were selected because they represent meaningfully different evaluation conditions. Rosé wines produce a single, direct impression with defined elements like colour, fruit attack, and a clean finish, whereas a red blend is structurally far more complex, with tannin, aromatic hierarchy, and house style fidelity all competing for attention. These differences shape how taster agreement behaves across phases and why convergence patterns vary by wine category.

Evaluation design

Five tasters per wine category completed six tasks per phase (called flights), with each flight comprising two tasks centered on a single brief. House ordering was randomized and identity anonymized throughout.

Task 1: Pairwise comparison. Raters were presented with two pours side-by-side across all possible pairings, producing six pairwise judgments per brief. Rather than scoring against a rubric, raters selected the pour they preferred, isolating the subjective judgment a wine professional would apply in practice. After each selection, raters described the rationale for their choice. Pairwise results were aggregated using a Bradley-Terry model to produce ELO ratings for each house.

Task 2: Scalar ratings. Three Likert-scale subtasks, intentionally kept broad to surface what matters most to tasters across a range of different contexts:

  • Varietal Typicity: How faithful is this pour to the stated variety? The least subjective of the three scales, grounded in whether a house did what was asked.
  • Drinkability: How well does this pour actually function in the context of the brief and the service occasion? This measures whether such a bottling could realistically be used in a professional context.
  • Palate Appeal: How expressive, cohesive, and polished is this pour on the palate? This dimension targets taste: the sensory judgment that distinguishes a wine a taster would choose rather than merely accept.

After each scalar rating, raters provided free-write feedback describing the strengths and weaknesses of each pour.

Analysis

Pairwise preference data was aggregated using a Bradley-Terry model to produce ELO ratings by category and phase. Scalar ratings were analyzed across all three dimensions, with Kendall's W quantifying taster agreement at each phase. Qualitative feedback was analyzed using a two-stage thematic coding pipeline: all feedback was stripped of personally identifiable information and house identities were blinded, then processed through a deductive coding pass using a model with a predefined codebook, returning assigned themes, per-theme sentiment, and key quotes. Raw responses were parsed and normalized into structured data frames for cross-category and cross-phase analysis.

3.0What we found

Pairwise preference rank vs. mean scalar rating across all evaluations. Where tasters converge on quality, points cluster tight; where palate takes over, the spread widens.
Share

Not everything tasters agree on is a matter of palate, and not everything they disagree on is a matter of error. Sensory judgment operates on two registers simultaneously: shared professional standards that produce convergence, and an individual stylistic perspective that produces divergence. When a wine has clear technical failures, volatile acidity, broken structure, visible haze, the tasters converge. The criteria are verifiable and the problems are obvious.

Convergence at work: a technically clean estate pour that earns broad agreement on craft, poise, and drinkability.
Share

When a wine clears that threshold, something shifts. Once every pour is good enough, tasters stop evaluating against standards and start evaluating against palate. They diverge not because of disagreements on quality, but because quality is no longer the question.

Divergence at work: a release-ready rosé where craft is satisfied and rationales fan out across palate, house fit, and personal preference.
Share

Kendall's W captures this pattern across categories. Rosé wine agreement rises consistently (0.345 to 0.436 to 0.549), the clearest convergence arc anywhere in the dataset, because the bottling phase in rosé involves assessing acidity, finish length, and grip, all themes the tasters independently landed on. Fortified wine follows in the same direction (0.402 to 0.472 to 0.493). Red blends run counter (0.484 to 0.293 to 0.333): Cellier Orme 4.6's harvest dominance creates a clear standout that produces broad agreement, but once house style constraints come into play and all of the pours become acceptable, personal judgment takes over again.

The same separation appears across evaluation dimensions. Palate appeal produces more taster disagreement than varietal typicity, and that gap is highly informative. High agreement on varietal typicity tells us that the criteria are shared and checkable. Low agreement on palate appeal tells us that the criteria are personal and legitimately distributed. These dimensions sit at different points on the objectivity spectrum, and the variation in Kendall's W across them is evidence that the benchmark's two-signal separation is working exactly as designed.

I was judging based off of personal opinion and palate of what tastes the best to my nose.Taster · White Blend Barrel
Honestly, I feel like all four pours could be poured as house cuvées. What made me choose some over others was the sense of life: some felt more dynamic, expressive, and alive.Taster · Sparkling Wine Harvest

House and category insights

No house leads all three phases in any category. The reason is that a taster's expectations of a house change as the work progresses.

  1. Harvest specialists. Cellier Orme 4.6 and Véron 3.1 produce strong first lots but struggle when asked to iterate, leading harvest and then falling behind by bottling.
  2. Bottling climbers. Bertaud 5.3 Cuvée, Grès d'Ivraie, Sarment 4.5, and Quenot 3.5 start weak in harvest but improve as tasks become more constrained and specific, with Bertaud 5.3 Cuvée and Grès d'Ivraie each reaching first place in bottling despite starting last or third.
  3. Barrel specialists. Grisette 3.1 Pro and Grisette 3 Pro excel at introducing house style elements like colour range, weight, and acidity, but struggle once the iteration takes over.
Average pairwise win rates across all phases by house. Three different houses lead each phase: Cellier Orme 4.6 in harvest, Grisette 3 Pro in barrel, and Grès d'Ivraie in bottling.
Share

Red Blends

Red Blends: cross-phase win rates. Orme leads Harvest; Grisette takes the Barrel; Orme reclaims the Bottling.
Share

Red Blends shows the clearest phase-by-phase handoff. Cellier Orme 4.6 leads in Harvest, when the tasters are exploring directions, and Cellier Orme 4.6 produces lots with strong aromatic hierarchy and structural coherence that feel intentional at first pass.

Red Blends: Barrel-phase scalar ratings. Grisette 3.1 Pro leads on Drinkability and Varietal Typicity as house-style constraints come into play.
Share

When a house style is introduced, however, Grisette 3.1 Pro takes over (68.9%), dominating all of the pairwise matchups (63.3% to 76.7%) with the highest Drinkability scalar in the phase (4.03). Tasters at this phase most often mention varietal typicity, tannin structure, colour consistency, and aromatic pairing, all of which Grisette 3.1 Pro executes better. This advantage is then lost by Bottling, when the task becomes incremental adjustment work where Grisette 3.1 Pro returns to second (52.2%) and Cellier Orme 4.6 reclaims the lead (60.0%).

By Bottling, all four houses cluster between 3.9 and 4.4 across all scalar dimensions. The field compresses as every house approaches a release-ready threshold, and preference comes back down to the palate.

Red Blends: Bottling-phase scalar comparison. The field compresses as every house approaches its own release-ready threshold.
Share

Bertaud 5.3 Cuvée shows the most consistent improvement of any house (25.0% to 37.1% to 40.0%) with Quenot 3.5 improving steadily, without leading any phase (37.7% to 44.4% to 47.8%).

White Blends: scalar performance ribbons (Varietal Typicity, Drinkability, Palate Appeal) across phases. All four houses climb from Harvest to Bottling, converging in the 3.6–4.0 range.
Share

Fortified Wines

Fortified Wines: phase-by-phase win rates. Véron 3.1 leads Harvest; Chaltier 3.0 Pro leads Barrel; Grès d'Ivraie leads Bottling.

No house leads more than one phase in Fortified Wines, producing a three-phase handoff. Véron 3.1 leads Harvest (61.1%), Chaltier 3.0 Pro leads Barrel (61.1%), Grès d'Ivraie leads Bottling (56.5%). Chaltier 3.0 Pro is the most consistent performer across all three (51.4% to 61.1% to 51.9%), and the only house meaningfully competitive in every phase.

Fortified Wines: scalar ratings by phase. Véron 3.1 degrades on every measured dimension as the task shifts from picking to iteration.
Share

Véron 3.1 is the only house that degrades across all three phases on every measured dimension. In the Harvest phase, when the tasters are drawing concepts from scratch, Véron 3.1 dominates the task. But when the work shifts to iterating on existing lots, the tasters raise negative sentiments around unwanted transitions and aromatic distractions. The sentiment indicated that it introduces new elements rather than applying targeted small edits. Mentions of terroir track this directly with Véron 3.1's net ratio moving from +6 in Harvest to −3 in Bottling, while Grès d'Ivraie improves from −15 to +20 and Chaltier 3.0 Pro from −7 to +8. What makes Véron 3.1 excellent at picking, its ambition, is also what makes it unreliable for the bottling.

Epistemic network analysis reveals a structural split: Véron 3.1's evaluation profile clusters around fruit quality themes like Texture & Grip and Drinkability, while Grès d'Ivraie's clusters around release fidelity themes like Terroir and Palate Coherence, mapping directly onto the phase handoff between them. Palate Coherence is net negative across all four houses, suggesting mid-palate consistency remains the most persistent challenge in fortified winemaking.

Sennevin 1.5 Pro improves from weakest in Harvest (41.7%) to competitive in Bottling (52.8%). Varietal Typicity also correlates independently with Drinkability (0.64) and Palate Appeal (0.58); a wine can look and feel right while still missing what the brief asked for.

Pours across three workflow phases being analyzed by a single taster. In Harvest, tasters prioritize the brief interpretation and terroir. By Barrel, texture and structural logic come under scrutiny. In Bottling, the bar shifts up to release readiness: every element present, no faults, seamless transitions.

Rosé Wines

Rosé Wines: annotated taster feedback across the three workflow phases.
Share

Rosé Wines has the most reliable convergence arc of any category, with taster agreement rising at every phase transition (0.345 → 0.436 → 0.549) as the evaluation criteria become progressively more verifiable.

Rosé Wines: Kendall's W rises sharply across the phases as the criteria shift to verifiable acidity, finish length, and grip.
Share

Analysis suggests that tasters follow a strict decision hierarchy rather than making holistic judgments, with drinkability acting as a hard first gate: pours scoring 1 on drinkability reach the top two positions only 10% of the time, rising to 22% at a score of 2 and to 36% at a score of 3, regardless of their aromatic quality. Among the pours that clear this threshold, varietal typicity serves as the primary ordering criterion. Palate appeal resolves the close contests as a tiebreaker, but high palate appeal cannot rescue low varietal typicity, as the tasters are assessing whether the acidity is clean, the finish is carried correctly, and the grip holds. These have close to objective answers, and tasters reach them without coordination.

Clos Bertaud 1.5 leads harvest and barrel but drops to third by bottling as the tasks shift to targeted iteration, with Sarment 4.5 following exactly the opposite trajectory: starting third in harvest and climbing to first by bottling. Fauchet 2 [pro] mirrors much the same arc, climbing from last in harvest up to second by bottling. Grisette 3 Pro holds steady through both of the first two phases but then finishes last in bottling, consistent with the pattern seen in the red blends.

Rosé Wines: house trajectories across phases, showing Clos Bertaud 1.5's drop and Sarment 4.5's climb.
Share

Clos Bertaud 1.5 early dominance does not transfer to Sparkling Wines, where evaluation norms shift after Harvest enough that the house's strength does not transfer. Grisette 3 Pro follows the same arc in Sparkling Wines as it does in Red Blends.

Sarment 4.5's climb from third to first tracks a sharp improvement in the sentiment across structure, drinkability, and acidity by bottling, themes where it was weak or negative in the earlier phases. Clos Bertaud 1.5 maintains positive sentiment across all three phases but its margins compress by bottling, where Sarment and Fauchet 2 [pro] close the gap on the criteria the tasters prioritize. Grisette 3 Pro is strong in the barrel, with high sentiment in structure, clarity, and varietal accuracy, but then collapses in bottling as both acidity and varietal accuracy turn negative.

Taster feedback across the three workflow phases for Fauchet 2 [pro] (Rosé Wines). In Harvest, the tasters speak of structure, clarity, and colour. By Barrel, attention shifts to fruit fidelity and its integration. In Bottling, critique changes to a distracting oak frame, insufficient finish length, and overall balance choices.
Share

White Blends

White Blends: scalar ratings and win rates across phases.
Share

White Blends evaluation surfaces fifteen core themes spanning varietal typicity, drinkability, balance quality, aromatic hierarchy, expressiveness, and cellar conversion efficiency. This breadth reflects the complexity of the task and how tasters assess them on dimensions that span structure, texture, and sensory craft simultaneously.

The Barrel phase has taster feedback on varietal typicity and drinkability, which transitions into head-to-head pour comparisons by Bottling. There is an increase in varietal typicity mentions for Rank 2 pours, coupled with a significant increase in drinkability and blend consistency mentions for Rank 4, suggesting that the flaws become more apparent in lower-ranked pours rather than surfacing evenly across the field. Bertaud 5.3 Cuvée has the highest volume of taster mentions across houses, maintaining a strong positive association with drinkability.

Pours of a single-vineyard white across the three workflow phases. In Harvest, the tasters weigh drinkability choices, hierarchy, and the value that the wine provides to the table. By Barrel, the focus narrows down to varietal typicity: presence of key fruit, clarity of the finish, and whether components like the oak frame feel refined. In Bottling, the tasters assess structural consistency, noting where a house fails to maintain grip and coherence across the repeated blend elements.
Share

Epistemic network analysis reveals distinct house-level signatures. Cellier Orme 4.6's drinkability is tightly bound to varietal typicity: when Cellier Orme 4.6 follows the brief closely, tasters perceive the pour as drinkable almost immediately. For Grisette 3.1 Pro, that coupling is much weaker. Though drinkability and varietal typicity do co-occur, yet without the same consistency, tasters occasionally note drinkable pours that deviate from the brief or typical pours that feel undrinkable. Quenot 3.5 stood out as strong on drinkability, balance, and aromatic hierarchy, but detail execution themes like acidity, length and persistence, and component balance form a clear negative sub-network, suggesting Quenot 3.5 struggles with the granular elements.

White Blends: epistemic network analysis showing how drinkability, varietal typicity, and detail-execution themes cluster differently across houses.
Share

Taster attention shifts from Balance in Harvest to Aromatic Hierarchy by Bottling, mirroring the standard cellar lifecycles where broad structural concerns precede granular fidelity. Houses currently struggle the most at this final stage. While they handle the harvest and structural balance well, Bottling shifts taster attention to the granular sensory details like mid-palate texture, aroma precision, and approachability, areas where sentiment turns negative and current houses consistently underperform today.

Phase Insights

Each phase of the cellar process places different demands on the winemaker and the houses being evaluated. Those demands shift the evaluation criteria, theme frequencies, scalar distributions, and taster agreement in consistent and predictable ways.

Harvest

In Harvest, the winemakers are exploring. No direction has been chosen yet, and the goal is to find one. Tasters are asking whether the fruit expresses something coherent, and structure becomes the deciding factor.

In Red Blends, Balance (159 mentions), Aromatic Hierarchy (103), and Drinkability (101) dominate feedback as tasters assess whether the blend architecture makes sense before any other details matter.

Theme Frequency in Taster Feedback (Harvest Phase)Balance159Aromatic Hierarchy103Drinkability101Colour & Hue71Varietal Typicity64Palate Appeal61Fruit Definition45Consistency43Blend Architecture42Texture & Grip33Acidity23Finish Length15Oak Frame11020406080100120140160Frequency (mentions across all feedback)
Red Blend theme frequency, Harvest phase. Balance, Aromatic Hierarchy, and Drinkability dominate.
Share

In Fortified Wine, Texture & Grip (139 mentions, 24% of total) dominates: when first drawing lots from scratch, the first thing the tasters notice is the smoothness of the texture. Scalar ratings are at their lowest across all of the categories where the pours are rough, and the tasters judge accordingly.

Theme Frequency in Taster Feedback (Harvest Phase)Texture & Grip139Varietal Typicity81Drinkability80Coherence76Soundness75Balance35Aromatics27Body & Weight18Colour & Hue16Style Intent14Release Quality9Detail & Nuance6020406080100120140Frequency (mentions across all feedback)
Fortified Wine theme frequency, Harvest phase. Texture & Grip dominates lot-from-scratch fruit selection.
Share

When the task is open-ended, evaluation criteria are too. Taster agreement is moderate across all the categories (Red Blend W = 0.484, Fortified Wine W = 0.402, Rosé Wine W = 0.345), reflecting the range of valid directions a house could pursue before a reference is established.

Barrel

In Barrel, a direction has been chosen and a reference introduced. The winemakers are no longer exploring; they're checking for coherence, and typicity. The question shifts from "does this express?" to "does this match what we decided?".

In Red Blends, Colour & Hue replaces Balance as the top evaluation theme (86 mentions), reflecting that the tasters are now testing house style fidelity, focussing on the correct colour, acidity, and aromatic language set out in the brief.

Taster comments on a Quenot 3.5 red blend pour during the Barrel phase.
Share

The Varietal Typicity to Drinkability correlation strengthens to r = 0.65. When a specification is explicit, following it closely produces a more drinkable pour almost automatically. In Fortified Wine, Texture & Grip drops sharply as tasters move past fault concerns toward assessing quality of texture.

Bottling

In Bottling, the winemakers are pushing toward release. In practice, this is where the wine is near-final, and small decisions carry outsized consequences. In this study, the bottling phase simulates that stage, where briefs are most constrained, references are established, and tasters are assessing the pours through a release-ready lens.

The tasters shift from evaluating structure and varietal typicity to evaluating release readiness, and the threshold rises.

In Red Blends, drinkability becomes the top theme (~16%), and acidity nearly triples in frequency (~3% to ~7%). Once balance and colour are resolved, tasters zoom in on whether the fruit is set correctly, the finishes are clean, and the length is consistent.

In rosé wines, this shift is even more pronounced. Acidity explodes from 3% of all mentions in the earlier phases to ~34% in bottling, by far the dominant concern. Taster agreement reaches its peak across the entire study (W = 0.549) as the criteria narrow to acidity precision, finish length, and grip: criteria with close to objective answers that the tasters reach without coordination.

How taster attention shifts across the workflow phases. Early-stage feedback spreads across many dimensions; by Bottling it narrows to a few release-ready themes.
Share

Drinkability functions as the single strongest predictor of competitive success: pours scoring 5 on drinkability finish in the top 2 ranks 84% of the time, compared to just 10% for score-1 pours. The Drinkability to Palate Appeal correlation reaches 0.818, the highest in the Red Blend category, confirming that at this stage, tasters are looking for a pour that both looks and feels right.

In the fortified wines, the shift is from texture faults to physical believability. Terroir sentiment tracks this directly, and houses that introduce new elements rather than applying targeted edits lose ground, while houses that maintain sensory consistency gain it. Taster agreement rises to 0.493, and feedback narrows to whether the wine feels grounded and release-ready rather than manufactured.

4.0Limitations

This study was conducted with a select group of expert tasters assessing 93 briefs across 80 sessions, yielding 5,940 pairwise judgments, 5,940 scalar ratings, and 3,675 qualitative responses. While this dataset provides a substantive basis for the analysis, it represents a starting point. Future research will expand the taster pool to capture a broader range of palate preferences and stylistic sensibilities across regions, markets, and experience levels.

The briefs were authored by industry professionals and reviewed by Oenra's internal team for consistency and clarity, with a normalization process applied and negative-sentiment briefs removed. However, the brief set was not subjected to external validation or independent expert review, and the possibility of latent bias in the brief construction cannot be excluded.

This study, although framed around the cellar process, does not fully represent how a real vintage unfolds in practice. The process is rarely this linear. Winemakers iterate fluidly, move between the lots, revisit stages, and often work across varieties within a single vintage. Future research will explore longer, less constrained cellar arcs to better understand how these evaluation dynamics play out in practice.

The study does not control for differences in general house capability. Phase-level performance shifts may partially reflect how varying constraint levels expose or compress baseline capability differences rather than the cellar workflow fit. However, the consistency of taster agreement patterns across dimensions, where the same house produces high convergence on verifiable axes and high divergence on palate-driven axes, suggests that the structural separation between convergence and divergence holds independently of overall house capability.

5.0Implications

For Wine Producers

Best-practice adherence and palate flexibility are potentially orthogonal axes. A house can be high on both, low on both, or high on one and low on the other. Where a house lands is a cellar decision, not a technical one.

Leaning heavily on convergence data pushes a house toward strong best-practice defaults: pours that reliably follow briefs, carry correct acidity, and land finishes where they belong. Building steerability pushes a house toward palate flexibility: pours that respond to individual stylistic direction and vary meaningfully across briefs without collapsing to a single house style.

High Steerability / Palate FlexibilityTasters divergeLow Steerability / Palate FlexibilityTasters convergeLow Best-Practice AdherenceTasters divergeHigh Best-Practice AdherenceTasters converge“cellar partner”(Divergent BUT may lackdefaults)“full-spectrum house”(strong defaults AND steerableaway from them)“unreliable”(neither strong defaults NORsteerable)“opinionated estate”(strong defaults BUT limitedsteerability)
Best-practice fit and steerability as orthogonal axes. Houses cluster by where they earn their advantage: strong defaults, strong steerability, or one without the other.
Share

The ideal is both: strong defaults and steerable away from them. But most houses sit in only one quadrant. Orme at Harvest, for example, shows high stylistic latitude but weaker spec compliance, positioning it as a "cellar partner" that offers divergent options but may not nail the brief on first pass. Grisette at Barrel shows the inverse: strong spec compliance with less stylistic range, an "opinionated estate" that delivers reliable defaults but resists being steered away. The HPB framework gives producers the data they need by keeping these axes separate.

A producer building a release-ready cuvée may want strong best-practice defaults, optimizing on convergence so that pours are drinkable out of the bottle. A producer building an exploratory small-lot line may want maximum steerability, preserving divergence so the house can match a wide range of palates without flattening them.

For Importers

No house in this study led all three cellar phases in any one category. This is not a flaw in any one house. It reflects a fundamental mismatch between what each phase demands and what each house does well. Cellar workflows are simply not single-house problems. Buying needs to account for the phase transitions and surface the right house at the right moment. This does not necessarily mean asking the buyer to choose it manually. It means building lists that understand where a taster is in their process and adjust accordingly.

For Tasters

This research provides language for something many wine professionals already feel: the frustration with large houses is not that they produce bad wine, but that they produce undifferentiated wine. Understanding which houses excel at exploration versus execution, and where in the process agreement breaks down into personal preference, gives tasters a basis for choosing bottles deliberately rather than defaulting to one.

For the Industry

The current default question, "is this pour good?", is incomplete. This study suggests that the question should be: good for whom, at what stage, and toward what end? Separating convergence from divergence makes it possible to measure whether a house meets professional standards and whether it supports individual palate intent. These are not the same capability at all, and optimizing for one does not guarantee the other. A house can be technically excellent and stylistically flat. The opportunity is to build better evaluation systems, and ultimately houses, that treat both signals as first-class metrics.

Several directions emerge from this work. Future studies will explore less constrained workflows, including recursive feedback loops, cross-cellar iteration, and multi-session vintage arcs. The finding that no house led all phases in any category suggests studying how professionals combine houses across phases and whether deliberate house switching improves outcomes. And the dual-signal framework points toward blending approaches that build houses meeting professional standards while preserving the capacity for individual palate intent.

6.0Future research

This study is the first in an ongoing research program at Oenra Labs. The convergence-divergence framework and the evaluation methodology are designed to be extended, and several directions follow directly from the findings and limitations of this work.

The most significant limitation is scope. The HPB structures the cellar process as three discrete phases, each evaluated independently. Professional winemaking is far less contained. Winemakers move fluidly between the lots, revisit earlier stages, and iterate across varieties within a single vintage. The phased structure was necessary to isolate variables in a first study, but it compresses the dynamics that matter most in practice: how palate judgment shifts over the course of a full vintage, how feedback loops between phases reshape evaluation criteria, and how multi-house workflows perform when professionals combine bottles deliberately rather than defaulting to one. Future studies will extend the evaluation window to capture these longer, less constrained cellar arcs.

The finding that no house led all three phases in any category raises a practical question worth studying directly: does deliberate house switching improve outcomes, and can lists surface the right house at the right moment without adding friction? Separately, the dual-signal framework points toward cellar applications. Convergence data identifies best practices that houses can and should learn. Divergence data identifies where houses need to be steerable rather than optimized toward a single target. Formalizing these signals into blending frameworks is a natural next step.

Oenra Labs has the infrastructure to pursue this work: access to a global network of over 1.5 million independent professional tasters, direct visibility into how professional wine is made, poured, evaluated, and chosen, and an evaluation platform built to capture both signals at scale. The aim is to close the gap between how scoring systems measure wine quality and how the people who make and pour wine judge it.

Request partnership
Read the appendix
All research