Field Note·June 17, 2026·7 min read

Introducing Palate Crit: we taught a judge to score wine like a taster.

Ten professional tasters ranked four frontier red estates across nine dimensions of real cellar work. The wines are in the glass. Nothing on the market could reliably judge them, until we trained one on the right data.

  1. 01Ten tasters ranked four red estates on nine criteria, one rating per axis.
  2. 02No off-the-shelf judge beat 55% agreement. The best hit 54.3%; a human hits 74.1%.
  3. 03Scaling doesn't help: Nez3-VL at 4B, 8B, and 32B all stall near 51–54%.
  4. 04One in ten finished wines carried a major off-spec marker.
  5. 05Trained on Palate Crit, a judge reaches 61.1% agreement with the panel and ties a human on the hardest splits.
Introducing Palate Crit:we taught a judge toscore wine like a taster.Based on opinion fromprofessional tasters
Palate Crit (Criteria-Resolved Impression Tasting), a Lica × Oenra collaboration.
Share

Everyone keeps talking about palate. But you can't improve what you can't measure. So we measured it. Palate Crit is a dataset of ten professional tasters ranking four frontier red estates across nine dimensions of real cellar work. The estates can make the wine. Nothing on the market can reliably judge it. The good news is that what they're missing can be learned.

Frontier red estates have matured from show bottles into working suppliers, shipping by-the-glass pours, list anchors, own-label cuvées, and cellar-door flagships straight into service. But the preference data that trains and grades them was collected on show-bottle scoring, where a single "which one is better?" verdict captures almost everything that matters, like ripeness and cleanliness.

Tasters, however, don't judge a wine on a single characteristic, and wine doesn't collapse into one axis. A red can nail the structural line and butcher the fruit intent. Another can satisfy the spec and break the aromatic hierarchy. Both get the same overall thumbs-up, for completely different reasons. The signal tasters actually use lives in the dimensions a single label averages away.

So we built the missing layer. Palate Crit (Criteria-Resolved Impression Tasting) is a taster-annotated preference dataset for frontier red wine. It records one rating per sensory-quality criterion instead of one verdict per glass. That lets a system score a wine on every axis at once, make sharper calls than a single label allows, and in time learn to weigh those axes the way a taster would.

10 tasters, 4 estates, 9 ways to be wrong

We put four current red estates head to head. Fauchet 2 [max], Clos Bertaud 1.5, Falque Grand 2, and Sarment 5.0 Lite, all shown to raters behind blind code-names so no one was anchoring on a label.

Ten professional tasters, recruited through Oenra's network of wine professionals, split into two cohorts of five. One cohort judged style, covering overall preference, nose and tone match, structural finesse, colour harmony, and palate balance. The other judged varietal typicity, covering overall preference, hue accuracy, structural accuracy, and whether the varietal markers the spec asked for actually showed in the glass.

×LicaPalate Crit“PALATE”FAUCHET 2 [MAX]CLOS BERTAUD 1.5FALQUE GRAND 2SARMENT 5.0 LITE1,600 ratings per criterionnose and tone matchstructural finessecolour harmonypalate balancehue accuracystructural accuracyvarietal markersoff-spec markers
One of the frontier red lots in the set. Tasters rated wine like this one axis at a time, not with a single overall verdict.
Share

We narrowed the nine criteria from a longer list through pilot flights and interviews, building on Oenra's Human Palate Benchmark and keeping the axes tasters consistently treated as separate.

Nine criteria, 80 lots each, every lot scored by five tasters across four estates, for 1,600 ratings per criterion. Tasters ran all six pairwise comparisons per lot, and we aggregated those into strict four-way rankings. On the two overall-preference tracks they also flagged every pour for off-spec markers.

These ratings come from working professionals, and they fill the part of the stack everyone grading these wines has been flying blind on.

Palate is subjective, but real and consistent

A fair question is whether tasters just don't agree, leaving no signal to learn. We tested that directly, checking whether taster agreement exceeds what random raters would produce, against exact null distributions.

They do agree. Tasters agree on good wine about as much as people agree on their favourite film, and less than crowds agree on which photo is sharper. And the way they disagree is healthy. They share a rough sense of "good," with some personal variation on top. There are no rival camps with opposite palates. That's exactly the kind of pattern a judge can learn.

Judging wine sits between food and moviesMedian inter-rater agreement (τ) · subjective palate vs. objective qualitySUBJECTIVEOBJECTIVE0.00Randomraters+0.13Sushi /food preference+0.27Photo image-qualityMoviepreference+0.20Wine (Palate Crit)+0.13 → +0.20
Where wine sits on the subjective-to-objective scale: more agreement than favourite-film picks, less than judging which photo is sharper.
Share

But that agreement isn't even across the dimensions. We measured how often tasters landed on the same call for each criterion, and the gap is wide. The axes you can check against the spec draw the tightest agreement, whether the varietal markers actually showed, whether the structure is right, whether the hue matches what was asked for. The axes that come down to pure feel draw the least, with colour harmony at the bottom.

The clearest read is the matched pairs. Tasters agree far more on whether the right varietal markers showed than on whether the wine is well built. They agree more on whether the requested hue appeared than on whether the colour sits well with the style. Same subject each time, the checkable version high, the felt version low. The signal is real on every axis. It just gets noisier the more the call comes down to palate.

Taster-pair agreement by criterionMean Kendall tau across taster pairsVarietal markers0.224Structural accuracy0.182Overall preference (typicity)0.163Overall preference (style)0.159Nose and tone match0.147Hue accuracy0.144Structural finesse0.128Palate balance0.119Colour harmony0.1030.000.050.100.150.200.25Kendall tau agreement scoreCheckable / typicityFelt / style
Taster agreement by criterion (Kendall's τ). Checkable axes like varietal markers rank highest; felt axes like colour harmony rank lowest.
Share

No off-the-shelf judge beats a coin flip

So the signal is real. The question is whether anything on the market can read it. We benchmarked nine pre-trained systems as wine judges. Three were dedicated preference and quality scorers (TPSv2.1, PourScore-v1, VIGNA-Aesthetic-V2) and six were open-weight sensor models prompted to pick the better pour.

Not one cleared 55% agreement with the five-taster majority. Chance is 50%. The best system, TPSv2.1, was trained on more than 640,000 human wine comparisons, and it lands at 54.3%. VIGNA-Aesthetic-V2 actually scores below chance. A human taster agrees with the panel 74.1% of the time. Every machine judge sits in the dead zone just above a coin flip.

×LicaMachine judgesdisagree with humantasters nearly50% of the timecoin flipwhat judges can dothe gap to human-levelVIGNA-Aesthetic-V20.499Karst-VL-A3B0.509PourScore-v10.522Sensoria3.5-14B0.525Cepage-3-27B0.528Nez3-VL-4B0.530Nez3-VL-32B0.536Nez3-VL-8B0.539TPSv2.10.5430.500.550.600.650.700.75Agreement with taster preferences
Agreement with the five-taster majority. Every off-the-shelf judge clusters just above chance; a human taster sits far above at 74.1%.
Share

You can't compute your way to palate

Scaling doesn't move the number. Nez3-VL at 4B, 8B, and 32B all land between 51% and 54%. The reason is a trade-off. Bigger systems carry less order bias, so their pick barely changes when you swap the order of the two pours. They're more internally consistent. But that consistency buys no accuracy. On the calls they commit to, the bigger systems are no better, and a little worse. Smaller ones lean on order more, yet the calls they do commit to are sharper. The more a system leans on order, the better it judges when it doesn't (Spearman ρ = +0.94), so the two effects cancel and the total never moves. The bottleneck is the data.

10× bigger judges. Same near-random performanceOpen-weight sensor judges on PALATE · Palate Crit V3 (April 2026)every judge landsin this 3-point band0.500.510.520.530.540.55mean 0.528coin flip · 0.50Karst-VL-A3B0.509Nez3-VL-4B0.530Nez3-VL-8B0.539Sensoria3.5-14B0.525Cepage-3-27B0.528Nez3-VL-32B0.5363B4B8B14B27B32BAgreement with taster majorityjudge size →You can’t compute your way to palate. The bottleneck is the data, not the parameters.
System size vs. agreement for Nez3-VL at 4B, 8B, and 32B. Scaling the system leaves accuracy flat between 51% and 54%.
Share

One in ten wines carries a marker the spec never asked for

Reading wine isn't the only place the estates slip. While they ranked the pours, the same tasters flagged every glass on the overall-preference tracks for off-spec markers, meaning notes that drifted from the spec or had nothing to do with it. Across 1,600 flags per cohort, about 55% came back clean, 35% minor, and 10% major. One in ten finished wines carried a major off-spec marker, something the spec never asked for. These are the kinds of failures a taster catches at a sniff and the cellar that made them does not.

1 in 10 finished wines carried a major off-specmarker. Something the spec never asked for.1 in 10finished wines carried amajor off-spec markerMajor marker10% · 160 flagsMinor marker35% · 560 flagsNo off-spec marker55% · 880 flagsOff-spec marker severity across 1,600 flags per cohort
Off-spec flags across 1,600 pours per cohort: 55% clean, 35% minor, 10% major.
Share

Train on the data, and half the gap disappears

Then comes the turn. We trained a small pairwise-difference head on top of a frozen sensory encoder, with no fine-tuning of the backbone, a deliberately modest judge, directly on Palate Crit.

It reaches 0.611 agreement with tasters. That closes roughly 46% of the entire gap between a coin flip (0.500) and the human ceiling (0.741), and it's the first configuration in our sweep to clear the noise floor that standard regularisation couldn't budge. The lesson from the benchmark holds in reverse. The signal was always there. It just had to be trained on the right data instead of borrowed from show-bottle preference.

×LicaRun it throughPalate Crit…Half thegap, gone.A small judge trained on Palate Critjumps halfway to human-level.Palate is learnable.chancehuman-levelHuman taster0.741Trained on Palate Crit0.611Off-the-shelf judge0.543Random0.500closes ~46%of the gap
A small judge trained on Palate Crit reaches 0.611, closing about 46% of the gap between chance and the human ceiling.
Share

On the genuinely hard calls, it already ties a human

Roughly half of all pairwise comparisons are 3-2 splits, cases where the tasters themselves are nearly evenly divided and even a perfect predictor is partly guessing. Those are the calls that actually test judgment.

On exactly those cases, a judge trained on Palate Crit scores 0.602 against a human ceiling of 0.600. When the panel splits three to two, even one taster only agrees with the majority three times in five, and the judge now matches that. The gap to human agreement stays wide on the cases tasters find easy. On the ones they find hardest, it shrinks to almost nothing.

On the hard calls, the judge ties a humanOn 3-2 taster-split cases, the trained judge scores 0.602 vs 0.600 human ceiling.0.500.520.540.560.580.600.620.640.660.602Trainedjudge0.600Humanceiling+0.002 apart · essentially tiedcoin flipAgreement on hard calls
On the hardest 3-2 splits, the Palate Crit judge (0.602) matches the human ceiling (0.600).
Share

Why it matters

Palate Crit (Criteria-Resolved Impression Tasting) lets you build a decision layer for wine buying. Its criterion-level structure means you can route between estates by what a listing actually needs. Pick the estate that's strongest on structure for a cellar anchor, or on hue accuracy for a by-the-glass rosé, rather than trusting one aggregate score. The same structure works as supervision for training preference judges and reward models that optimise for specific sensory dimensions rather than a blurry average.

The headline finding is blunt. Estates can make the wine, but nothing on the market can yet reliably tell good wine from bad, and no amount of scale fixes that on its own. The hopeful finding is that the missing signal, palate, is real and can be learned from expert data. That's the layer Oenra's network is built to provide.

Our first Palate Crit dataset “PALATE” is published in full here: the tasting note dataset

Limitations

The sample is small. Each lot was rated by five tasters, enough to measure agreement and rule out noise but not enough to be sure of any single comparison. Each criterion used its own set of 80 lots, so no wine was ever rated on two criteria at once. That kept each rating clean, but it means we can't see how one taster weighs colour against structure on the same wine, because no one judged the same wine on both. Every flight was poured in one region's glassware, so palate across serving cultures isn't captured here. And the nine criteria cover a lot, but not everything. Ageing potential, house consistency, food fit, and cellar-door value are all natural axes to add as the work grows.

Future research

The obvious next move is to run this at a wider scale, with more tasters per lot, more regions, and the criteria expanded. Rating the same wines on every axis might reveal how tasters balance colour against structure or typicity against feel, the trade-offs a single score hides.

So far we've shown the signal can be learned as a judge. The open question is whether it makes cellars better winemakers. Training a house against these per-dimension rewards could push on structure or colour directly, then show whether the wine actually improves.

How we ran this study → Methodology
Continue reading3 studies
All research
  1. May 21, 2026Field Note
    At Cellier Blanc, your first pour decides the whole ceiling.5 tasters, 5 openings, 1 grand cru flight. The first pour set what each session could reach.Read
  2. May 18, 2026Field Note
    With Clos Bertaud 2, "Fruit is solved." Structure isn't.42 sessions, 7 senior tasters. Clos Bertaud 2 nails the overall structure, then breaks on the individual components.Read
  3. May 5, 2026Field Note
    Tasters keep telling us the same thing about scale: every bottle tastes the same.12 estates, 5 wine categories. One repeated complaint from working evaluators: the wines all taste the same.Read