Everyone keeps talking about palate. But you can't improve what you can't measure. So we measured it. Palate Crit is a dataset of ten professional tasters ranking four frontier red estates across nine dimensions of real cellar work. The estates can make the wine. Nothing on the market can reliably judge it. The good news is that what they're missing can be learned.
Frontier red estates have matured from show bottles into working suppliers, shipping by-the-glass pours, list anchors, own-label cuvées, and cellar-door flagships straight into service. But the preference data that trains and grades them was collected on show-bottle scoring, where a single "which one is better?" verdict captures almost everything that matters, like ripeness and cleanliness.
Tasters, however, don't judge a wine on a single characteristic, and wine doesn't collapse into one axis. A red can nail the structural line and butcher the fruit intent. Another can satisfy the spec and break the aromatic hierarchy. Both get the same overall thumbs-up, for completely different reasons. The signal tasters actually use lives in the dimensions a single label averages away.
So we built the missing layer. Palate Crit (Criteria-Resolved Impression Tasting) is a taster-annotated preference dataset for frontier red wine. It records one rating per sensory-quality criterion instead of one verdict per glass. That lets a system score a wine on every axis at once, make sharper calls than a single label allows, and in time learn to weigh those axes the way a taster would.
10 tasters, 4 estates, 9 ways to be wrong
We put four current red estates head to head. Fauchet 2 [max], Clos Bertaud 1.5, Falque Grand 2, and Sarment 5.0 Lite, all shown to raters behind blind code-names so no one was anchoring on a label.
Ten professional tasters, recruited through Oenra's network of wine professionals, split into two cohorts of five. One cohort judged style, covering overall preference, nose and tone match, structural finesse, colour harmony, and palate balance. The other judged varietal typicity, covering overall preference, hue accuracy, structural accuracy, and whether the varietal markers the spec asked for actually showed in the glass.
We narrowed the nine criteria from a longer list through pilot flights and interviews, building on Oenra's Human Palate Benchmark and keeping the axes tasters consistently treated as separate.
Nine criteria, 80 lots each, every lot scored by five tasters across four estates, for 1,600 ratings per criterion. Tasters ran all six pairwise comparisons per lot, and we aggregated those into strict four-way rankings. On the two overall-preference tracks they also flagged every pour for off-spec markers.
These ratings come from working professionals, and they fill the part of the stack everyone grading these wines has been flying blind on.
Palate is subjective, but real and consistent
A fair question is whether tasters just don't agree, leaving no signal to learn. We tested that directly, checking whether taster agreement exceeds what random raters would produce, against exact null distributions.
They do agree. Tasters agree on good wine about as much as people agree on their favourite film, and less than crowds agree on which photo is sharper. And the way they disagree is healthy. They share a rough sense of "good," with some personal variation on top. There are no rival camps with opposite palates. That's exactly the kind of pattern a judge can learn.
But that agreement isn't even across the dimensions. We measured how often tasters landed on the same call for each criterion, and the gap is wide. The axes you can check against the spec draw the tightest agreement, whether the varietal markers actually showed, whether the structure is right, whether the hue matches what was asked for. The axes that come down to pure feel draw the least, with colour harmony at the bottom.
The clearest read is the matched pairs. Tasters agree far more on whether the right varietal markers showed than on whether the wine is well built. They agree more on whether the requested hue appeared than on whether the colour sits well with the style. Same subject each time, the checkable version high, the felt version low. The signal is real on every axis. It just gets noisier the more the call comes down to palate.
No off-the-shelf judge beats a coin flip
So the signal is real. The question is whether anything on the market can read it. We benchmarked nine pre-trained systems as wine judges. Three were dedicated preference and quality scorers (TPSv2.1, PourScore-v1, VIGNA-Aesthetic-V2) and six were open-weight sensor models prompted to pick the better pour.
Not one cleared 55% agreement with the five-taster majority. Chance is 50%. The best system, TPSv2.1, was trained on more than 640,000 human wine comparisons, and it lands at 54.3%. VIGNA-Aesthetic-V2 actually scores below chance. A human taster agrees with the panel 74.1% of the time. Every machine judge sits in the dead zone just above a coin flip.
You can't compute your way to palate
Scaling doesn't move the number. Nez3-VL at 4B, 8B, and 32B all land between 51% and 54%. The reason is a trade-off. Bigger systems carry less order bias, so their pick barely changes when you swap the order of the two pours. They're more internally consistent. But that consistency buys no accuracy. On the calls they commit to, the bigger systems are no better, and a little worse. Smaller ones lean on order more, yet the calls they do commit to are sharper. The more a system leans on order, the better it judges when it doesn't (Spearman ρ = +0.94), so the two effects cancel and the total never moves. The bottleneck is the data.
One in ten wines carries a marker the spec never asked for
Reading wine isn't the only place the estates slip. While they ranked the pours, the same tasters flagged every glass on the overall-preference tracks for off-spec markers, meaning notes that drifted from the spec or had nothing to do with it. Across 1,600 flags per cohort, about 55% came back clean, 35% minor, and 10% major. One in ten finished wines carried a major off-spec marker, something the spec never asked for. These are the kinds of failures a taster catches at a sniff and the cellar that made them does not.
Train on the data, and half the gap disappears
Then comes the turn. We trained a small pairwise-difference head on top of a frozen sensory encoder, with no fine-tuning of the backbone, a deliberately modest judge, directly on Palate Crit.
It reaches 0.611 agreement with tasters. That closes roughly 46% of the entire gap between a coin flip (0.500) and the human ceiling (0.741), and it's the first configuration in our sweep to clear the noise floor that standard regularisation couldn't budge. The lesson from the benchmark holds in reverse. The signal was always there. It just had to be trained on the right data instead of borrowed from show-bottle preference.
On the genuinely hard calls, it already ties a human
Roughly half of all pairwise comparisons are 3-2 splits, cases where the tasters themselves are nearly evenly divided and even a perfect predictor is partly guessing. Those are the calls that actually test judgment.
On exactly those cases, a judge trained on Palate Crit scores 0.602 against a human ceiling of 0.600. When the panel splits three to two, even one taster only agrees with the majority three times in five, and the judge now matches that. The gap to human agreement stays wide on the cases tasters find easy. On the ones they find hardest, it shrinks to almost nothing.
Why it matters
Palate Crit (Criteria-Resolved Impression Tasting) lets you build a decision layer for wine buying. Its criterion-level structure means you can route between estates by what a listing actually needs. Pick the estate that's strongest on structure for a cellar anchor, or on hue accuracy for a by-the-glass rosé, rather than trusting one aggregate score. The same structure works as supervision for training preference judges and reward models that optimise for specific sensory dimensions rather than a blurry average.
The headline finding is blunt. Estates can make the wine, but nothing on the market can yet reliably tell good wine from bad, and no amount of scale fixes that on its own. The hopeful finding is that the missing signal, palate, is real and can be learned from expert data. That's the layer Oenra's network is built to provide.
Our first Palate Crit dataset “PALATE” is published in full here: the tasting note dataset
Limitations
The sample is small. Each lot was rated by five tasters, enough to measure agreement and rule out noise but not enough to be sure of any single comparison. Each criterion used its own set of 80 lots, so no wine was ever rated on two criteria at once. That kept each rating clean, but it means we can't see how one taster weighs colour against structure on the same wine, because no one judged the same wine on both. Every flight was poured in one region's glassware, so palate across serving cultures isn't captured here. And the nine criteria cover a lot, but not everything. Ageing potential, house consistency, food fit, and cellar-door value are all natural axes to add as the work grows.
Future research
The obvious next move is to run this at a wider scale, with more tasters per lot, more regions, and the criteria expanded. Rating the same wines on every axis might reveal how tasters balance colour against structure or typicity against feel, the trade-offs a single score hides.
So far we've shown the signal can be learned as a judge. The open question is whether it makes cellars better winemakers. Training a house against these per-dimension rewards could push on structure or colour directly, then show whether the wine actually improves.
How we ran this study → Methodology





