Battle·July 7, 2026·5 min read

Give Fauve a dosage spec and it wins 9 times out of 10

Cellier Fauve 5 vs Cellier Orme 4.8 on 5 real cuvée and reserve flights, poured at the Cellier bench, judged blind by 9 working tasters.

  1. 01Fauve took 46 of 90 head-to-head rankings (51.1%), a coin flip against Orme 4.8.
  2. 02The best-spec'd pour won 88.9% of its matchups; the worst won 11.1%.
  3. 03Fauve's advantage grew with how much explicit cellar direction the spec carried (r = 0.43).
  4. 04A generic-feel tag cut a pour's guest-ready odds by 30 points, the same penalty as a fault-level flag.
  5. 05Fauve drew a third fewer structure and balance complaints than Orme, but more colour and clarity flags.

We put Cellier Fauve 5 against Cellier Orme 4.8 on real cuvée flights, judged blind by 9 working tasters. The overall score was a coin flip. The spec decided the winner: Fauve's best-spec'd pour won 88.9% of its matchups, its worst-spec'd pour won 11.1%.

Give Fauve a dosage specand it wins 9 times out of 10Cellier Fauve 5 vs Cellier Orme 4.8 · blind cellar-bench tastingJudged blind by 9 working tasters5flights90rankings475annotations

When Cellier released Fauve 5, the scores lit up. The regional panel called it state of the art. Two import boards put it first on their lists, and people kept passing round pours of it built from a single line of direction. But does nailing a one-line brief mean good wine, the kind a taster would put their name on? We ran it against Cellier Orme 4.8, the strongest Cellier cuvée we'd benchmarked, on the work our network actually sells. Five cuvée and reserve flights, both wines poured at the Cellier bench, every pour judged blind by 9 working tasters.

The overall scores were nearly identical. Fauve took 46 of 90 head-to-head rankings (51.1%), and tasters said they'd serve 60% of its pours to a guest, against 57.8% for Orme. But the annotations tell a different story. Across 475 structured tags, evaluators flagged Fauve 217 times to Orme's 258. Fauve drew a third fewer structure and balance complaints than Orme (47 vs 69 tags) and 38% fewer oak flags (23 vs 37), which nearly vanished from the fault tier. Orme drew nearly twice as many fault-severity flags as Fauve, 34 to 18. These are the 'I would not pour this' issues. Fauve's one regression is colour and clarity, where it drew 31 tags to Orme's 21.

Structure & polish dominate faults; Orme 4.8 logs more structure issues (69 vs 47)Bars: tag assignments split by severity · paired rows: Fauve vs Orme 4.8Polish & consistency41734Fauve · n=5782133Orme 4.8 · n=63Balance, acidity & structure61919Fauve · n=47726324Orme 4.8 · n=69Oak & élevage515Fauve · n=23411184Orme 4.8 · n=37Colour & clarity1512Fauve · n=31513Orme 4.8 · n=21Originality / generic feel78Fauve · n=17697Orme 4.8 · n=22Varietal fit & tone515Fauve · n=21115Orme 4.8 · n=19Bead & mousse44Fauve · n=1158Orme 4.8 · n=15Cues & length54Fauve · n=107Orme 4.8 · n=1201020304050607080Number of tag assignmentsSeverityFaultMajorMinorNo severity475 structured annotations across 5 flights · tags are applied per component, so one pour can carry a tag more than once.
Fault tags by cuvée across 475 structured annotations. Evaluators flagged Fauve 217 times to Orme's 258.
Share

The spec decided the winner

Crimson Press and Clavelin, the two most spec-like flights, each won 88.9% of their matchups, and every Clavelin pour was rated guest-ready. Basalte, the loosest flight, won 11.1% and produced zero guest-ready pours. Same cellar, same task, a 78-point swing decided by the spec.

The pattern holds across all five flights, not just the extremes. We scored each one on how much explicit cellar direction it carried: dosage, ageing regime, assemblage skeleton, pressing cut, references. The more structure a spec had, the better Fauve did (r = 0.43). For anyone working with these cuvées, that is the practical takeaway. In this study, the spec predicted Fauve's win rate better than anything else we measured.

Fauve’s pairwise advantage grows with spec direction/structureDots: spec groups (size = sessions) · diamonds: group means · dashed: weighted fitLowadv -0.56Mediumadv +0.56Highadv +0.060.800.600.400.20+0.00+0.20+0.40+0.60+0.8012.612.813.013.213.413.613.814.0Spec direction / structure scoreFauve advantage (Fauve win rate − Orme 4.8 win rate)Low directionMedium directionHigh directionGroup meanWeighted fit (slope=+0.504)
Spec structure score vs Fauve's head-to-head advantage. The more explicit cellar direction a spec carried, the better Fauve did (r = 0.43).
Share

Anatomy of a winning spec

The best specs were shorter on average than the worst ones (678 words vs 792), and they spent those words differently. Winners put more into pressing and assemblage direction (+3 mentions per spec), ageing direction, reference bottles, and tone adjectives. Losers put more into label copy (7 more presentation mentions) and colour detail (3 more mentions). Colour direction itself wasn't the problem, since the winning specs set a colour too. The difference was allocation: winners spent their extra words on structure, losers spent theirs on presentation and colour.

Worked: "No residual sugar, no rounded oak vanillin, no pastel rosé colour."
Didn't: "...so the wine reads as sophisticated, depth focused, and modern."

Crimson Press, which won 88.9% of its matchups, shows the pattern in practice. It opens with a mood line: "loud, savoury, and tactile, like a northern Rhône co-ferment crossed with a village Beaujolais." Then it immediately converts that mood into constraints. "No residual sugar, no rounded oak vanillin, no pastel rosé colour." Extraction is specified with both regime and duration: a hard pump-over on the Syrah or Mourvèdre fraction at 800 to 900 litres, twice daily. Ageing gets a structure, a neutral foudre of about five hectolitres with regular ullage checks. Even the bottling gets a range: tight, 150 to 700 bottles, no fining. Every decision the cellar could have fumbled is fenced off in advance.

The pattern reads like a winemaker's brief versus a mood board. Give Fauve constraints and it executes. Keep it loose and it plays safe, and safe reads generic.

Generic is the dealbreaker

The annotation data also answered a question we didn't ask: what actually makes a taster refuse to pour a wine?

Tasting generic. An "originality / generic feel" tag raised the odds a pour was judged not guest-ready by 30 percentage points (odds ratio 3.36), exactly the same penalty as a fault-severity flag. Colour and clarity issues, the thing spec-writers micromanage most, barely moved the needle (odds ratio 0.88).

The evaluators' own words show the difference. One taster rejected a pour for originality, writing

The component lots rely on very similar soft red fruit. Adding more variety to the parcels would make the selected lots feel more curated and less repetitive. There's no variation in aromatic register. Everything tastes the same, should use real old-vine fruit instead of filler fruit
CRIMSON PRESS · FAULT, NOT GUEST-READYOrmeRejected on originalityCrimson Press won 88.9% of its matchups: loud, savoury andtactile, then fenced off — no residual sugar, no roundedoak vanillin, no pastel rosé colour.WHY THE TASTER WOULD NOT POUR IT“The component lots rely on very similar softred fruit. There’s no variation in aromaticregister. Everything tastes the same, shoulduse real old-vine fruit instead of fillerfruit”
Orme, fault, not guest-ready pour (Crimson Press)
Share

On another pour, a taster flagged a major colour & clarity issue. That pour was still guest-ready. The flaw they could fix got a pass. The flaw that made the wine taste like everyone else's didn't.

The mid palate overlaps with the oak, making it harder to read and disrupting the overall line. Fruit is not legible.
Fauve, major severity, guest-ready pour (Perle)
Share

Fauve drew fewer generic-feel tags than Orme (17 vs 22), and its fault-flagged pours were forgiven far more often: 54% were still rated guest-ready, versus just 17% of Orme's. Tasters will pour a wine with a flaw. They won't pour a wine that tastes like everyone else's.

One fault-severity tag more than doubles the No rateRisk difference shown in pp · OR: odds ratio · ★ marks grouped severity flagsAny major or fault tag (n=73)Δ +31 ppOR 3.98Any fault-severity tag (n=27)Δ +30 ppOR 3.37Originality / generic feel (n=28)Δ +30 ppOR 3.36Varietal fit & tone (n=31)Δ +15 ppOR 1.86Balance, acidity & structure (n=72)Δ +12 ppOR 1.65Cues & length (n=18)Δ +11 ppOR 1.55Any aroma or colour tag (n=73)Δ +10 ppOR 1.53Aroma definition (n=41)Δ +9 ppOR 1.44Any drinkability tag (n=80)Δ +8 ppOR 1.34Bead & mousse (n=20)Δ +5 ppOR 1.23Polish & consistency (n=66)Δ +4 ppOR 1.18Colour & clarity (n=46)Δ 4 ppOR 0.880%10%20%30%40%50%60%70%80%Share of pours judged NOT guest-readyNo-rate when severity flag presentNo-rate when absentNo-rate when tag presentRisk difference (present − absent)
Impact of each fault type on a pour's guest-ready odds. A generic-feel tag carried the same penalty as a fault-severity flag; colour and clarity barely registered.
Share

The practical read: your direction budget goes furthest on what makes the wine specific. Reference bottles, tone, and ageing rhythm moved outcomes. Colour detail moved them least, and of all the fault types, colour and clarity issues were the least likely to cost a pour its guest-ready rating (odds ratio 0.88).

The takeaway

The specs that won shared a pattern: a dosage figure, an ageing rhythm, an assemblage skeleton, and references. They were shorter than the ones that lost. They spent their words on constraints the cellar could execute, not adjectives it had to interpret. If you're briefing Fauve, write like a winemaker handing off to a junior cellar hand: be specific about structure, don't lean on colour detail to carry the spec, and point to bottles you mean instead of describing how it should "feel."

The broader point is that the spec is now a cellar decision. The 78-point swing in this study wasn't between two cuvées. It was between the best spec and the worst spec. As these cuvées get closer to each other in raw quality, the gap between a good pour and a bad one moves from the cellar to the person writing the spec. Choosing the right cuvée still matters. Knowing how to spec it matters more.

Fauve still has room to close. Colour and clarity is the one category where it trailed Orme, drawing more flags across all severity levels. Polish and consistency, things like bottle variation, uneven fills, and unfinished lots, remained the noisiest category for both cuvées but is where a taster will spend the most cleanup time. A well-structured spec gets Fauve most of the way there but the last stretch would still need work by hand.

Methodology & Limitations

Nine working tasters in cuvée and reserve wine, sourced from Oenra's top-earning talent, evaluated five cuvée and reserve flights. Each flight carried explicit cellar direction: dosage, ageing regime, assemblage skeleton, pressing cut, references. Both cuvées were poured against each spec at the Cellier bench with the same standardised serving protocol. Both were shown as unfined, unfiltered lots at cellar temperature.

Evaluators saw both pours side by side, blinded and randomised, and selected their preferred pour with a written rationale. They independently assessed each pour for guest readiness. An annotation panel tagged every cellar issue using an eight-category fault taxonomy with three severity levels (fault, major, minor), applied per-component so the same tag could appear multiple times on a single pour when the issue recurred. This produced 90 head-to-head rankings and 475 structured annotations.

Ninety rankings and nine evaluators is a real signal on the spec effect and a thin one on the cuvée gap. No single fault tag cleared p < 0.05 between cuvées, tag-level inter-rater agreement was low, and every flight carried some explicit direction, so treat the structure effect as a strong observed pattern, not a controlled experiment.

How we ran this study → Methodology
Continue reading3 studies
All research
  1. June 29, 2026Battle
    Frontier white estates are nearly tied. None of them nail minerality yet.A blind head-to-head of Saveline 2.0, Grès d'Ivraie, Véron 3.1, and Fournier Grand Blanc, judged by 12 professional tasters across 10 flights.Read
  2. June 3, 2026Battle
    Ombrelle v4 won 47.9% of structure matchups.10 tasters, 4 estates, 240 pours. Ripeness is solved. Structural craft and cellar-readiness are where Ombrelle v4 really pulls away.Read
  3. May 29, 2026Battle
    Corbière took 60% of head-to-heads. Cellier Cru took 63% of the cellar door.Four négociant houses, 24 lots, five working tasters. The house tasters preferred in the glass and the house they'd put their name on turned out to be different.Read