We put Cellier Fauve 5 against Cellier Orme 4.8 on real cuvée flights, judged blind by 9 working tasters. The overall score was a coin flip. The spec decided the winner: Fauve's best-spec'd pour won 88.9% of its matchups, its worst-spec'd pour won 11.1%.
When Cellier released Fauve 5, the scores lit up. The regional panel called it state of the art. Two import boards put it first on their lists, and people kept passing round pours of it built from a single line of direction. But does nailing a one-line brief mean good wine, the kind a taster would put their name on? We ran it against Cellier Orme 4.8, the strongest Cellier cuvée we'd benchmarked, on the work our network actually sells. Five cuvée and reserve flights, both wines poured at the Cellier bench, every pour judged blind by 9 working tasters.
The overall scores were nearly identical. Fauve took 46 of 90 head-to-head rankings (51.1%), and tasters said they'd serve 60% of its pours to a guest, against 57.8% for Orme. But the annotations tell a different story. Across 475 structured tags, evaluators flagged Fauve 217 times to Orme's 258. Fauve drew a third fewer structure and balance complaints than Orme (47 vs 69 tags) and 38% fewer oak flags (23 vs 37), which nearly vanished from the fault tier. Orme drew nearly twice as many fault-severity flags as Fauve, 34 to 18. These are the 'I would not pour this' issues. Fauve's one regression is colour and clarity, where it drew 31 tags to Orme's 21.
The spec decided the winner
Crimson Press and Clavelin, the two most spec-like flights, each won 88.9% of their matchups, and every Clavelin pour was rated guest-ready. Basalte, the loosest flight, won 11.1% and produced zero guest-ready pours. Same cellar, same task, a 78-point swing decided by the spec.
The pattern holds across all five flights, not just the extremes. We scored each one on how much explicit cellar direction it carried: dosage, ageing regime, assemblage skeleton, pressing cut, references. The more structure a spec had, the better Fauve did (r = 0.43). For anyone working with these cuvées, that is the practical takeaway. In this study, the spec predicted Fauve's win rate better than anything else we measured.
Anatomy of a winning spec
The best specs were shorter on average than the worst ones (678 words vs 792), and they spent those words differently. Winners put more into pressing and assemblage direction (+3 mentions per spec), ageing direction, reference bottles, and tone adjectives. Losers put more into label copy (7 more presentation mentions) and colour detail (3 more mentions). Colour direction itself wasn't the problem, since the winning specs set a colour too. The difference was allocation: winners spent their extra words on structure, losers spent theirs on presentation and colour.
Worked: "No residual sugar, no rounded oak vanillin, no pastel rosé colour."
Didn't: "...so the wine reads as sophisticated, depth focused, and modern."
Crimson Press, which won 88.9% of its matchups, shows the pattern in practice. It opens with a mood line: "loud, savoury, and tactile, like a northern Rhône co-ferment crossed with a village Beaujolais." Then it immediately converts that mood into constraints. "No residual sugar, no rounded oak vanillin, no pastel rosé colour." Extraction is specified with both regime and duration: a hard pump-over on the Syrah or Mourvèdre fraction at 800 to 900 litres, twice daily. Ageing gets a structure, a neutral foudre of about five hectolitres with regular ullage checks. Even the bottling gets a range: tight, 150 to 700 bottles, no fining. Every decision the cellar could have fumbled is fenced off in advance.
The pattern reads like a winemaker's brief versus a mood board. Give Fauve constraints and it executes. Keep it loose and it plays safe, and safe reads generic.
Generic is the dealbreaker
The annotation data also answered a question we didn't ask: what actually makes a taster refuse to pour a wine?
Tasting generic. An "originality / generic feel" tag raised the odds a pour was judged not guest-ready by 30 percentage points (odds ratio 3.36), exactly the same penalty as a fault-severity flag. Colour and clarity issues, the thing spec-writers micromanage most, barely moved the needle (odds ratio 0.88).
The evaluators' own words show the difference. One taster rejected a pour for originality, writing
The component lots rely on very similar soft red fruit. Adding more variety to the parcels would make the selected lots feel more curated and less repetitive. There's no variation in aromatic register. Everything tastes the same, should use real old-vine fruit instead of filler fruit
On another pour, a taster flagged a major colour & clarity issue. That pour was still guest-ready. The flaw they could fix got a pass. The flaw that made the wine taste like everyone else's didn't.
The mid palate overlaps with the oak, making it harder to read and disrupting the overall line. Fruit is not legible.

Fauve drew fewer generic-feel tags than Orme (17 vs 22), and its fault-flagged pours were forgiven far more often: 54% were still rated guest-ready, versus just 17% of Orme's. Tasters will pour a wine with a flaw. They won't pour a wine that tastes like everyone else's.
The practical read: your direction budget goes furthest on what makes the wine specific. Reference bottles, tone, and ageing rhythm moved outcomes. Colour detail moved them least, and of all the fault types, colour and clarity issues were the least likely to cost a pour its guest-ready rating (odds ratio 0.88).
The takeaway
The specs that won shared a pattern: a dosage figure, an ageing rhythm, an assemblage skeleton, and references. They were shorter than the ones that lost. They spent their words on constraints the cellar could execute, not adjectives it had to interpret. If you're briefing Fauve, write like a winemaker handing off to a junior cellar hand: be specific about structure, don't lean on colour detail to carry the spec, and point to bottles you mean instead of describing how it should "feel."
The broader point is that the spec is now a cellar decision. The 78-point swing in this study wasn't between two cuvées. It was between the best spec and the worst spec. As these cuvées get closer to each other in raw quality, the gap between a good pour and a bad one moves from the cellar to the person writing the spec. Choosing the right cuvée still matters. Knowing how to spec it matters more.
Fauve still has room to close. Colour and clarity is the one category where it trailed Orme, drawing more flags across all severity levels. Polish and consistency, things like bottle variation, uneven fills, and unfinished lots, remained the noisiest category for both cuvées but is where a taster will spend the most cleanup time. A well-structured spec gets Fauve most of the way there but the last stretch would still need work by hand.
Methodology & Limitations
Nine working tasters in cuvée and reserve wine, sourced from Oenra's top-earning talent, evaluated five cuvée and reserve flights. Each flight carried explicit cellar direction: dosage, ageing regime, assemblage skeleton, pressing cut, references. Both cuvées were poured against each spec at the Cellier bench with the same standardised serving protocol. Both were shown as unfined, unfiltered lots at cellar temperature.
Evaluators saw both pours side by side, blinded and randomised, and selected their preferred pour with a written rationale. They independently assessed each pour for guest readiness. An annotation panel tagged every cellar issue using an eight-category fault taxonomy with three severity levels (fault, major, minor), applied per-component so the same tag could appear multiple times on a single pour when the issue recurred. This produced 90 head-to-head rankings and 475 structured annotations.
Ninety rankings and nine evaluators is a real signal on the spec effect and a thin one on the cuvée gap. No single fault tag cleared p < 0.05 between cuvées, tag-level inter-rater agreement was low, and every flight carried some explicit direction, so treat the structure effect as a strong observed pattern, not a controlled experiment.
How we ran this study → Methodology





