We pulled every pour five working tasters called at Cellier Blanc across one grand cru flight. The data points the same direction in every session: the first pour matters more than the fifteen that follow it.
Our first piece on Cellier Blanc tested the wine. The house delivered the first blend fast, then stumbled when tasters tried to push it further. Balance, colour, and aromatics were the consistent failures, and tasters settled on 40% at Cellier, the rest on the bench.
This piece moves from lot to call. Every pour was coded for structure and intent. What follows: how five different openings produced five different ceilings, where regression starts, and why one taster called sixteen pours and finished worse than they started.
The first pour
Every session reached a moment where the house system was built, the fractions were assembled, and the taster had to call the first thing on a blend that didn't exist yet. The cellar already had context: a house, a colour, a fraction library the taster had just approved. The flight, reference samples, and inspiration sat alongside it. What tasters called was the first instruction on top of all of that.
The shape is what carries the story. The opening is doing two things at once: a brief that sets context, buyer, and style, and a directive that names what to build, what fractions to use, and what to avoid. That hybrid shape was the most specific opening of the five. It's one of four shapes tasters used to open their sessions, and the variation in those shapes is where the story starts.
The five openings
The Aetheon call above was Participant 5's, a Brief-Then-Directive Hybrid at 263 words. We scored each opening for specificity on a 0–1 scale: how many concrete cellar decisions the pour fixed in place versus left to the house. Participant 5's opening scored 0.75, the highest of the five, and was the only one to specify sensory values from the house system the taster had just approved.
Participant 4 wrote nearly as much, 242 words, organised as a Comprehensive Cellar Brief with every category a trade document would include: role, overview, objective, buyer, fractions, sensory system, structure, style, and constraints. Participant 3 cut that scope in half: 161 words, a Context-Action Framework with vintage, buyer, pain points, and an assembly directive, but no sensory or structural specifics.
The other two openings looked nothing like briefs. Participant 2 wrote twelve words: a single corrective directive asking the house to swap a fraction in what was already in the glass. Participant 1 wrote one word, an affirmation: the Minimal Confirmation, accepting what the house had produced and moving on.
What tasters chose to include tells a clearer story than how specific they were. Every one of them included a flight brief or task instruction in some form. Most named a target buyer and a goal. But across all five openings, only Participant 5 specified sensory values (exact references to the house system they'd already built). The rest assumed the cellar would carry the system over. None of the five provided analytical figures or extraction targets in their first pour, and none asked for a specific bottling format. The most consistent gap, even in the most comprehensive briefs, was the same one the first article identified from the other side: precision the house was being asked to produce without being given any of its specifics.
The five pour categories describe what kind of move a taster made on any given turn. Assembly builds something new. Additive adds to what exists. Refinement adjusts. Corrective fixes what broke. Meta asks the house to step back and reconsider.
No two tasters used the same mix. Participant 4, who had written the second-most-comprehensive opening, lived almost entirely in refinement, five pours adjusting what the house had produced. Participant 5, the hybrid opener, spread their session across every category, with a notable cluster of corrective pours. Participant 1, the Minimal Confirmation opener, never called a refinement at all; their session was a small number of corrective and assembly moves and nothing else. Participant 2 made one corrective and stopped.
Two patterns hold across the variation. Refinement is the most common category overall, fourteen pours across the five tasters. Corrective is second, at eleven. Together they account for roughly two-thirds of everything anyone called. Assembly, by contrast, was almost entirely a first-pour behaviour. Only one taster called an assembly pour anywhere except the opening. The first pour was the only time most tasters stepped back and built; after that, they were adjusting and fixing what the house returned.
Did opening well help?
The two tasters with the strongest opening moves (Participant 4's Comprehensive Cellar Brief and Participant 5's Brief-Then-Directive Hybrid) reached the two most finished outcomes, cellar sample and bottling-ready. The Minimal Confirmation opener landed in trial. The Iterative Fraction Swap landed at cellar sample, despite a twelve-word opening, by working with what the house had already produced rather than rebuilding from scratch.
The pattern is directional. More structure in the opening correlated with more finished lots, with one exception. Participant 3 opened with a real assembly pour (middle of the pack on specificity, a real Context-Action Framework) and ended in trial, the same outcome as the taster who'd opened with a single word. The contrast shows up in the glass itself.
The session that proves it
Participant 3 is the most informative session we ran. The opening predicted a cellar sample or better. The session ended in trial.
Across the session, Participant 3 called sixteen pours, more than three times as many as any other taster. They kept going. They got more specific over time. Their average pour by mid-session was more precise than their opening.
The red line is lot quality, scored across each round on a −1 to +1 scale. It sits below zero for twelve of the sixteen pours. The peak was a single refinement at round four, just above +0.2. The trough was pour two, a meta pour asking the house to step back and reconsider, which pulled quality down to −0.7. The cellar treated it as an instruction to discard: the structure from the opening pour was lost, and the session never fully recovered. Every subsequent corrective was a fight to rebuild ground that had been there a pour earlier.
Look at the pour-role labels along the bottom. Eight of the sixteen pours are corrective. Half of this taster's session was spent telling the house to fix what it had broken. Refinements made up most of the rest. The only pours that produced positive quality scores at the end were two additives, small accretions, late in the session, after the structural fight was already lost.
This is the regression story from the first article, observed at the resolution of a single session. Participant 3 kept going for sixteen rounds, climbing in specificity, and the quality score never held positive territory until the last two pours, by which point the wine had drifted far from the brief. The opening pour sets the ceiling. What comes after determines whether the ceiling holds. In Participant 3's session, it didn't.
Verdict
The opening pour is the most consequential move in a Cellier Blanc session. The two tasters who opened with structured briefs reached the two most finished lots; the taster who opened with one word ended in trial. The correlation isn't subtle.
But structure at the start doesn't survive iteration on its own. Refinement and corrective pours made up roughly two-thirds of everything tasters called, and sessions regressed through those moves, even when tasters kept trying. Participant 3 called sixteen pours, got more specific over time, and finished worse than they started.
Specify sensory values in the opening, even when the house system is already built. Only one taster did this, and it's the one who reached bottling-ready. The system you approved is not context the cellar can be trusted to read on its own. Be cautious with meta pours. The single largest quality drop in the study came from a pour asking the house to step back. Treat “reconsider” as a request to discard. Assembly belongs at the start. Only one taster called an assembly pour anywhere except the opening. After the first move, the session is a refinement and correction game. Plan accordingly.
The open question is whether the regression we observed is a property of the house or a property of how tasters learn to call it. A second round, with the same tasters and a different flight, would start to answer that.
Methodology
Five tasters each ran a single session at Cellier Blanc against the same cellar brief: a blend for a fictional ultra-premium upland estate called Aetheon. Each taster was given the brief, reference samples, and candidate colour targets before the session began. They worked through Cellier Blanc's standard flow (entering a house name, submitting reference samples, generating a house system, and reviewing the fractions Cellier Blanc produced) before calling their first pour against the blend itself.
Sessions were recorded end-to-end using Rollout, Oenra Labs' session-capture tool. Rollout records the bench, video, audio, pour trails, decants, and note input, including everything tasters did inside and outside Cellier Blanc.
We collected every pour across every session and coded each one for structure and role against a fixed rubric applied by two Oenra analysts. First pours were classified into five structural shapes: Minimal Confirmation, Iterative Fraction Swap, Context-Action Framework, Comprehensive Cellar Brief, and Brief-Then-Directive Hybrid. All pours across the session were classified into five categories of intent: assembly (build something new), additive (add to what exists), refinement (adjust), corrective (fix what broke), and meta (step back and reconsider). Each taster's final lot was self-rated for cellar readiness on a four-point scale: restart, trial, cellar sample, and bottling-ready.
Limitations
Feedback was self-reported from the same tasters who produced the wine, carrying the usual self-assessment bias. The sample is small, and we treat these findings as directional. Only one flight was tested, so brief-specific effects can't be cleanly separated from the house's general behaviour. The patterns, however, hold across tasters and across adjustment rounds in the same direction: a strong system-level start, weak precision execution under iteration.
How we ran this study → Methodology





