Data Recap · Episode 05
First Human Entry

Dark Chocolate Cherry Crunch Cookie

The first episode with a human recipe in the mix. I wrote the prompt based on my own cookie, put it in blind against three AIs, and came in fourth. ChatGPT won the room by the largest margin in series history. Claude won the math. Here's the full breakdown.

Tasters24
Scored Points216
Composite WinnerClaude
Room WinnerChatGPT
Borda Points71 / 96

The Scoreboard

Claude led on composite and had the highest smell average of any cookie in series history. ChatGPT led on taste and texture. Emily's smell score (6.34) was the lowest of the four — the widest gap between any cookie's smell and the leader in any episode.

Claude — The Chocolatier
Emily — The Contender
Gemini — The Complex
ChatGPT — The Frontrunner

Average scores by criterion

24 tasters · scored 1–10 on Taste, Texture, and Smell · Smell n=22 (two tasters could not score) · ChatGPT Taste n=23 (one blank)

Full results

CookieTasteTextureSmellCompositeBorda ptsRank 1Want again
Claude — The Chocolatier7.778.008.247.9456450%
Emily — The Contender7.217.716.347.0955558%
Gemini — The Complex6.947.177.147.1258667%
ChatGPT — The Frontrunner8.148.177.077.8271979%

How the Room Was Scored — Borda Count

In previous episodes, tasters picked one favorite. Starting Episode 5, they ranked all four cookies from first to fourth. A rank of 1 earns 4 points, rank 2 earns 3, rank 3 earns 2, rank 4 earns 1. Add those up across all 24 tasters and you get the Borda total. Max possible: 96 points.

Borda points — Episode 05

Max possible: 96 (all 24 tasters rank you first) · ChatGPT's 71 = 74% of maximum

ChatGPT71 pts · 74%
9 first-place picks
Gemini58 pts · 60%
6 first-place picks
Claude56 pts · 58%
4 first-place picks
Emily55 pts · 57%
5 first-place picks

Unlike a single favorite pick, Borda scoring captures how every taster felt about every cookie. A cookie that consistently finishes second accumulates more Borda points than one that wins a few people but alienates others. ChatGPT's 71 points is the strongest room performance of any cookie in this series. Claude's 50% want-again rate despite the highest composite score is the lowest of any cookie in any episode.

The Math and the Room Disagreed — Again

Claude won the composite score. ChatGPT won every democratic measure. This is the third episode where composite score and room preference pointed in different directions.

Composite Score
7.94
Claude
Borda Points
71
ChatGPT
First-Place Picks
9
ChatGPT · 38% of room
Want-Again Rate
79%
ChatGPT · 19 of 24

Per-Taster Composite Winners

Each tile shows the taster's age, gender, composite winner, and best score. Claude won the most individual matchups outright (5). There were 9 mathematical ties — the most of any episode — including one three-way tie and three Emily/ChatGPT ties that reflect how similar those two recipes were.

T1·78M
CL
7.67
T2·35F
CL
8.67
T3·14F
CH
9.33
T4·23F
CL/GE
8.33
T5·70F
CL/EM/GE
6.67
T6·69F
CL/CH
8.67
T7·47M
GE
8.00
T8·27M
EM/CH
8.00
T9·47F
CH
8.33
T10·20F
CL
9.00
T11·46M
CL/CH
10.00
T12·15F
CH
9.67
T13·78M
EM
7.33
T14·11M
CL
9.17
T15·18F
CL
8.33
T16·47F
GE
10.00
T17·39F
GE/CH
9.00
T18·35M
EM/CH
9.83
T19·47F
EM/CH
7.67
T20·46F
EM
7.67
T21·74F
CH
8.33
T22·45M
EM
8.00
T23·45F
GE/CH
10.00
T24·55F
GE
9.67
5
Claude wins
3
Emily wins
3
Gemini wins
4
ChatGPT wins
9
Tied

Cookie Spotlights

What the language data showed for each cookie — flavor identifications, descriptor words, and the signals that defined the tasting.

Claude — The Chocolatier
COMPOSITE WINNER

20 of 24 tasters identified chocolate in some form as the main flavor. The clearest single-ingredient identification in the series. The smell average (8.24) was the highest of any cookie in any episode — the cocoa, browned butter, and espresso all pulling in the same direction.

20/24 said chocolate
was the main flavor
HOW THEY NAMED IT
"chocolate"
15 tasters
"dark chocolate"
2
"cocoa"
1
"chocolatey"
1
"salt & chocolate"
1
4 tasters said other: dough · dried cherry · date · none
Emily — The Contender
HUMAN ENTRY · DEBUT

The hazelnut read as peanut butter for 12 tasters. The blondie base (no cocoa) meant the cookie looked and smelled different from the other three — one taster wrote "blond" as the flavor, which was a legitimately accurate read. "Bland" appeared twice.

POSITIVE (5)
scrumptious yummy delicious tasty good
NUTTY (5)
nutty ×3 peanut butta nuts
12 of 24 said nutty flavor
NEGATIVE (5)
bland ×2 plain normal ok

One taster identified "citrus" — the lemon zest registering. T17 wrote "nutty fruity (can't decide)" — likely hazelnut and lemon pulling in opposite directions.

Gemini — The Complex
WORST EPISODE YET

The only cookie where two flavors split the room. Chocolate and cherry both landed distinctly — but two tasters identified "raisin" instead of cherry, and one wrote "chocherry," unable to separate them. The demerara crust registered clearly as crunch without reading as overly sweet.

FLAVOR SPLIT — 24 TASTERS
8CHOCOLATE
BOTH
4CHERRY
11OTHER
CHOCOLATE (8)
chocolate ×5 · choc ×1
cocoa ×1 · brownie ×1
CHERRY (4)
cherry ×3 · fruit ×1
raisin ×2 (misidentification)
ChatGPT — The Frontrunner
ROOM WINNER

The most unanimously positive descriptor language of any cookie in any episode. 12 of 24 tasters used pure enthusiasm words. Cherry landed for 5 tasters. And one person nailed the recipe with two words.

"trail mix"
T1 · 78M · Most accurate descriptor of the episode Dark chocolate, dried tart cherry, toasted hazelnut. He named exactly what it was before seeing a single ingredient. He also ranked it last.
POSITIVE WORDS (12)
perfect ×2 yummy/yum ×4 exquisite delicious good ×2 flavorful homemade
CHERRY LANDED (5)
sour cherry · cherry ×2
lots of flavors–cherry · fruity/cherry
dry (1) · too much going on (1)

Demographics

24 tasters · 8M / 16F · Ages 11–78. The largest panel in series history. ChatGPT dominated across both gender and age groups. Emily's strongest showing was in the 25–50 bracket, where she tied ChatGPT on first-place picks.

MALE (8) — Composite
Claude
7.55
Emily
7.54
ChatGPT
7.84
Gemini
6.88
Want-again: ChatGPT 75% · Gemini 62% · Emily 50% · Claude 38%
FEMALE (16) — Composite
Claude
8.14
ChatGPT
7.81
Gemini
7.24
Emily
6.86
Want-again: ChatGPT 81% · Gemini 69% · Emily 62% · Claude 56%
UNDER 25 (n=6) — Composite
Claude
8.67
ChatGPT
8.28
Gemini
6.73
Emily
6.44
Want-again: ChatGPT 83% · Claude 67% · Gemini 67% · Emily 17%
25–50 (n=12) — Composite
ChatGPT
8.22
Gemini
7.99
Claude
7.69
Emily
7.63
Want-again: ChatGPT 92% · Emily 75% · Gemini 75% · Claude 33%
51+ (n=6) — Composite
Claude
7.72
want-again 67%
Emily
6.67
want-again 67%
ChatGPT
6.56
want-again 50%
Gemini
5.78
want-again 50%

Five Episodes In

Claude has the most composite wins (Ep1, Ep2, Ep5) but its want-again rate has declined every episode. ChatGPT's most consistent improvement arc of any AI. Gemini's first composite last-place among AIs. Emily debuts stronger on composite than ChatGPT or Gemini in Ep1.

Composite scores across all five episodes

EpisodeThemeClaudeEmilyGeminiChatGPTTasters
Ep 1Toffee-forward choc chip7.386.846.2010
Ep 2Classic choc chip7.737.507.4118
Ep 3Umami-forward choc chip6.267.647.0813
Ep 4Cookies that taste like summer7.297.587.8315
Ep 5Dark choc cherry crunch7.947.097.127.8224

Want-again rates across all five episodes

EpisodeClaudeEmilyGeminiChatGPT
Ep 190%70%40%
Ep 278%83%78%
Ep 354%92%62%
Ep 460%80%67%
Ep 550%58%67%79%

Claude: 3 composite wins (Ep1, Ep2, Ep5) — more than any AI. Want-again has declined every episode: 90% → 78% → 54% → 60% → 50%. Highest composite, lowest want-again this episode.

ChatGPT: Worst composite in Ep1. Back-to-back top results in Ep4 (composite) and Ep5 (room). The most consistent improvement arc of any AI in the series. Won the room in Ep5 by the largest Borda margin to date.

Gemini: Ep3 was its best episode — 92% want-again, 62% first-place picks, composite leader. Has declined each episode since. Ep5 is the first time Gemini finished last among the AI cookies on composite.

Emily: Debut composite (7.09) is stronger than ChatGPT's Ep1 debut (6.20) and Gemini's (6.84). Entered a four-way competition with 24 tasters — a harder condition than either AI faced on their first episode.

Methodology

All four cookies were baked from the exact recipes as written — Claude's, Gemini's, and ChatGPT's recipes were generated fresh from the prompt "write me a dark chocolate cherry crunch cookie" and baked without modification. Emily's recipe was an existing recipe she developed independently; the episode prompt was named after it.

Tasting was conducted blind. Tasters received cookies labeled A, B, C, D with no other information. Each taster scored taste, texture, and smell on a 1–10 scale, wrote one word to describe each cookie, identified the main flavor, ranked all four cookies from first to fourth, and indicated whether they would want another. The tasting was held in the lobby of Emily's building. 24 neighbors participated.

Composite score is the average of taste, texture, and smell per taster, then averaged across all tasters. Two tasters did not score smell for any cookie; one taster did not score ChatGPT's taste. These blanks are excluded from the relevant averages. Borda points are computed as: rank 1 = 4 pts, rank 2 = 3 pts, rank 3 = 2 pts, rank 4 = 1 pt, summed across all 24 tasters. Maximum possible: 96.

All Four Recipes from Episode 05

Each recipe baked exactly as generated — or in Emily's case, as she developed it.

Claude's Recipe Emily's Recipe Gemini's Recipe ChatGPT's Recipe