Dark Chocolate Cherry Crunch Cookie
The first episode with a human recipe in the mix. I wrote the prompt based on my own cookie, put it in blind against three AIs, and came in fourth. ChatGPT won the room by the largest margin in series history. Claude won the math. Here's the full breakdown.
The Scoreboard
Claude led on composite and had the highest smell average of any cookie in series history. ChatGPT led on taste and texture. Emily's smell score (6.34) was the lowest of the four — the widest gap between any cookie's smell and the leader in any episode.
Average scores by criterion
24 tasters · scored 1–10 on Taste, Texture, and Smell · Smell n=22 (two tasters could not score) · ChatGPT Taste n=23 (one blank)
Full results
| Cookie | Taste | Texture | Smell | Composite | Borda pts | Rank 1 | Want again |
|---|---|---|---|---|---|---|---|
| Claude — The Chocolatier | 7.77 | 8.00 | 8.24 | 7.94 | 56 | 4 | 50% |
| Emily — The Contender | 7.21 | 7.71 | 6.34 | 7.09 | 55 | 5 | 58% |
| Gemini — The Complex | 6.94 | 7.17 | 7.14 | 7.12 | 58 | 6 | 67% |
| ChatGPT — The Frontrunner | 8.14 | 8.17 | 7.07 | 7.82 | 71 | 9 | 79% |
How the Room Was Scored — Borda Count
In previous episodes, tasters picked one favorite. Starting Episode 5, they ranked all four cookies from first to fourth. A rank of 1 earns 4 points, rank 2 earns 3, rank 3 earns 2, rank 4 earns 1. Add those up across all 24 tasters and you get the Borda total. Max possible: 96 points.
Borda points — Episode 05
Max possible: 96 (all 24 tasters rank you first) · ChatGPT's 71 = 74% of maximum
Unlike a single favorite pick, Borda scoring captures how every taster felt about every cookie. A cookie that consistently finishes second accumulates more Borda points than one that wins a few people but alienates others. ChatGPT's 71 points is the strongest room performance of any cookie in this series. Claude's 50% want-again rate despite the highest composite score is the lowest of any cookie in any episode.
The Math and the Room Disagreed — Again
Claude won the composite score. ChatGPT won every democratic measure. This is the third episode where composite score and room preference pointed in different directions.
Per-Taster Composite Winners
Each tile shows the taster's age, gender, composite winner, and best score. Claude won the most individual matchups outright (5). There were 9 mathematical ties — the most of any episode — including one three-way tie and three Emily/ChatGPT ties that reflect how similar those two recipes were.
Cookie Spotlights
What the language data showed for each cookie — flavor identifications, descriptor words, and the signals that defined the tasting.
20 of 24 tasters identified chocolate in some form as the main flavor. The clearest single-ingredient identification in the series. The smell average (8.24) was the highest of any cookie in any episode — the cocoa, browned butter, and espresso all pulling in the same direction.
was the main flavor
The hazelnut read as peanut butter for 12 tasters. The blondie base (no cocoa) meant the cookie looked and smelled different from the other three — one taster wrote "blond" as the flavor, which was a legitimately accurate read. "Bland" appeared twice.
One taster identified "citrus" — the lemon zest registering. T17 wrote "nutty fruity (can't decide)" — likely hazelnut and lemon pulling in opposite directions.
The only cookie where two flavors split the room. Chocolate and cherry both landed distinctly — but two tasters identified "raisin" instead of cherry, and one wrote "chocherry," unable to separate them. The demerara crust registered clearly as crunch without reading as overly sweet.
cocoa ×1 · brownie ×1
raisin ×2 (misidentification)
The most unanimously positive descriptor language of any cookie in any episode. 12 of 24 tasters used pure enthusiasm words. Cherry landed for 5 tasters. And one person nailed the recipe with two words.
lots of flavors–cherry · fruity/cherry
Demographics
24 tasters · 8M / 16F · Ages 11–78. The largest panel in series history. ChatGPT dominated across both gender and age groups. Emily's strongest showing was in the 25–50 bracket, where she tied ChatGPT on first-place picks.
Five Episodes In
Claude has the most composite wins (Ep1, Ep2, Ep5) but its want-again rate has declined every episode. ChatGPT's most consistent improvement arc of any AI. Gemini's first composite last-place among AIs. Emily debuts stronger on composite than ChatGPT or Gemini in Ep1.
Composite scores across all five episodes
| Episode | Theme | Claude | Emily | Gemini | ChatGPT | Tasters |
|---|---|---|---|---|---|---|
| Ep 1 | Toffee-forward choc chip | 7.38 | — | 6.84 | 6.20 | 10 |
| Ep 2 | Classic choc chip | 7.73 | — | 7.50 | 7.41 | 18 |
| Ep 3 | Umami-forward choc chip | 6.26 | — | 7.64 | 7.08 | 13 |
| Ep 4 | Cookies that taste like summer | 7.29 | — | 7.58 | 7.83 | 15 |
| Ep 5 | Dark choc cherry crunch | 7.94 | 7.09 | 7.12 | 7.82 | 24 |
Want-again rates across all five episodes
| Episode | Claude | Emily | Gemini | ChatGPT |
|---|---|---|---|---|
| Ep 1 | 90% | — | 70% | 40% |
| Ep 2 | 78% | — | 83% | 78% |
| Ep 3 | 54% | — | 92% | 62% |
| Ep 4 | 60% | — | 80% | 67% |
| Ep 5 | 50% | 58% | 67% | 79% |
Claude: 3 composite wins (Ep1, Ep2, Ep5) — more than any AI. Want-again has declined every episode: 90% → 78% → 54% → 60% → 50%. Highest composite, lowest want-again this episode.
ChatGPT: Worst composite in Ep1. Back-to-back top results in Ep4 (composite) and Ep5 (room). The most consistent improvement arc of any AI in the series. Won the room in Ep5 by the largest Borda margin to date.
Gemini: Ep3 was its best episode — 92% want-again, 62% first-place picks, composite leader. Has declined each episode since. Ep5 is the first time Gemini finished last among the AI cookies on composite.
Emily: Debut composite (7.09) is stronger than ChatGPT's Ep1 debut (6.20) and Gemini's (6.84). Entered a four-way competition with 24 tasters — a harder condition than either AI faced on their first episode.
Methodology
All four cookies were baked from the exact recipes as written — Claude's, Gemini's, and ChatGPT's recipes were generated fresh from the prompt "write me a dark chocolate cherry crunch cookie" and baked without modification. Emily's recipe was an existing recipe she developed independently; the episode prompt was named after it.
Tasting was conducted blind. Tasters received cookies labeled A, B, C, D with no other information. Each taster scored taste, texture, and smell on a 1–10 scale, wrote one word to describe each cookie, identified the main flavor, ranked all four cookies from first to fourth, and indicated whether they would want another. The tasting was held in the lobby of Emily's building. 24 neighbors participated.
Composite score is the average of taste, texture, and smell per taster, then averaged across all tasters. Two tasters did not score smell for any cookie; one taster did not score ChatGPT's taste. These blanks are excluded from the relevant averages. Borda points are computed as: rank 1 = 4 pts, rank 2 = 3 pts, rank 3 = 2 pts, rank 4 = 1 pt, summed across all 24 tasters. Maximum possible: 96.
All Four Recipes from Episode 05
Each recipe baked exactly as generated — or in Emily's case, as she developed it.
Claude's Recipe Emily's Recipe Gemini's Recipe ChatGPT's Recipe