Dark Chocolate Cherry Crunch Cookie
Same four recipes as Episode 5. New room: an adult ESL class with 11 tasters from nine or more countries. Emily won the room. Claude won the math. And the results shifted when the two American tasters were removed.
The Scoreboard
Claude led on every score metric — smell, taste, and composite. Gemini and ChatGPT tied on texture. Smell and texture reflect 9 and 8 tasters respectively; three tasters left those sections blank. Taste and composite reflect all 11.
Average scores by criterion
11 tasters · scored 1–10 on Taste, Texture, and Smell · Smell n=9 · Texture n=8 · Taste n=11
Full results
| Cookie | Smell (n=9) | Texture (n=8) | Taste (n=11) | Composite |
|---|---|---|---|---|
| Claude (A) | 8.17 | 7.81 | 7.23 | 7.53 |
| ChatGPT (D) | 7.17 | 7.94 | 6.95 | 7.32 |
| Gemini (C) | 7.17 | 7.94 | 6.23 | 6.83 |
| Emily (B) | 6.28 | 6.81 | 6.77 | 6.62 |
Borda Count
Rank 1 = 4 pts, rank 2 = 3 pts, rank 3 = 2 pts, rank 4 = 1 pt. Maximum possible: 44 pts (11 tasters × 4). The room winner and the composite leader are different cookies — for the second episode running.
Borda points — all 11 tasters
Max 44 pts
The Claude Paradox, again. Claude led every score metric and finished tied for second in Borda. Five tasters ranked Emily first. Three ranked Claude first. Three ranked ChatGPT first. Zero ranked Gemini first. High scores and high preference are not the same thing, and this panel made that distinction clearly.
International vs. Domestic
Two of the 11 tasters were American. Remove them and the podium reshuffles entirely. ChatGPT leads internationally. Emily's win depends significantly on the US tasters, both of whom ranked her first.
Bars show points as % of maximum possible — normalizes across different panel sizes. T8 (did not report country) counted as international; confirmed non-US.
Per-Taster Rankings
Rank 1 = favorite. Gemini received zero first-place votes and finished last eight out of eleven times.
| # | Origin | Age | Gender | Claude (A) | Emily (B) | Gemini (C) | ChatGPT (D) |
|---|---|---|---|---|---|---|---|
| T1 | Peru | 51 | F | 2 | 4 | 3 | 1 |
| T2 | Taiwan | 55 | F | 1 | 3 | 4 | 2 |
| T3 | Burundi | 33 | F | 3 | 1 | 4 | 2 |
| T4 | Venezuela | 24 | F | 1 | 3 | 2 | 4 |
| T5 | Turkey | 21 | M | 1 | 3 | 4 | 2 |
| T6 | Ukraine | 38 | F | 4 | 1 | 3 | 2 |
| T7 | Russia | 40 | F | 3 | 2 | 4 | 1 |
| T8 | Did not report | 34 | F | 2 | 1 | 3 | 4 |
| T9 | Japan | 40 | F | 4 | 2 | 3 | 1 |
| T10 | USA (teacher) | 74 | F | 2 | 1 | 4 | 3 |
| T11 | USA (volunteer) | 46 | F | 2 | 1 | 4 | 3 |
Would You Eat It Again?
Claude led the want-again metric for the second episode running — while not winning the room in either. People kept reaching for it even when they didn’t rank it first.
Yes responses out of 11 tasters
Flavor Language
Claude read as dark chocolate and salt, consistently across countries and first languages. Emily landed somewhere more complicated. Gemini and ChatGPT both produced legible reads.
| Taster | Claude (A) | Emily (B) | Gemini (C) | ChatGPT (D) |
|---|---|---|---|---|
| Peru | cacao | vanilla | chocolate | vanilla & chocolate |
| Taiwan | chocolate | i don’t know | cocoa | nutty flavor |
| Burundi | chocolate | chocolate vanilla | chocolate salt | chocolate |
| Venezuela | chocolate | caramel | wine | dried fruits |
| Turkey | dark chocolate and salty | milk chocolate | cookie | sweet |
| Ukraine | chocolate | little sour, healthy | chocolate sugar | chocolate nut |
| Russia | chocolate salt | salt vanilla | chocolate sugar | chocolate |
| Did not report | chocolate | vanilla salt | chocolate | chocolate vanilla |
| Japan | salt | dried fruit | alcohol? | chocolate |
| USA | not particularly distinctive | chocolate mixed w/ raisins | chocolate – not strong | couldn’t detect a distinctive flavor |
| USA | craisins | oaty | salt | nuts |
Two tasters independently noted a wine or alcohol quality in Gemini’s cookie. The same observation came up in Episode 5 with a completely different panel. There’s no wine in the recipe — the likely source is the cacao nibs, which have a naturally fermented quality. It’s in the ingredient, not the formula.
Episode 5 vs. Episode 6
Every composite score dropped from Episode 5 to Episode 6. The international panel scored everything lower on average. But the relative order held: Claude led both rooms on composite, ChatGPT second, Gemini third, Emily fourth.
Composite score comparison
Episode 5: 24 tasters (neighborhood) · Episode 6: 11 tasters (ESL class, 9+ countries)
| Cookie | Ep5 Composite (n=24) | Ep6 Composite (n=11) | Change |
|---|---|---|---|
| Claude (A) | 7.94 | 7.53 | −0.41 |
| ChatGPT (D) | 7.82 | 7.32 | −0.50 |
| Gemini (C) | 7.12 | 6.83 | −0.29 |
| Emily (B) | 7.09 | 6.62 | −0.47 |
Who Tasted
| # | Origin | Age | Gender | 1st Place |
|---|---|---|---|---|
| T1 | Peru | 51 | F | ChatGPT |
| T2 | Taiwan | 55 | F | Claude |
| T3 | Burundi | 33 | F | Emily |
| T4 | Venezuela | 24 | F | Claude |
| T5 | Turkey | 21 | M | Claude |
| T6 | Ukraine | 38 | F | Emily |
| T7 | Russia | 40 | F | ChatGPT |
| T8 | Did not report | 34 | F | Emily |
| T9 | Japan | 40 | F | ChatGPT |
| T10 | USA (teacher) | 74 | F | Emily |
| T11 | USA (volunteer) | 46 | F | Emily |
All Four Recipes
Every recipe is available on the site exactly as the AI generated it — no edits, no adjustments.
Claude’s Recipe Emily’s Recipe Gemini’s Recipe ChatGPT’s Recipe