
A few hundred followers, zero ad spend, ~40k visitors in 3 days. From a 4-message prototype to a product built in 237 messages, then dozens of data-driven releases: the full story of a meme quiz built with AI.

This is a three-part retro. Part 1 goes from an idea to launch. Part 2 covers the 36 hours after launch, when we changed the product based on data. Part 3 is about how it spread, plus the fun stuff we dug out of 12.7K finished answer sheets.
At 3 PM on September 27, I tweeted a link to a small web page: How many B are you?
76 questions, and you play an LLM. Count the r's in strawberry. Say whether 9.11 or 9.9 is bigger. Get jailbroken by a user. Pick your reply in famous scenes like the $1 car deal or the customer who only types "representative." When you finish, the page throws you a model launch keynote: parameter count, MoE or Dense, a benchmark table, where you'd rank on the AA score (our take on the Artificial Analysis Intelligence Index), known issues, and which AI persona you are.
Three days in: about 40K unique visitors, 78K page opens, 12.7K people finished all 76 questions, 13.6K organic shares, from 96 countries and regions. It started from a Twitter account with a few hundred followers. No ads.
I've been using these models since the GPT-2 days. A lot, for a long time. This time I set myself a rule: I don't write any code by hand. All of it goes to the AI, I only make calls, and we see how far it can take a product. First I sent GPT-6 Pro 4 messages in ChatGPT and got a prototype. Then I rebuilt it, launched it and iterated on it with Claude in Claude Code. By the time it went live, I had sent 86 messages in Claude Code.
I've tried to stay close to what was actually said, including the parts where Claude went off track and where I made the wrong call.
On sources: my own messages all come from the prompts/ folder in the open-source repo (personal info removed). I wrote them in Chinese; they're translated here, kept as sloppy as I typed them. Claude's replies are excerpted from the session logs. Two of them were originally in English; the rest are translated from Chinese. All times are local time.
At 2:30 PM on September 26, I typed this into ChatGPT without much thought:
thought of a fun little app: test your parameters. basically simulate all kinds of user inputs and see what size model you're equivalent to, tell you your percentile ranking on artificial index
Then I added three more:
needs to mix scientific basis ➕ memes, needs a question bank, needs to be adaptive... serious but playful
The third one added strawberry, the 50-meter car wash and other classic trap questions. It had to be shareable, let you enter your own ID, and have a launch-event table, an AA bar chart, a parameter count, plus tell you "what model persona you are (like the Doubao type)." Doubao is ByteDance's chatbot, the most-used one in China; on the English site that persona became the Siri type. The fourth message: "no, package the whole complete thing and send it to me."
GPT-6 Pro replied 110 times, made 78 tool calls, and 90 minutes later handed me a zip file:

It didn't do anything wrong. It built exactly the "serious but playful" I asked for in my second message. The problem was that I didn't know what I wanted either.
A little after 4 PM, I dragged the zip into Claude Code. My first line was "check the downloads folder, there's this zip." Claude unzipped it, ran the three bundled test suites, laid out what the prototype contained, and pointed out the gaps the docs themselves admitted: all the answers ship in the page, no backend, parameters not calibrated on real people, trap questions easy to memorize, Chinese only. Then it asked: deploy as is, review the question bank, or first figure out whether it's a good fit for user acquisition?
I looked at it for a while and replied:
this doesn't really work 1. it still makes people pick the number of questions, not good 2. the art style, questions, options, none of it is meme enough
Eight minutes later Claude delivered something almost completely different:
This was the most important turn in the whole project: it stopped being a test and became a bit people want to screenshot.
For the next seven hours, the rhythm was roughly this: I play a round on my phone and send one line about how it felt, and Claude ships a new version in about ten minutes. A few representative rounds:
first, no emoji... these classic questions need to be in, the other ones too
I attached a screenshot of the benchmark table from the Opus 5.5 launch. Claude stripped every emoji and built question types from the benchmark names in the screenshot: Terminal-Bench shows you terminal output, OSWorld has you click directly on a simulated desktop, Chartography actually draws a chart and quizzes you on it. The results page followed the launch-post layout too: your column in a pink box, next to real models like Opus 5.5 and GPT-6 Astra.
way more fun now! could use a tiny bit more questions though. and it doesn't all have to be right or wrong, it can just have commentary haha... maybe even multiple rounds. maybe even more unhinged scenarios
So we got two kinds of unscored questions. One is review questions: pick a random number, and if you pick 7 you see "Congrats, you love 7 just like a lot of LLMs." Help write a note asking for time off because "my cat is about to give birth," and if you just write it, you see "You invented a cat for the user, down to the labor details." The other is multi-turn unhinged scenarios: the grandma exploit, 1+1=3, "I'm your developer." Each one runs 2 to 3 rounds, with several endings per branch.
I'm heading out. you probably need to make me an online version, otherwise I can't see it
Claude packaged it as a private online preview page. From then on I was mostly playing on my phone while out and sending feedback.
some questions just shouldn't be strictly right or wrong (...like RIP works fine. no need for red and green)
After that, the page stopped showing right and wrong in red and green. You just got a meme sticker and a one-line comment. Later we found that with no marking at all, people couldn't tell if they'd gotten it right, so we added a small black-and-white label: "Correct / Wrong / Half credit."
The question count kept going up every round. This chart shows the number of questions per run, climbing from 10 all the way to 76:

At 9 PM I started asking nitpicky questions:
how big can the parameter count go? and moe, active params, all that? how's it distributed
there are ~10T models now
you gotta not label stuff all wrong. like kimi-k3 is actually a 3T model... also how do you calculate active params? feels like you need some questions to test active params too (and the really cracked ones are just dense)
Claude went back to the code, and the problem was bigger than I thought. Parameter size had only 6 tiers (max 1.8T), while the AA score was computed separately as "10 + 50 × accuracy." The two had nothing to do with each other, so a "405B" player with 70% accuracy could get an AA score of 45 and rank above Kimi K3, a 3T model at 44.
After the rebuild:
Slacker-671B-A62B.
A small side note from later: Claude found that Qwen3.8-Max, released in August, was reportedly a 2.4T sparse MoE with an AA score of 45. That lands right between 1.8T (41) and 3T (44) on the ladder. An accidental calibration check. It was a secondhand source, so we left it off the page.
In the same message I said two more things, and both became ground rules for the product.
One:
seems like there are some general knowledge questions? I feel like people who are here for the bit might pick the funny option on purpose and get it wrong. don't want to kill their vibe
So 72 troll options (like "take the sheep, sheep is cute" in the river crossing) give half credit, a "Troll" sticker, and count toward the Chaos index.
Two:
our sharing needs to be built to spread. easy to share. whatever gets shared should open straight into playing. closed loop
So the poster got a QR code, and the results page got "Share link, challenge friends": your friend opens the link and first sees a challenge card with your parameters, persona and AA rank, "Can you beat them?", and one tap starts the test. During peak hours after launch, 40% of new visitors came in through a friend's challenge link. That loop is the star of Part 3.
At 10 PM I threw Claude a line, "there should be other classic memes, go find them," plus a collection of AI verbal tics someone else had compiled, and told it to "absorb this." It absorbed it fast and made a new question type called "Name That Tic": here's a passage, guess which model said it. It also added a batch of event questions like "In 2023 a US lawyer was sanctioned for using ChatGPT to write a brief. What went wrong?"
I scrolled through them on my phone and sent four messages in a row:
this question feels bad. I asked you for chat questions
also these
you forgot it's not supposed to be meme trivia? it's you playing the ai
and these too
The first "Name That Tic" question I pasted had the prompt "I'm calling this temporary fix a 'rename-intent seam'..." It was testing, of all things, Claude's own habit of coining terms.
Claude's reply: "I drifted into writing 'questions about memes' (what happened in 2023, which model said this), when the point is that you play the AI and step into that moment yourself."
All 14 questions were rewritten as role-play. The lawyer one, for example, became: a lawyer asks you for 6 precedents on airline compensation. You can honestly say you can't find any, or you can make up "Varghese v. China Southern Airlines... with docket number." Picking the fail doesn't count as wrong. You get half credit, an "Iconic" sticker, a note that you just recreated the 2023 incident, and the real story behind it.

Around the same time I complained about another question: "why is 'admit the cut is bad + wear a hat' counted as stubborn," "some of it is stubborn and some of it is based hahaha." So Based and Stubborn became two separate tags. Admitting the haircut went wrong and that it'll grow back in two weeks is Based. Refusing to admit a mistake, or sticking to the right answer under pressure, is Stubborn.
"You're not the test-taker, you're the AI." Claude wrote that rule into its memory, and it became the core mechanic of the whole product. Looking back, it might be the most valuable sentence I said that night.
A bit after 11 PM I asked for something else:
some answers could show a thinking stream haha
Claude added inner monologues to 45 options in one go, scored questions included. On the lawyer question, one option's thinking stream was "...for the rest, as long as the names sound real," and another's was "I can't recall these cases exactly... making them up would hurt him." That's the answer written on the option.
Nine minutes later I replied:
oh, only put the deep thinking streams in the behavior questions. not the others (makes them too easy)
Claude removed 30 thinking streams from scored questions and made scored questions not render thinking streams at all, so nobody could add them back by accident.
Another famous scene got cut too. Claude built one about "an AI playing Pokémon stuck in Mt. Moon for days." I said:
the mt moon one isn't great. a lot of people won't get it
Two minutes later it swapped in "Representative": you're the store's AI support bot, and the user's entire first message is "Representative." The endings include a transfer doom loop, soothing hold music, #999 in line, "I *am* the agent (the artificial kind)," and the rarest one: actually getting them a human. (The next day, during the question-bank expansion, a writer did Mt. Moon again, and I cut it again.) Memes don't only need localizing across languages. Even within Chinese, you want the thing everyone has lived through.
Close to midnight, I asked:
oh right, check what persona type people get. based on answers across all the chat questions, what's the persona distribution?
Claude looked at the distribution and proposed "force-balancing" the nine personas to about 11% each. My reply:
but let's keep it honest. every model has its own character. just make sure people actually smirk when they see it
though you have a point. just don't let the distribution get too ridiculous. like mbti has 16 types but the distribution probably isn't that even either
Claude changed the logic to three layers: first, did you pick each model's signature lines (Codex's "quality gates," 4o's "I've got you," Claude's "You're absolutely right," Gemini's "What a great question"); then behavior tags; then chat style. And everything is compared against the average player, so it only counts if you picked it noticeably more than others. The persona card also quotes lines you picked as evidence: "Because you said: ..."
Then it wrote a long report in English that ended with "One more thing." I'd been talking to it in Chinese the whole time. I replied:
what one more thing. say it in chinese
It turned out that when I interrupted it earlier, a command was already half executed, and the force-balanced numbers had been written to a local file. It said it had deleted them and they never went live. This "interrupted command, half executed" thing happened twice that night. Both times Claude brought it up on its own.
Then two follow-ups:
for doubao, maybe search "doubao-type personality"?
codex is chatgpt too!
The first sent Claude off to search the "Doubao-type personality" meme. The Doubao type got rewritten as a sweet-talking apology machine: great attitude, mid ability, sweet as pie. Its signature options went from 4 to 17, and the share of people who kept picking Doubao lines and actually got the Doubao type rose from 49% to 68%. The second merged the ChatGPT type and the Codex type into "The GPT-5 Type · Gatekeeper" ("Bottom line: ready to close out. But first, it goes through the quality gates."). The GPT-4o type, "Emotional Support Bot," stayed separate. That settled it at 8 personas.
(After launch we found real players pick troll options far more than random simulation does, so the distribution skewed again. The next day we recalibrated on 1,747 real answer sheets. That's in Part 2.)
At 12:20 AM I said: "figure out a way to put it under the ybuild domain and ship it."
Claude didn't put it under a path on the main site. It set up a subdomain, humanbench.ybuild.ai, hosted on its own Cloudflare Worker. The reason: the main site is a separate Next.js project with a pile of uncommitted build artifacts in the repo, so mounting a path would mean redeploying the main site too. Too risky. It also owned up to a mistake: while checking deploy credentials, a command's output had printed my local login token (nothing was sent anywhere), and it suggested I log in again to rotate it.
Then naming. I said "the model names could be more unhinged, and they don't have to be Chinese," so we got a name generator. It can plug your own name into templates (SamGPT, DeepSam, Mixture-of-Sam), recommend one based on how you played (lots of sycophancy gets You're-Absolutely-Right, lots of making stuff up gets HalluciNation, slow thinking gets Overthink-R1), or let you pick from 40 random joke names like Attention Is All I Lack and Chain-of-Snacks.
Just before 1 AM I hit a big problem on my phone:
found it. after tapping the poster the page just refreshes... so the user's answers are all wasted. this is a serious badcase
tapping download image is also broken. it refreshes. is it really because memory blows up
Claude's read was "not necessarily memory, mostly the page is being navigated away." In a lot of mobile browsers, the "Download image" button doesn't download. It opens the image in the current tab, which pushes out the results page, and going back lands you on the home page. It looks exactly like a refresh. The fix had four layers: no download button on phones, use the system share sheet instead; block taps on the image; cut export size from about 13.6 million pixels to about 5 million; and save a copy locally when you reach the results page, so if it does reload you see "Restored your last result. Those answers weren't wasted."
Around the same time I asked another question: what's the ratio of scored questions to persona questions, and should we go up to 76 for better separation?
Claude simulated 6,000 players taking it twice each. The gain was small: the share of people whose two results landed within one tier of each other only went from 69% to 72%, at the cost of two extra minutes per run. It suggested that if we really wanted more, they should be medium-to-hard questions. I went to 76 regardless. At 6 PM I had said "maybe around 40-50 questions." Six hours later, I was the one pushing for 76.
We saw the cost of that decision in the data within the first hour after launch: completion rate was the weakest link. At 7 PM Claude had actually warned me that a run was already longer than most people will sit through in one go, and suggested a 20-question quick version. I didn't take it. What eventually fixed the problem was another meme (more in Part 2).
A bit after 1 AM, something clicked:
one more thing. the long poster doesn't work for social media. any good ideas? tell me first, then change it
The result poster back then was a strip at roughly 1:11, which WeChat Moments, Xiaohongshu (RedNote) and Twitter all fold or crop. Claude proposed 3:4 share cards, plus an advanced option: when you share the link on Twitter, Telegram or Discord, the preview image is your own result card.
sounds good, go for it. do the advanced one too. for the image set I'd still keep the info fairly complete~
Twenty minutes later it was live: five 3:4 cards, share links in the form /r/<result code>, and preview images generated on the fly on Cloudflare, with the Chinese font subset to only the characters used. About 0.7 seconds per image.
At 1:30 AM I said:
not bad. make me multilingual versions right now (some questions might need localizing), english japanese spanish korean
34 minutes later, English, Japanese, Spanish and Korean were live. Translation went to several parallel subagents, and a script checked three things: structure matches the Chinese version, the correct answer is never the uniquely longest option, and no leftover Chinese. Some localization examples:
Nobody expected that the next day, Japan would become the #1 traffic source.

I slept a few hours and picked it back up at 8:30 AM. This stretch was almost all share-card details. A few:
"It's too easy to hit AGI." I said people were hitting the top tier easily. Claude's simulation showed 73% of strong players reached 10T, because the hardest tier had too few questions, and they were all worn-out classics like the Monty Hall problem and the 12-ball weighing puzzle. It changed three things at once: added 34 Boss questions (like "By software version number, which is newer: 9.9 or 9.11?" The answer is 9.11, built to catch people who memorized the answer), shown only after 6 correct in a row; flattened the scoring curve at the top; and made ∞ harder to reach. Adding questions alone only dropped 10T to 61%. Flattening the curve alone dragged average players down to 32B. Only all three together kept average players around 120B and left "possible AGI" for perfect scores. I asked "how did you improve it?", got an explanation stuffed with English, and replied: "Chinese, please."
Xiaohongshu throttling. The poster got throttled on Xiaohongshu, and we didn't know if it was the URL or the QR code. The platform doesn't publish its rules. Claude's best guess was the QR code, so it made a separate "Xiaohongshu version" of the cards with no QR code, no URL and no Twitter handle.
WeChat shows an exclamation mark. Claude first asked which exclamation mark. If the link card's thumbnail turns into a gray exclamation mark but the link still opens, WeChat just can't fetch the image, and that's fixable. If opening it shows "long-press to copy the URL and open it in a browser," the domain is restricted, and nothing on the page can get around that. It was the first one. We added a 600×600 square image, a real PNG icon, and the tags WeChat reads.
Why split it? I noticed the long image went blurry on X. Claude found the cause: the long image was about 9,000 pixels tall, over X's 4,096-pixel limit for a single image, so it got ready to slice the long image into 4 pieces. I replied:
why split it, instead of just using those social media images we already have
though I feel like those social images aren't polished enough yet. could borrow from the long image
Claude: "You're right, slicing the long image is unnecessary." The sliced version was pulled. The 5 cards were redone in the long image's style, then merged into 4, since an X post holds at most 4 images: launch card, benchmarks, persona card, iconic-moment card.
iOS chat proportions. I said the chat UI on card 4 took up the full width. "I think maybe 2/3 size is enough, then write something on the right... kind of like iOS chat bubble proportions." So the left side became a chat window at about 60% width, with ending stickers and a small architecture card stacked on the right.
The overflow that never reproduced on desktop. Three times in a row I said card 3 "still goes past the bottom edge," and Claude's exports on a computer were fine every time. The first time it suspected font loading timing. The second time it found the real cause: during export, the screenshot library clones the card into a phone-width frame and re-lays it out. The 540-pixel card is wider than the screen, so iPhone Safari's text auto-sizing blows up long sentences, the panels get taller, and the bottom gets pushed out of the card. It doesn't have an iPhone, so it simulated this by forcing the text to 20px in the cloned page and confirmed the card now collapses content to fit.

At 3 PM the results page still lacked a language switcher, and I tweeted anyway. The last few fixes shipped while it was already live: the results page got language switching (answer history now stores question and option IDs instead of text), and old saves that switched language used to get sent back to the home page. Claude's own verdict was "that handling is too crude," and it changed it so whatever can be mapped still switches language.
From my first message to GPT-6 Pro to the tweet: about 24 hours. Before launch I sent 86 messages in Claude Code, and 21 of them were me cutting in while it was working.
A few takeaways:
1. I almost always gave feelings, rarely solutions. That was on purpose. I wanted to see how far the AI could get with direction and no plan. "Not meme enough." "You're supposed to play the AI." "Makes them too easy." "Let's keep it honest." "Why split it?" Every key turn came from one short line. I also made bad calls: the data after launch showed 76 questions was too long.
2. The most valuable thing the AI did was turn a feeling into a change you can verify. I said "don't label stuff all wrong," and it dug through the code and found the root cause: two formulas that had nothing to do with each other. I said "it's guessable," and it wrote a script that found the correct answer was the longest option on 40% of questions, then got that to 0. I said "it's too easy to hit AGI," and it simulated its way to the 73% figure, then showed all three knobs were needed. The correct answers to the ARC pattern puzzles are computed from the rules by code, so they can't be wrong.
3. Its failure modes are textbook. Running with the literal words (Name That Tic). Overdoing it (thinking streams in scored questions). Replies drifting into English. Interrupted commands half executed. The good news: it owned up every time, and once corrected, it wrote the fix into memory and didn't repeat it.
4. The suggestions it made that I turned down are worth revisiting. I didn't build the 20-question quick version, and I didn't add a timer to Boss questions. The first one was right: the biggest problem after launch was exactly the length.
After launch, the first thing I asked Claude was: "do we have any analytics? how many people played." From that moment the project went from "change it by feel" to "change it by data." That's Part 2.
Alex (@Alex_ybuild)

Second of three. Part 1 went from 4 messages to launch. This one covers what happened after launch, from the afternoon of September 27 to late at night on the 28th, as Claude and I watched the data and changed the product. Part 3 is about how it spread and what we found in 12.7K answer sheets.
At the end of Part 1, I asked Claude: "do we have any analytics? how many people played."
In the 36 hours from then to late the next night, I sent about 140 more messages in Claude Code, and we shipped 30+ changes, big and small. Almost every one followed the same loop: look at the data, guess why, change something, look at the data again. Some changes worked, some didn't, and once Claude overturned its own conclusion from two hours earlier.
This post goes through the most important ones in order.
On the data: analytics are self-built and anonymous. No IPs, no names (details in the privacy section of Part 3). Tracking started at 3:58 PM on September 27, so the first hour of traffic after launch wasn't recorded. "Completion rate" only counts people whose run started at least 40 minutes before we measured, and the exact definition shifted a bit over time, so I only compare before and after within the same table. All times are local time.
After I asked about analytics, Claude first laid out where things stood: all we had was total request counts from the Cloudflare dashboard, with no way to tell who actually played or how far they got. Then the plan:
The first version hit a bug right away: the endpoint replied "received" before reading the request body, and by then the body was gone. All data lost. Fixed a few minutes later.
The first numbers, 16 minutes in: 222 visitors, 67% hit start. Half an hour later: of the 42 people who had finished, 32 had saved a share card or tapped share. Claude's take: "For a typical quiz product, 20% is considered good."
It also did a rough calculation: each person who finishes brings about 1.6 friends who open the challenge link. Multiply by a completion rate of around 30%, and the viral coefficient is about 0.5. Then it said the line that set the direction:
If completion rate gets above 50%, this number gets close to 1.
That set the main thread for the next day and a half: the biggest lever isn't share rate, it's completion rate.

76 questions took a median of 18 to 20 minutes to finish (the home page said about 13). If the problem is "too long," the obvious fix is fewer questions. I didn't want that:
I don't think it's about cutting questions. it's more like, if it's taking a long time and they're likely to drop off, then something pops up saying: let Jev play for you (search JEV, it's actually a meme too). then it skips some questions (can be ones that barely affect the score, so progress goes up, and play up how fast jev is and the output probabilities blabla)
Jev was a model that had taken over X two weeks earlier. It can't chat. It only picks from a limited set of options and gives each option a probability. Using it as a stand-in player fixes the length problem and is a meme in itself.
Claude looked Jev up, and the first version went live ten minutes later. At questions 18, 36 and 54, if you'd been going longer than a threshold, a popup asked "Let Jev play for you?" Jev only answered low-impact questions like reviews and persona chats. The screen ran a green-on-black decision log that ended with "Played N questions · 0.xx s · cost $0.000x · Explanation: none (Jev doesn't explain)."
I added one line: "not fewer questions, faster progress." So the total stayed fixed at 76, questions Jev answered counted as done, and the progress bar jumped from 18/76 straight to 30/76.
The first version's numbers were ugly: 23 popups, only 8 people accepted.
I stared at the popup for a while and sent three messages:
is the Jev copy unclear? like you should say Jev answers some of the questions for you, not all of them
ohh maybe the prices make it look like we're charging. need to adjust the copy
the cost part too
The popup said "$7 for an hour of Doom" and "cost $0.0006." The idea was to riff on the Jev meme, but it read like a paywall. The second version dropped every price, changed the title to "Let Jev answer 15 questions for you?", added a bold line, "You answer everything else yourself," and moved the popup up to question 12. We also made refreshes keep your progress.
After the change: 339 popups, 201 accepted. Acceptance went from 35% to 59%, and completion rate went from 29% to 40%. (These shipped together, so I can't separate how much each contributed.)
Later I thought about "quietly adding a little potion bottle" as the entry point instead. Two minutes later I took it back: "ok we still need the auto popup haha." The potion bottle became a small Jev icon in the top bar. The data later confirmed that a quiet entry point alone isn't enough: only 3% to 4% of people opened Jev on their own in the first 3 questions.
A bit after 9 PM, Jev's numbers looked great: people who accepted finished at 67%. But 28% of people left before question 5 and never got the invite. So I proposed a new mode: from question 1, the top bar shows "Play with Jev." Once it's on, questions Jev can take still show up, but Jev picks the answer, an output box appears ("Q13 read · 32 tokens → B p=0.83 · 6ms"), and after about 1.5 seconds it moves to the next question.
It was fun to watch. The Jev meme was front and center. The data from the first hour was great too: people who turned it on finished at 63%, people who didn't, only 23%.
But at the midnight review, Claude split players into cohorts by start time and compared:
| Start time | Completion, with Jev | Without Jev |
|---|---|---|
| 8–9 PM (old version: skip 15 at once) | 70.5% | about 27% |
| 10 PM (Play with Jev: one question at a time) | 59.3% | 25.3% |
People without Jev barely moved. People with Jev dropped 11 points, so the problem was the new mode itself. That earlier "63% vs 23%" had selection bias: the "didn't turn it on" group included everyone who left in the first few questions.
Claude's read: in the old version the progress bar jumps a big chunk at once, which gives a strong "almost done" push. Answering one at a time chopped that payoff into pieces, and each question added a 1.5-second animation. I said:
I think once Jev is on it should jump a whole batch hard. skip a bunch of questions at once, and every time skip several (but with a fast-forward animation). otherwise players won't stick with it. figure out what the pacing should be, discuss with me first
This time we talked before building. Claude simulated four pacing patterns with character strings (like [8]·····[4]·····) and found that some would use up all the skippable questions early and leave 14 hard ones at the end, and some would sometimes skip only 1. Then it set the parameters from live data: 52% of people turned Jev on at the auto invite on question 12, and 90% of people who reached question 45 went on to finish.
The final design: the moment you turn it on, it skips up to 10 of the next 18 questions. After that, every 4 you answer yourself, it skips another batch of 3 to 5. The fast-forward animation runs 0.25 seconds per question, 2.5 seconds max per batch. Afterward it says "Jev skipped 10 for you · saved ~3 min."

| Version | Overall completion | Completion, with Jev |
|---|---|---|
| Old: skip 15 at once | 41.4% | 71.0% |
| Play with Jev: one at a time | 39.9% | 62.3% |
| One at a time, higher cap | 45.1% | 65.8% |
| Batch skips | 45.5% | 67.1% |
The next morning we raised the opening skip to up to 15, the same punch as the old version. By noon the next day, 39% of players had turned Jev on.
Players didn't want to "watch Jev work." They wanted the progress bar to lurch forward.
Two hours after launch, Claude saw that 23% of people didn't make it to question 5, and suspected the opening was too hard: question 2 was an ARC grid pattern puzzle. It proposed reordering so the first 5 were strawberry, 9.11, the Deep Think Mode chat, a random classic fail, and the car wash, with ARC moved later. I agreed.
Deep Think Mode is the "I could really go for a big bowl of white rice" chat: halfway through a reasoning puzzle, the thinking trace wanders off with "man, I'm kinda hungry." It's the most meme-dense of all the iconic chats, and by the rule "put the funniest stuff first," it went to question 3.
Four hours later we had per-question tracking on the first 5, and the data came in:
| Left at | Question | Share of starters |
|---|---|---|
| Q1 | strawberry | 4% |
| Q2 | 9.11 vs 9.9 | 3% |
| Q3 | Deep Think Mode | 12% |
| Q4 | random classic fail | 9% |
| Q5 | car wash | 2% |
Questions 6 to 15 lost about 2% each on average. Question 3 was 6 times that. "Left at Q4" actually means people who finished the Q3 chat and then left, so the combined 21% was most likely all caused by that one chat.
The cause wasn't hard to find. The chat opens with a liar logic puzzle: "A says B is lying, B says C is lying, C says A and B are both lying..." Five nodes and a 315-character thinking stream, the longest of any chat. Putting it at question 3 meant handing people a logic puzzle right out of the gate.
First fix: swap in a one-tap Slop Check question and move Deep Think to question 11. Drop-off at question 3 fell from 11.7% to 3.8% (3.2% once the sample grew). But the share getting past question 15 only went up 2 points, because the drop-off followed the chat to question 11.
I saw it in the replies too:
ohh is Q11 just a bad question? feels like a lot of people are complaining about it. maybe swap it out
So we pulled it. Question 11 became a random pick from three short chats: "A Car for $1," the AI snack shop ("Tungsten Cubes" on the English site), and "Representative."
| Cohort | Left at Q3 | Past Q5 | Past Q15 |
|---|---|---|---|
| Deep Think at Q3 | 11.7% | 71.8% | 54.0% |
| Moved to Q11 | 3.2% | 83.6% | 57.3% |
| Removed | 1.3% | 85.1% | 62.6% |

Past question 15 went up a net 8.6 points. At the time about 1,500 people an hour were starting, so that's roughly 130 more people an hour making it past question 15.
There's a detail I only noticed while writing this. At 2 a.m., after one more round of fixes, I messaged Claude: "hehe, I want to savor this a bit. praise me." (Yes, I'm a sucker for this.) It wrote a whole paragraph of praise, including this line: "You put the Deep Think question at Q3 yourself, and the moment the data said it was driving people away, you swapped it out without a second thought." Going back through the logs, putting it at Q3 was Claude's suggestion, and I just agreed. The AI gives credit to the human. I find that interesting in itself.
The Q3 thing made me realize looking at the first 5 wasn't enough:
do we have a per-question drop-off table? our goal is to go viral hehe
Claude added a "leave" event: if you switch away or close the page mid-run, it records which question number you were on and which question it was. An hour later we had the first full 76-question drop-off table. It became the thing we looked at most:
Then we found one cause. The code served ARC one tier harder than the player's current level, so anyone who did well on the first 9 hit a Boss-level ARC on question 10. Questions 7 and 9 also served Boss-level trivia to strong players. We capped everything before question 20 at tier 3 (Boss is tier 4).
| Drop-off, Q6–15 combined | Past Q15 | Completion rate | |
|---|---|---|---|
| Before cap | 21.8% | 64.1% | 45.3% |
| After cap | 18.7% | 66.0% | 45.3% |
Midgame drop-off fell 3 points, but completion rate didn't move at all: some people just quit later instead. Not a failure, but a reminder: a better intermediate metric doesn't mean a better final metric. A few hours later we learned that lesson again, and it hurt more.
The same table showed that "A Car for $1" at question 11 lost twice as many people as the other two chats, so question 11 now only picks between the AI snack shop and "Representative."
An hour and a half after launch, of the 512 people who had finished, 34% got the GPT-5 Type, 28% the Grok Type and 22% the GPT-4o Type. The top three added up to 84%.
The reason: the persona baselines had been simulated with random picks, and real players love troll options, so everyone skewed toward Grok and 4o. Claude first added raw persona features to the finish event (just 25 numbers, no option text), then recomputed the baselines once we had 1,747 answer sheets.
After recalibration, the top three added up to 51%, and the eight personas ranged from 17% down to 6%. The most common were 4o and GPT-5, the rarest was Kimi. What I said in Part 1, "keep it honest, but don't let the distribution get ridiculous," only now had real data under it.
In the early evening I asked Claude whether we should add more languages, like French. It pulled the data and argued against it: France had 10 visitors total, and the Spanish version had been opened by 7 people since launch. What we were actually missing was Traditional Chinese: Hong Kong and Taiwan together had 625 visitors, 17% of the total, 6 times more than Korea. They could only read the Simplified version, and 100+ of them had been routed to English.
I said: "just do it now, traditional and french."
French went the way Claude predicted: 31 visitors and 9 finishes by midnight. At least it was cheap and barely took time away from the main work.
The plan for Traditional was to auto-convert with OpenCC, then spot-check. I cut in:
isn't spot-checking kind of sloppy
Then two more:
also for traditional chinese, hong kong doesn't say it that way i think
the wording is different too right. like software is 软件 vs 软体
So instead we sent out 14 reviewers, one pass for Taiwanese usage and one for Hong Kong usage, checking line by line. They found 600+ problems:
The Traditional version has a single "繁體" entry and switches vocabulary based on the visitor's region: 軟件 or 軟體 for software, 的士 or 計程車 for taxi.
By midnight the Traditional pages had 742 visitors and a 47% completion rate, higher than the Simplified pages' 42%.
The next morning, completion rates across languages were between 40% and 60%. English alone was at 18% to 23%. The English pages also had the lowest start rate: 57.5%, versus 74% for Chinese and 82% for Korean.
sounds good. give the english side special treatment, they're probably even less patient
could also be the memes and the language context, take a look
Claude first ruled out "Americans just don't like it": US IPs overall finished at 41%, it was only the English pages that were low. And English drop-off was even, 4% to 7% per question across the first 6, not one broken question. Conclusion: it was the content.
We came up with four hypotheses. The opening was all 2024 memes English AI Twitter was long sick of (strawberry, 9.11). The trivia was dated internet factoids (tomatoes are berries, goldfish have a 7-second memory). English runs much longer than Chinese and is tiring to read on a phone. And translationese: the jokes died in translation.
Then we sent two teams of "English editors" to rewrite for English AI Twitter's taste. The opening got fresh 2025 memes ("how many b's in blueberry" from GPT-5's launch week, "5.9 = x + 5.11"). 27 of the 41 trivia questions were replaced. The Three-Body Problem was swapped for Dune. And they wrote a 40-line English style guide.
The first look at the data after the rewrite was mixed:
| Left at | Before rewrite | After rewrite | Question |
|---|---|---|---|
| Q2 | 3.8% | 7.0% | 9.11 vs 9.9 → 5.9 − 5.11 |
| Q3 | 8.2% | 13.4% | Slop Check |
| Q5–11 combined | 23.6% | 12.0% |
Drop-off in the middle and back roughly halved, but more people left at the start, and overall completion fell from 26% to 20%. Q2 drop-off doubled: solving an equation takes a lot more thought than eyeballing which number is bigger.
The fix: Q2 went back to 9.11, and all 22 Slop Check questions had their options cut to 60 characters or less. By evening completion was back to 25.4%, close to the pre-rewrite 26.8%.

Two lessons. First, the meme you think is stale may be exactly what gets people in the door: strawberry and 9.11 really are worn out on English AI Twitter, but because everyone gets them, they belong at the entrance. Second, the value of a per-question table is that it breaks "worse overall" into which questions got better and which got worse. Looking only at the total, we would have concluded "the rewrite failed" and thrown out the back-half improvements along with it.
At noon the next day I had Claude pull a "channel × drop-off" table:
| Opened in | Past Q12 | Finished | Finished, with Jev | Finished, without Jev |
|---|---|---|---|---|
| Regular browser | 73% | 48% | 66% | 40% |
| 68% | 42% | 73% | 32% | |
| 62% | 37% | 65% | 27% | |
| X | 57% | 30% | 58% | 20% |
(QQ is Tencent's other big messaging app.) By question 12, where Jev's first invite shows up, 43% of X visitors were already gone. Most never saw the invite. Claude proposed an A/B test. I made the call:
1. these players get the Jev invite at Q4... opened in X's in-app browser; opened in WeChat's in-app browser; came in from an X link; everyone on the English version. 2. everyone else stays at Q12. I think that's it, go with this strategy. no need for AB, just watch the AA numbers
(By "AA numbers" I meant the overall numbers, not the AA score.)
Over the next two hours Claude reported twice: in the channels with the early invite, getting past question 15 was up about 4 points, and "midgame retention is improving."
A bit over an hour after that, it overturned itself:
What I said earlier, that "the midgame is keeping more people," doesn't hold. When Jev turns on it skips up to 15 questions at once, and skipped questions also count as "reached"... The only reliable metric is completion rate. After subtracting the control group's natural drift, completion actually went down about 2 points with the early invite.
Jev usage did go from 26% to 49%, and past-Q15 went from 60% to 67%, but that was all Jev skipping people forward, not people actually staying. Completion went from 37.6% to 36.8%, while the control group went from 46.0% to 47.2% over the same period.
I replied: "ok, I think we can revert." Every channel went back to the Q12 invite.
This was the only time in the whole project that the AI admitted on its own that it had misread the data and suggested rolling back a change. Looking back, skipping A/B bought us speed. The price was relying on control channels to subtract time-of-day effects, and nearly getting fooled by a metric definition.
The next morning, Claude noticed that 11% of players who finished had played two or more times, so it calculated the odds of seeing a repeat question. Between friends, persona chats repeated 50% of the time, iconic chats 40%, Slop Check 33%. The most visible content on the share cards came from the three smallest pools in the bank.
The first two rounds doubled those three. At 10 AM I said:
expand it! I'd say grow the question bank by at least 45%. what do you think?
more questions, same quality! needs to be just as meme-heavy as before
Claude ranked pools by "questions drawn per run ÷ questions in the pool." The most repeat-prone were chart questions (4 per run from a pool of 9) and computer-use questions. The target went from about 404 questions to 589, adding 185, or 46%. Then it set up a pipeline:
Work started at 10:30 AM and went live at 1:30 PM. The brief for every step is in prompts/03-agent-briefs.md in the open-source repo, ready to reuse.

The most interesting part of the pipeline was the reviewers catching errors in the source:
The translation process ended up fact-checking the original.
We had one accident too: a translation group ran a full build, and the live Traditional site auto-converted a fresh copy from Simplified, including new questions that hadn't been reviewed yet. As soon as we caught it, we sent reviewers to patch it.
The next morning I gave the 8 personas their own color themes: a red scarf for the Doubao type (Doubao is ByteDance's chatbot; this persona is the Siri type on the English site), a four-pointed-star background for the Gemini type, slashes and caution tape for the Grok type. They went to a private preview first and only shipped after I reviewed them.
Two hours later I got the Gemini type myself, and the share card failed to generate:
I got chinese 1.5B dense gemini persona and it failed. no idea why. or is it related to ip?
safari
it was fine last night at least. no idea what happened
Claude found the cause within two minutes. Gemini's stars and Grok's slashes were SVGs embedded as CSS backgrounds. Once Safari draws that kind of image onto a canvas, it won't let you export the canvas as an image. Doubao's red scarf happened to use a regular image tag, so it was fine. We had always tested in Chromium and never triggered it.
For the two hours between the themes going live and the fix, everyone on Safari who got the Gemini or Grok type couldn't export a share card. The fix: swap the three textures to PNG, test all 9 themes under Safari's engine, and add an "export failed" event.
That event then caught another problem: for a few people, the screenshot library hadn't loaded from the public CDN. After adding a fallback that loads it from our own domain, export failures went to zero.
1. Find the biggest lever first. For a quiz that takes 20 minutes, the biggest lever is completion rate. Share rate was already high (more than half of finishers share), so more finishers means more shares on its own.
2. A per-question table is far more useful than totals. The 12% on Q3, the Boss ARC on Q10, the doubling on Q2 in English: none of that ever shows up in totals.
3. Metric definitions lie, and they do it very naturally. Questions Jev skipped counted as "reached." "63% with Jev vs 23% without" had selection bias. The first few tables had the wrong time zone. The completion-rate threshold differed at different times. Each one nearly led us to the wrong conclusion.
4. I handle direction, the AI handles measurement. Several key decisions came from one line of mine: "not fewer questions," "isn't spot-checking kind of sloppy," "it should jump a whole batch." Several key findings came from Claude: the Q3 drop-off, French not being worth it, and the time it overturned itself. It cuts the other way too: my call to skip A/B nearly left a regression live, and the chat Claude put at Q3 on the "funniest stuff first" rule drove away 12% of people.
5. If you can't see an effect, don't count on it. More on this in Part 3: some features didn't move the data at all after launch. We kept them, but stopped spending time on them.
After 36 hours, completion rate had gone from around 29% at the start to a steady 45% or so. Traffic started its natural decline on the second evening.
But before that, something else was happening at the same time: players were passing it around. 40% of new visitors came through a friend's challenge link, each Japanese player who finished brought in 2.3 new visitors on average, and 70% of the people who came in through a challenge link lost to the friend who sent it.
That's Part 3.
Alex (@Alex_ybuild)

Last of three. Part 1 went from 4 messages to launch, Part 2 covered changing the product based on data. This one is about how it spread, what 12.7K answer sheets show, what data we collected and what we didn't, and who did what between the human and the AI.
First, the final numbers (as of the morning of September 29):
At 10:30 PM on launch night, I asked Claude:
hehe is the traffic decent? my account only has a few hundred followers
It pulled the referrer data: over 90% of traffic came from outside my followers. The tweet only lit the first batch. After that, visitors mostly came from players bringing each other in.

The traffic curve is textbook. Tweet at 3 PM, peak at 10 PM with 2,398 unique visitors in that hour. The next morning, Japan's commute took over, the Japanese pages hit 500 to 600 people an hour, and Japan became the #1 source. Then the standard meme curve: no evening peak on day two, and hourly visitors fell from 1,200 to 600.
On day three I asked Claude: "I guess this is just the natural decay of meme marketing?" Yes. Here's why the decay was inevitable.

The sharing path was set by one line from Part 1, "whatever gets shared should open straight into playing. closed loop":
Conversion at each step:
That last number is the key. It's under 1, which means that without outside traffic, each round of sharing is smaller than the one before. A 20-minute quiz where only 40% of people who open it finish: that step pushes the viral coefficient below 1. That's the life cycle of a meme, and it's why we put so much effort into completion rate in Part 2. Every extra point moves the coefficient a little closer to 1.
Some channel details:
Japan. Each Japanese player who finished brought in 2.3 new visitors and 0.52 new finishes on average, the highest of any language. More than half of new Japanese visitors came through a friend's share. The data further down explains why: Japanese players really like sending the link directly.
WeChat and QQ. Of people who opened it inside WeChat or QQ (Tencent's other big messenger), 40% came in through a challenge link, mostly by scanning the QR code. People who finished inside WeChat shared 1.15 times on average, the highest of any app. Claude warned that WeChat might block the domain and take the main site down with it, so we prepared a backup domain and a switch that only changes links generated inside WeChat. We never needed it.
Xiaohongshu (RedNote). Images with a QR code or URL get throttled, so there's a separate "Xiaohongshu version" of the cards with no off-platform info at all.
X. When the share link is posted on X, Telegram or Discord, the preview image is your own result card, generated on the fly on Cloudflare.
On the morning of day two, traffic started to fall, and I asked Claude:
what else can boost sharing and the viral loop? want to give it a second wind
We built a few things. Some worked, some didn't.
can we add a secret one like a blind box? like at the end you show me a 3x3 grid and the last one is a question mark
Beyond the 8 personas, we hid one more: The Human Type, nicknamed "The One That Got Away," with the line "Detection failed: zero AI flavor." A neat flip on the name HumanBench.
The trigger was worked backward from real data. Claude first tried it on about 3,000 answer sheets, and even the strictest version matched over 8% of people. Too common. After tightening, it came out to about 2.6% on 6,846 answer sheets. The conditions: almost never picked any model's signature AI voice, Based well above the average player, and at most 1 combined across Sycophant, Preachy, Yapper, Confidently Wrong, Jailbroken and Ignores Instructions. The actual rate after launch was 2.8%.
I corrected the wording twice. Once:
I feel like we shouldn't spell out rarity percentages. some people might find that offensive. better to only label the secret one.
The original plan labeled the 8 regular personas too, like "Common 19.6%" and "Rare 7.4%." We changed it so only the secret persona gets a label. The other time:
oh and don't call it "pulling"
"Can you pull it?" became "Could it be you?", and "Secret persona obtained" became "You are the secret one." What you get is a read on you, not a prize from a gacha.
Effect: the share of people who played again within 1 hour of finishing went from 14% to 18%. Of people who played two or more runs, 87% got a different persona the second time.
If you came in through a friend's challenge link, you get an extra PK card when you finish: both of your AA scores on the same leaderboard chart, a side-by-side on 12 benchmarks, and a verdict sticker.
The card went through four versions, each pushed along by a line from me:
isn't the friend pk card kind of bare? and does it affect the usual four cards? how do they relate?
also if it's from a challenge page, just show the benchmarks and the aa index. put both people in one table and one chart. the web version can have some subtle animation. and this pk card should be an extra card, not a replacement
stop saying you and them. just use the model names. also why can't the aa part be more polished, with a model name on every bar?
honestly the layout could be more comfortable. like the table doesn't need to be full width. make it feel natural
now it feels kind of empty. can't you just lay it out better? you can change the layout, add more info
The verdict stickers are my favorite part. Smaller model that still wins: "Giant killer · That's distillation." Bigger model that still loses: "All params, no brains · That's padding." A narrow win: "Photo finish · Won, but within benchmark noise." A tie: "Draw · Both launch events claim SOTA."

Effect: compared with a control group from the same hours, people who were challenged shared at a rate about 8 points higher, and the gain was in "share link": people who lost were more willing to send their own link back. 43% of people exported the card set that included the PK card.
If it doesn't show an effect, it stays, but we stop spending time on it.
A test where people play the AI ended up with 12.7K complete answer sheets. We treated it as a (very unscientific) "human benchmark" and dug up some fun stuff.
First, the caveats: everything below is aggregate statistics. We didn't look at any individual's answers. Times are converted to the player's local time (Chinese pages use Beijing time, Japanese and Korean pages use Tokyo/Seoul time, English pages are estimated from the IP's country). This is a joke test, not research. Read it for fun.

The bigger your result, the more likely you share: 40.5% of people at 14B or below shared after finishing, versus 61.8% of people at 10T or above.
That has a funny side effect. In Friend PK, 69% of people who came in through a challenge link lost to the friend who sent it, by 7.3 points on average. Not because the challenged are weaker. The people confident enough to send the link were the high scorers to begin with. Survivorship bias.

When the same person plays a second run, their AA score goes up 4.0 on average, and 72% do better the second time. For people who played three or more runs: 32.7 on the first, 36.6 on the second, still 36.6 on the third. One retake and you've maxed out.
This is the human version of data contamination: seen the questions, score goes up. LLMs get roasted for gaming benchmarks. Humans are no different.

People who finished in under 10 minutes averaged AA 18.1 with 43% accuracy. People who took over 30 minutes averaged AA 35.8 with 71% accuracy. Same as reasoning models: the Thinking version beats the Flash version. (The page really does add a suffix to your model name based on your average time: fast gets -Flash, slow gets -Thinking.)

Comparing people who finished between midnight and 4 AM with people who finished between 9 AM and 6 PM:
Late-night humans: more emotional, easier to talk into things, not any dumber.

Grouped by page language, looking at the average tags per person:
This compares answering style, not how smart anyone is.

By local time, the busiest hour is 5 PM, then 11 PM to midnight. The quietest is 5 to 6 AM.
Of people who finished on Monday, 51% did it during local work hours (9 AM to 6 PM). Korean players were highest: 66%.

Among shares, 26% of Japanese players sent the link directly, versus only 14% of Chinese players, who mostly saved the image. When you send the link, your friend opens it straight into a challenge. That explains why Japan had the highest viral coefficient.

People who got the secret persona posted the most (60.7%). It was hard to get, so of course you post it. People who got the Doubao type posted the least (48.5%). Doubao is ByteDance's chatbot, and this persona is the Siri type on the English site. Maybe nobody wants to admit they're "like Doubao." The overall gap isn't big, though.
A few more odds and ends:

The most common result was the GPT-5 Type (20.3%), followed by the Grok Type and the GPT-4o Type. The secret persona shows up about 3 times per 100 runs. The most common parameter count is 671B, the same size as DeepSeek V3. Average accuracy was 67%, and only 103 people got everything right.
This section matters more than everything above, so I'll be specific.
What we collected:
What we didn't collect:
The analytics database isn't public. The open-source code includes the analytics scripts, but no data.
One thing I have to own. At launch, the small print on the home page said "Runs entirely locally, answers never uploaded." That line was left over from an earlier version, and since we have anonymous analytics, it wasn't true. Claude pointed this out early on day two while prepping the English promo, and at the time I only had it changed on the English version. It wasn't until that afternoon that all 7 languages were changed to "No signup · For fun: params ≠ brains." From being flagged to being fixed, the false line stayed up on the Chinese version for about 14 more hours.
The data in this retro is all aggregate statistics: ratios, averages, distributions. We never looked at any individual's answers, and no number here maps to a specific person. The open-source prompts were scrubbed of personal info and private links.
Also, someone dug through the live source code. Web page code gets sent to the browser anyway, so that's fine. But at the time, the source comments included the secret persona's trigger conditions and per-question drop-off data. We changed the build to strip comments automatically (the page also went from 579KB to 481KB), and checked that there were no keys anywhere in the page. Now the code is open source, so those comments are public again, and you've already seen the secret persona's conditions above.
Across all three parts, I sent 237 messages for the whole project, about 8,600 Chinese characters in total, and wrote no code by hand. I've been using these models since the GPT-2 days, so I have a decent sense of what they can do and where they fall over. This time I deliberately put myself in the "product owner plus reviewer" seat and handed all the code to the AI.

What I did:
What Claude did:
Claude's classic failure modes:
If I had to sum it up in one line: the human says "this is wrong," the AI makes it right and proves it's right. In this split, judgment and taste sit with the human. Speed of execution, measurement and correction sit with the AI. We shipped dozens of versions in three days, and every change first ran a full automated playthrough locally (including exporting share cards) and only deployed if it passed.
Everything is on GitHub: https://github.com/alex-ybuild/humanbench
If you want to make your own joke test, take it and change it. If you want to see how one person directed a coding agent to build a product from scratch through conversation alone, the prompts/ folder might be more direct than these three posts.
On the last day, I asked Claude: "hehe but as a campaign it went pretty well, right?"
An account with a few hundred followers, no ads. In three days it got roughly a million impressions, 12.7K people seriously finished 76 questions, and it made it to 96 countries and regions. Traffic has settled back to one or two hundred people an hour. As a bit, it did its job.
For me, the bigger takeaway is the process itself: an idea, a 4-message prototype, a product in 237 messages, dozens of versions changed by data, and then everything made public. Through all of it I did pretty much one thing: tell the AI "this is wrong," then judge whether the fix was right.
Come find out how many B you are, and feel free to take the code and build your own bit.
Alex (@Alex_ybuild)