How to Validate an AI Product Idea Before Paying for OpenAI Tokens

Validate an AI product idea before paying €0.01 in inference costs — landing pages, fake doors, pre-sales. The pre-token validation playbook.

10 min read

"Just give all of your tokens and all of your money to an AI Claw bot that will just waste millions and millions of tokens." That's Kevin McGrath, CEO of Meibel, on CNBC last month, warning founders that the default architecture — every input flows through an LLM, every step burns tokens — is unit-economics suicide. He's not the only one worried. Anthropic itself reportedly posted a negative 94% gross margin in 2024 (per The Information) — i.e., the model provider lost more on inference than it earned on revenue. Then it cut its 2025 margin projection from 50% to 40% because compute costs ran ~23% above plan.

If the people selling the tokens can't make the math work, an indie founder paying retail rates can't either. Not without checking, first, whether anyone wants the product at all.

The first inference call you should ever pay for is the one that fulfils a real customer order, ideally a paid one. Everything before that runs without a token. This piece is the playbook for the "everything before that" phase — five steps, real founder examples, with what to spend (mostly time and €100 of ads), what not to build (the API integration), and a kill criterion at the bottom of each step.

If you've already read validate or build an MVP first, this is the AI-specific overlay. If not, start there for the order-of-operations argument; this article assumes you've decided to validate before you build, and you want the AI-flavoured version.

AI products fail two ways, not one

Generic SaaS validation asks one question: do people want this? AI validation asks two: do people want it, and can the model actually do it well enough? The Andrés Max framework calls these "problem validation" and "technical validation" and runs them in parallel; that bit is right and most articles get it wrong.

Skipping the second one is how you end up with Builder.ai. Pitched "AI-built apps" via an assistant called Natasha, raised ~$445M, then Rest of World revealed that ~700 human engineers in India were doing the work — instructed to align communications to UK business hours and avoid Indian-English colloquialisms to maintain the AI illusion. Revenue was reportedly inflated 300% ($55M actual vs $220M reported). Creditor Viola Credit seized $37M in May 2025; the company collapsed into insolvency.

The mistake wasn't using humans behind the curtain. Wizard-of-Oz is a valid technique — we'll come back to it. The mistake was raising $445M on the lie that the AI worked when it couldn't. Builder.ai inverted the validation order. They sold the AI before they had it. The right order is: fake the AI to validate demand, then build the AI to fulfil it. Never invert.

Skipping the first question — do people want it? — is how you end up Jasper. AI copywriting wrapper, $125M raised, then ChatGPT shipped at $20/month and Jasper's $49/month "templates and a marketing-focused UI" stopped justifying the gap (Felix Neumann's analysis is the cleanest writeup). Layoffs and revenue contraction in 2023–2024. The demand existed; the defensibility didn't. Validation should have asked: if OpenAI ships this in their next release, do I still have a business?

Two failure modes, two questions. The five steps below test both, before the meter starts.

Step 1. Run twenty manual outputs and read them

Total cost: zero of yours, maybe €5 of your own ChatGPT Plus quota. Time: an evening.

Take 20 realistic inputs — actual emails, actual contracts, actual whatever-your-AI-product-eats. Run them through ChatGPT or Claude yourself, in the web UI, and grade the outputs. Hand-grade. Three buckets: "good enough to ship", "bad but fixable with prompt work", "the model can't do this yet."

This is the technical-validation step. If 14 out of 20 are in bucket 3, the model isn't ready for your product. Stop. Wait six months for the next generation of models, or pick a different feature where the model already works. We've seen founders skip this step, integrate the API, and discover at customer #3 that the failure rate is 30% — too high for the use case, but invisible until real users hit it.

The 2024–2025 winners came in around 4 out of 20 in bucket 3, sometimes lower. The 2024–2025 graveyard was littered with founders who built around 14-out-of-20 capability, told themselves "GPT-5 will fix it", and burned through their seed before GPT-5 fixed it. Some model gaps stay open for years (legal contract review, multi-step financial reasoning, medical triage as of writing); some close in eighteen months. Your roadmap can't depend on the timing.

Kill criterion: more than 8 out of 20 outputs are in bucket 3. Try a different feature, or wait.

Step 2. Wizard-of-Oz the output to a real customer

Don Norman coined "Wizard of Oz" in 1973 for IBM's speech-to-text experiments — a researcher manually faked the system's responses while the user thought they were talking to a computer. The technique is forty years old and still the highest-signal validation method for an AI product. Most founders skip it because they feel it's "cheating." Builder.ai cheated investors. You're cheating your backlog.

The play: take the same 20-input set from Step 1, run them through ChatGPT yourself, paste the outputs into emails, and ship them to 5–10 named prospects who said the problem was real. Don't tell them the back-end is you. The cost: zero tokens of yours (the prospect's request goes through your own ChatGPT Plus subscription). Time: two evenings of copy-paste.

Aardvark is the textbook precedent. The peer-to-peer Q&A app routed every user question to a human respondent and pasted answers back, simulating an algorithm that didn't exist. Validated the matchmaking premise on real users. Acquired by Google for $50M in 2010. Early Wealthfront ran the same play pre-robo-advisor, hand-recommending investments to early clients, validating both demand and pricing before automating; they now manage roughly $50B AUM. The concierge phase was the validation phase.

What the test surfaces: whether the output, in the form your customer would actually receive it, prompts the action you need. Did they reply? Did they share it? Did they ask for the next one? Did they offer to pay? Spontaneous "where do I sign up?" within five concierged outputs is the strongest possible signal — stronger than any landing-page CVR.

Kill criterion: zero unprompted "I'd pay for this" replies across 10 hand-delivered outputs.

Step 3. Build the page for the AI feature you haven't built

This is the demand-validation step. You've confirmed (Step 1) the model can do it and (Step 2) ten humans liked the output. Now: do strangers care?

Build a one-page pitch describing the AI feature. Headline, three benefits, a screenshot or short video of the output (which exists, because you faked it manually in Step 2), a CTA. The CTA matters: don't make it "Join the waitlist." Make it filter for willingness-to-pay — early-access pricing with a refundable deposit, or a pre-sale link.

Wire €50–€100 of paid traffic to it from Reddit, Meta, or Google Search. Watch what happens.

The honest measurement: CTR validates curiosity, CVR validates intent, paid pre-orders validate willingness-to-pay. The order matters. Most fake-door articles celebrate high CTR. The sharper take: for AI features specifically, CTR validates curiosity, not willingness to pay. A B2B SaaS that fake-door-tested "AI-powered analytics" got high signups, then 89% of waitlist subs cancelled within 30 days when the feature actually shipped. AI buttons get clicked. Paid AI features get scrutinised.

This is also where LemonPage fits. The cheapest-method version of this step is a Carrd page wired to a separate ad account — page on one tool, ads on another, measurement in a spreadsheet. We built LemonPage because we lost four hours of plumbing per test, and the whole pre-token validation math only holds if running the test stays cheap. Page + Reddit/Meta/Google ads + measurement in one workflow. Carrd Pro Lite at $9/year and Framer at $15/month are real alternatives if you don't mind the plumbing — pick whichever lets you run more tests, faster.

For the budget question — what €100 of Meta Ads actually buys you in 2026 — see the $100 Meta Ads test. The pre-sale CTA mechanic is in take payments before code.

Kill criterion: under 2% CVR after 1,000 visitors, or no deposit conversion at all on a list of 100+ qualified leads.

Step 4. The first €50 of tokens go on Haiku, not Opus

The "validate before tokens" framing has a quiet flaw: most founders read it as "avoid token spend entirely until launch." That's not what the math says. The math says avoid the token spend whose only purpose is your own learning. By the time you've passed Step 3, you have signed-up customers expecting an output. They'll get it. The question is which model and which budget.

Default to the cheapest tier that survived Step 1's 20-input test. As of writing: GPT-4o-mini, Claude Haiku, Gemini 2.0 Flash Lite (free 15 RPM). A 100-user validation test on Haiku costs single-digit dollars. The actual budget question isn't "should I avoid tokens" — it's can I get to first paying customer for under €50 of API and €100 of ads?

If the answer is no — if your use case genuinely requires Opus or GPT-5 to produce a shippable output — flag it. That's a unit-economics warning before you launch. The Cursor investment-firm analysis reportedly showed roughly $650M paid to Anthropic against ~$500M revenue in the bleeding window — a negative 30% gross margin. Cursor crossed $1B ARR by November 2025 at a $29.3B valuation, and analysts now project margin recovery to 74–85% by 2027 by mixing in cheaper / open models. The recovery is real. The bleeding window was also real, and Cursor had ten figures of capital and a category lead. An indie founder doesn't.

The compounding fact: a POC that costs $50 in tokens during validation can scale to $2.5M/month at production volume, depending on per-user calls and the model tier. Treat that as illustrative, not as a measured single case — but treat it. Step 4 is where you find out, on cheap tokens, what the eventual scaled bill looks like before the cheap tokens become expensive ones.

Kill criterion: customer-acquisition cost (paid traffic) plus inference cost per first-month customer exceeds 2× their paid price. The unit economics aren't going to fix themselves at scale.

Step 5. The Levels counter — validate demand, then plug in the API

Pieter Levels' Photo AI is the cleanest counter-example to "AI startup = bleeding margin." Launched February 2023, hit ~$5.4K MRR in week one, $132K+ MRR by November 2025. He doesn't run his own Stable Diffusion infra; he uses Replicate API at ~$40/month in infrastructure for compute. Validated demand on a cheap landing page and his Twitter audience first; plugged in the API as a variable cost, not a fixed-infra bet.

His repeated principle, paraphrased from the Indie Hackers case study: ship within two weeks, then check if there's demand and if real people pay — only paying customers validate an idea.

Two patterns from the Levels playbook that translate:

  1. Variable-cost inference, not fixed infra. Replicate, Hugging Face Inference, OpenAI/Anthropic direct APIs — pay per call, not per GPU-hour. Don't build your own inference stack until you have a paying-customer problem big enough to justify it.
  2. Audience first, model second. Levels had a Twitter audience before he had Photo AI. The audience reduced the cost of Step 3 (paid traffic) toward zero. If you don't have one, the meta-ads test is the substitute. If you do, use it.

The third pattern, less often quoted: Levels has openly killed 70+ products to keep the five that now generate over $3M/year combined. The kill rate isn't a failure — it's the validation working. The five steps above exist so you can kill the bad AI ideas cheaply enough to free the calendar for the next one.

What total spend looks like, end to end

A clean run through the playbook:

StepOut-of-pocketTokens used
1 — 20 manual outputs (graded)€0Your own ChatGPT Plus / Claude Pro subscription
2 — Wizard-of-Oz to 5–10 prospects€0Same
3 — Landing + €50–€100 paid traffic€50–€100 ads + €0–€20 pageNone
4 — First Haiku/Flash Lite calls for paying users€5–€20 in APICheap tier only
5 — Audience-led, variable-cost inferenceWhatever scalesPer-call, not per-GPU

End-to-end pre-launch spend: €55–€140, plus your time. The first OpenAI invoice that matters is the one a paying customer's request triggers. Everything before that is your own subscription quota and €100 of ads.

Compare that against the CodeParrot end of the distribution: YC W23, $500K raised, Figma-to-code AI tool that pivoted multiple times, peaked at ~$1,500 MRR, shut July 2025. Or Builder.ai's $445M end. CB Insights' canonical "Why Startups Fail" research reports 42% of startup post-mortems cite "no market need" as the failure reason — and the 2024 update on 431 VC-backed shutdowns since 2023 found 43% failed for poor product-market fit. The five-step playbook is a €100 hedge against being part of that statistic in the AI category.

Recap and the kill funnel

Five steps, in order, with kill criteria written down before the test:

  1. Capability test — 20 hand-graded outputs. Kill if more than 8 are bucket-3.
  2. Wizard-of-Oz — 10 hand-delivered outputs. Kill if zero unprompted "I'd pay" replies.
  3. Demand smoke test — page + €100 ads. Kill if under 2% CVR after 1,000 visitors.
  4. Cheap-tier fulfillment — Haiku / Flash Lite for first paying users. Kill if CAC + inference exceeds 2× paid price.
  5. Audience-led scale — variable-cost inference, no own infra until paying-customer demand justifies it.

The order matters. Step 3 before Step 1 means you sell something the model can't deliver — Builder.ai. Step 1 before Step 3 means you've built something nobody wants — Jasper, in the form of the wrapper that ChatGPT subsumed. Run Steps 1 and 2 in parallel if you must, but pre-commit each kill criterion before you start.

For the wider field — the ChatGPT-wrapper question, the unit-economics shift since 2023, and the seven pre-MVP validation methods you can mix into Steps 2 and 3 — see the AI cluster on the blog.

Validate your AI idea on LemonPage — page + ads + measurement, in one workflow. We built it because we kept losing four hours of plumbing per test, and pre-token validation only holds if the test itself stays cheap.

OpenAI doesn't care if your AI product has buyers. They charge by the token either way. Make sure the meter only starts when someone else has already paid.

FAQ

How much should I spend on tokens during AI product validation?

Under €20 of API spend, in total, before your first paying customer. Use your own ChatGPT Plus or Claude Pro subscription for Steps 1 and 2 (manual outputs and Wizard-of-Oz). Run Step 3 (paid-traffic landing page) on €50–€100 of ads, no API spend at all. The first real API spend happens in Step 4, on the cheapest tier — Haiku, GPT-4o-mini, Gemini Flash Lite — and only for paying users. End-to-end pre-launch budget: €55–€140 plus your time.

What if my AI product needs a frontier model like Opus or GPT-5 to work at all?

That's a unit-economics warning, not a stop sign — but treat it like one. Cursor reportedly ran a negative 30% gross margin while leaning on Anthropic's frontier tier, and Cursor had ten figures of capital and a category lead to absorb the bleeding. An indie founder doesn't. Check whether your use case can be split — frontier-tier for one critical step, cheap-tier for the rest — or whether the price tag of the eventual product can carry frontier-only inference at 80–90% gross margin. If neither math works, the product isn't viable yet.

Is using ChatGPT manually behind the scenes the same lie as Builder.ai?

No, and the difference is fundraising. Wizard-of-Oz validation, where a founder manually generates outputs for 5–10 prospects to test demand, is a forty-year-old technique with a working name and an academic paper behind it. Builder.ai's failure wasn't that 700 humans were doing the work. It was that the company raised roughly $445M on the lie that the AI worked, and reportedly inflated revenue 300%. As a pre-launch validator, Wizard-of-Oz is gold. As a $445M pitch, it's fraud. Validate with humans, build with AI, never invert.

How do I validate a B2B AI product where customers won't pay before seeing it work?

Combine Wizard-of-Oz delivery (Step 2 here) with cold ICP outreach from the seven pre-MVP methods. The ask isn't "buy this." It's "we're testing a tool that drafts X for you — want to see it on your real input?" The output is real (because you produced it manually in ChatGPT). The decision metric is whether they ask for the next one and whether they're willing to scope a pilot. B2B contracts at €500+ ACV justify the slower signal that comes with this method.

Can free-tier APIs really cover a 100-user validation test?

Often, yes. Gemini 2.0 Flash Lite has a free tier of 15 requests per minute as of writing. Anthropic Haiku and OpenAI GPT-4o-mini run a few cents per million tokens — well under €5 for 100 typical text-output requests. The exception is image/video generation, where a 100-user run on Replicate-hosted Stable Diffusion will run into the tens or low hundreds of euros. Check your specific use case before assuming "API spend ≈ €0", but for text-only AI features the answer is genuinely close to zero.

Which step do most founders skip, and why does it matter?

Step 2 — Wizard-of-Oz delivery to real prospects. It feels like cheating, and it's slow (two evenings vs. five minutes to wire the API). Skipping it means founders learn at customer #3 — after the API is integrated, the page is up, and the deposit is taken — that the output, not the API, was the part that didn't land. The Wizard-of-Oz step surfaces the output-quality problem before the integration cost is sunk. Two evenings is the cheapest insurance in the playbook.