Journal

Design Doing: An Early Version of an AI-Accelerated Innovation Framework

Why AI makes user behavior cheaper to test than opinions are to collect.

Why AI Makes User Behavior Cheaper to Test Than Opinions Are to Collect

TL;DR: The cost of building a testable prototype has dropped by ~90% thanks to AI. It’s now cheaper to build and test than to research extensively upfront. Design Doing is a workflow where the prototype itself becomes the experiment: functional software deployed to real users to observe actual behavior, not opinions. Discovery and empathy still matter, but the confidence threshold to start building is lower (30–40%). The new bottleneck isn’t ideas or engineering. It’s user attention.

I’ve sat through enough Design Thinking workshops in German corporations to recognize the pattern.

  • Day 1: 8 people in a room. Sticky notes. Sharpies. A facilitator leads through “Empathize, Define, Ideate.” Everyone leaves energized. “We really understand our users now!”
  • Day 10: A prototype gets built. Beautifully designed. Internally validated. Users love it.
  • Day 50: MVP development starts. And gets delayed.
  • Day 310: Launch day.
  • The week after: 12 sign-ups. 0 returning users.

We workshop. We prototype. We validate. We ship. We fail.

But here’s what I believe changed in the last months: the entire economic equation of digital product innovation is flipping.

The Traditional Economics of Innovation

For years, the math was straightforward:

  • User research (recruiting, interviewing, synthesis): €10,000–€20,000 + 3–4 weeks
  • Building a functional MVP: €40,000–€100,000 + 2–12 months of developer time (depending on complexity)
  • Fully loaded costs: ~€50,000–€120,000

You needed at least 60% confidence before building because being wrong was expensive. So we built elaborate research processes. Recruit users, create personas and empathy maps, synthesize findings into reports, validate with stakeholders, get approval, then maybe start building.

This made sense when engineering was the bottleneck.

The problem? You can never achieve certainty until you launch the product. There are simply too many variables that no workshop and no clickable prototype can address. What does the product feel like when it’s real? Does the brand convey the right signal? Does the interaction feel premium or cheap? Does the onboarding create trust or friction? Desirability isn’t just about solving the right problem. It’s about tone, texture, timing, and a hundred micro-decisions that only surface when someone uses a real thing in a real moment. A Figma mockup can’t carry brand energy. A sticky-note prototype can’t simulate trust.

On top of that, by the time you’ve achieved these expensive insights, the market has moved. Your competitor tested three hypotheses while you were synthesizing interview themes.

The New Economic Reality

Here is what it costs to build a testable prototype today, and it will get even crazier in the future:

  • AI-generated behavioral prototype: €400 (API costs + tool subscriptions)
  • Your time: 5 days (~€2,000–€3,000 at typical innovation manager rates). Complex prototypes might take longer, simpler concepts can be testable in a weekend.
  • Deploy to 100 real users: €500–€1,000 in ad spend per week (for ~4–6 weeks)
  • Fully loaded costs (incl. time): ~€4,400–€7,400

The cost of testing a hypothesis has dropped by roughly 90%. Time-to-learning: weeks instead of months.

But most innovation teams are still using the decision-making framework from when testing was expensive.

The inversion: It is becoming cheaper to build and test than to research exhaustively upfront.

Think about that. A focus group costs €10,000 and tells you what people say they’ll do. For a similar budget, a behavioral prototype shows you what people actually do. Which data would you trust?

Why the Economics Flipped

The cost inversion didn’t happen by accident. It happened because building software is fundamentally changing. And the trends are clear. Bear with me, here comes the AI part ;)

Developers using tools like Cursor, GitHub Copilot, or Claude consistently report around 30–50% productivity gains. What used to take two weeks takes one. Boilerplate gets generated in seconds. Context switching drops. Learning curves compress. That alone shifts the build-vs-research calculation.

But here’s what matters more for innovation: the people who couldn’t code before can now build functional prototypes. Andrej Karpathy, former Director of AI at Tesla, called it “vibe coding.” Describing what you want in plain language and getting working software back. Tools like Replit Agent, Lovable, and Claude Artifacts make this real. Not mockups. Not wireframes. Working software that users can interact with.

A marketing manager with a hypothesis about a new customer portal no longer needs to write a requirement document, get budget approval, and wait six months for IT to build an MVP. She can build the prototype herself in a weekend and put it in front of real users on Monday.

For production systems, this isn’t enough. You wouldn’t run a bank on vibe-coded software. But for answering “will anyone actually use this?” It changes the game entirely.

The cost of testing a hypothesis has collapsed. Whether you’re a developer moving faster or a non-technical person building your first prototype, the old “research extensively before building” logic no longer holds.

The Code Is Disposable. The Behavioral Data It Generates Is Not.

I call these behavioral prototypes: functional software built not for production, but to observe what real users actually do in real contexts. Not a mockup. Not a demo. A working tool deployed to real people, designed to generate behavioral data rather than opinions.

To be clear: this is not an MVP. The term has always carried two competing definitions. Frank Robinson, who coined it in 2001, meant a real product: “the right-sized product for your company and your customer, big enough to cause adoption, satisfaction, and sales.” Eric Ries redefined it in 2011 as a learning instrument: “the fastest way through the Build-Measure-Learn feedback loop with the minimum amount of effort.” Corporations mostly went with Robinson’s version, whether they knew it or not. A behavioral prototype builds on Ries’s version but goes further: it’s not even meant to become the product. It’s meant to answer a question and then die. Confusing the two is how teams end up optimizing code that should have been thrown away. Or throwing away products that should have been refined.

Design Doing: An Evolved Workflow

I didn’t set out to reinvent anything. I just noticed that the tools changed, so I changed my workflow. Over the past year, this has evolved into what I’m calling Design Doing. I don’t know exactly where this leads. But the direction feels clear enough to share.

The term isn’t new. The shift from “Design Thinking” to “Design Doing” has been discussed for years, something we’d also explored with colleagues at Dark Horse. I like it because it captures exactly what changed: the “doing” itself became the research method. But my use goes beyond just “stop theorizing, start executing.” It describes a process where building is testing, and the behavioral prototype is the experiment.

The name isn’t about speed, though speed is a byproduct. It’s about a fundamental shift in what counts as evidence.

Now, if you’ve practiced Design Thinking seriously, you’ll push back: “DT was always about observing real behavior. Empathize means contextual observation, not focus groups.” And in theory, you’d be right. But think about what actually happened in corporate practice. When building cost €100,000+, teams couldn’t afford to use real products as research tools. Interview-based empathy wasn’t a methodological choice. It was an economic constraint. The gap between DT’s promise (“observe real behavior”) and its practice (“ask people in workshops”) wasn’t caused by bad facilitators. It was caused by budget reality.

Design Thinking validates with: Prototypes that simulate. Reactions in controlled settings.

Design Doing validates with: Prototypes that function. Behavior in real contexts.

That’s not a subtle distinction. It’s the difference between collecting opinions and observing behavior. Between hypothetical preference and revealed preference. Between “I would definitely use that” in an interview and silence after launch.

The philosophy remains human-centered. The focus stays on real needs. But the method of validation changes entirely. From asking to watching, from researching to testing, from confidence-building to rapid falsification.

This framework is a first draft, not a finished model. I’ve been working this way for about a year now, adjusting the process with every project. But what I can say with growing confidence: this workflow has been significantly more efficient and effective than what I did before. Not because it skips steps, but because it turns the most expensive step—building—into the research instrument. Instead of spending months building confidence before anything gets made, you make something that delivers confidence. The sequence flips. And with it, the way teams learn, decide, and iterate flips too. Consider this an invitation to experiment, not a recipe to follow.

Discovery: Understanding the User

The first step hasn’t changed: you need to understand the people you’re building for. Their needs, their frustrations, the problems they’ve stopped noticing because they’ve lived with them so long. That part is non-negotiable.

What has changed is the toolkit. In the framework, this is the Discovery phase: AI Research and Human Depth working iteratively, not sequentially.

Traditionally, discovery meant interviews, field visits, maybe a survey. Small samples. You’d talk to 10–15 people, synthesize the findings, and hope they were representative. It was the best we could do with the time and budget we had.

AI Research opens up new dimensions here. You can crawl Reddit threads, forum discussions, app store reviews, support tickets, and social media conversations. Thousands of real, unfiltered voices talking about their problems in their own words, unprompted. You can start spotting friction patterns in existing products that users might never mention in an interview because they don’t consciously notice them. It’s still early, and it’s not perfect. But the direction is clear.

What I’ve found is that AI Research doesn’t replace interviews… it reframes them. That’s where Human Depth comes in. Instead of walking in hoping to discover what the problems are, you walk in with a better sense of the landscape and use the conversation to understand the why behind the patterns. The emotional context. The lived experience. The things data can surface but never explain.

At its best, the combination works like this: AI Research gives you breadth. Many data points across forums, reviews, and usage patterns. Human Depth gives you exactly that. Depth. The trust, the nuance, the moments where someone reveals what they actually need rather than what they think you want to hear.

In a recent project in the sports sector, this played out exactly like that. Before a single interview, we used AI to scan forums, app reviews, and competitor landscapes. That gave us over 20 hypotheses to bring into conversations. Not guesses. Informed starting points. Five interviews later, we had nearly a hundred validated data points. But the moment that reframed the entire project wasn’t in any dataset. It was a user looking at her results and saying “Okay, but what do I do with this?” That pause, that frustration. That’s the depth no model can extract from a transcript.

Under the old economics, you needed high confidence before building because being wrong cost €100,000+ and months of engineering time. Aiming for 60% confidence before committing resources was rational risk management.

Under the new economics, being wrong costs €4,000–€7,000 and a few weeks of work. The calculus inverts.

And this changes the exploration phase itself. You no longer need to research until you’re sure. You need to research until you have a direction. An informed gut feeling, a strong hypothesis based on real signals, not exhaustive proof, is enough to start building. Because the prototype isn’t the commitment. It is the next round of research. The heavy upfront discovery process made sense when it was your last chance to course-correct before a six-figure bet. When the bet costs a few thousand euros and a couple of weeks, your gut feeling plus a working prototype will teach you more than another month of interviews ever could.

At the Confidence Threshold, 30–40%, you know enough to build something testable but haven’t yet invested the weeks of research that create emotional attachment to your hypothesis. You’re informed but not committed. Curious but not convinced.

  • Below 30%, you’re guessing randomly. There’s not enough signal to even know what to prototype.
  • Above 50%, you’re probably over-researching. Burning time to achieve certainty that only real-world behavior can provide anyway.

The goal isn’t to be reckless. It’s to recognize that additional research has diminishing returns once you’ve crossed the Confidence Threshold: the point where a behavioral prototype becomes cheaper than another round of interviews.

To be clear: this doesn’t mean skipping research. Below 30%, you’re building blind. No amount of cheap prototyping fixes a complete lack of understanding. Discovery still matters. You still need to talk to people, map the landscape, build hypotheses worth testing. The difference is what happens once you’ve crossed that threshold. That’s the Problem Reframing loop in the framework. Discovery doesn’t end when building begins. It continues through building. People struggle to articulate their problems in the abstract. But give them something concrete and they’ll show you where you were wrong. At 30–40%, you’re not committing to a solution. You’re creating the conditions for deeper discovery.

But one thing remains irreplaceable.

You cannot automate empathy.

You still need to be in the room. You need to see the micro-expressions, the pause before someone answers, the way they lean back when they say “yeah, it works fine” but their body says otherwise. You need to feel the energy shift when you touch on something that actually matters to them. You need to notice when someone is giving you the polished answer versus the real one. And know how to gently push past it.

AI can transcribe. AI can cluster themes. AI can summarize. But AI can’t sit across from a frustrated nurse at the end of a 12-hour shift and understand what she actually needs by watching how she navigates her workaround. It can’t recognize that the real problem isn’t the one she’s describing—it’s the one she’s accepted as normal.

The human skill in discovery isn’t gathering information. It’s creating the conditions where someone trusts you enough to show you the truth. That’s a relationship, not a data point.

Build, Deploy, Watch: Solution Testing

This is where things get genuinely different. And honestly, this is where I had to unlearn the most.

And yes, we built prototypes. These “learning” prototypes gave us something. We could show them to users. We could watch where they hesitated, where they said, “oh, that’s not what I expected.” We gathered insights. We built confidence.

But here’s what those prototypes couldn’t tell us: Would anyone actually use this in their real life?

A Figma prototype doesn’t answer whether someone will open the app on a Tuesday morning when they’re stressed and busy. A clickable mockup doesn’t show you if they’ll come back tomorrow. A demo doesn’t reveal if they’ll tell their colleague about it or quietly forget it exists.

We always had question marks left. Big ones. The kind that only real usage in real contexts could answer. But real usage required real software, and real software required real investment. So we made decisions based on incomplete data and hoped for the best.

I saw this play out directly in the sports sector. A team had built a technically solid product—full process, research, prototyping, internal validation. Everyone agreed it solved a real problem. They launched. The CEO told me later: “We put a tool where a solution was needed.” Users got results but nobody showed them what to do next. The product quietly disappeared. Not because the technology was wrong. Because the real need had never been tested with a real product in a real context.

That gap… between “users liked the prototype” and “users actually use the product”… is where most innovation projects went to die.

That gap has now collapsed.

Where the second diamond used to produce mockups in 5–10 days, you can now build functional software in the same timeframe. Not because we’re cutting corners. Not because we’re being reckless. But because the tools have changed so fundamentally that we can now Build Behavioral Prototypes, functional software, not mockups, and Deploy to Users in real contexts. In the time it used to take to schedule a user testing session.

And here’s what quietly disappeared in that shift: ideation as a separate phase. There’s no whiteboard moment anymore where you brainstorm concepts and then hand them off to someone who builds them. You have an idea, you start building it, you feel that it’s wrong while you’re building it, you delete it, you rebuild. The conversation with the tool is the ideation. You think by doing. Three ideas that would have been sticky notes on a wall are now three prototypes you can put in front of people. The thinking didn’t go away. It merged with the making.

That same sports project? We went from insight synthesis to 15 behavioral prototypes in under three weeks. Not mockups. Formats users could actually interact with. A card showing three personal action zones. A 60-second voice message with one concrete recommendation. A system that surfaces the single most important lever and fades everything else. Each one testing a different hypothesis about what “useful” actually means. And these weren’t demos we talked people through. They delivered real, personal results that users could share, save, or ignore. Same prototype, two completely different settings. Not asking “do you like this?” Watching Behavior. Observing whether it triggers the moment of understanding.

The Real Shift

The empathy phase was never supposed to be a workshop exercise. It was supposed to be deep, contextual, behavioral. Good DT practitioners did exactly that: shadowing, contextual inquiry, ethnographic observation. But even the best empathy research couldn’t answer whether people would actually use a solution in their real lives. That required a real product, and real products required real investment. So there was always a leap of faith between “we deeply understand the problem” and “this solution works in practice.”

That’s the gap AI collapses. Not the empathy. The testing.

The workshops and sticky notes were never the point. They were the constraint. Now that building is cheap enough to use as a research method, the exploration of human needs can happen where it always should have: in the lived reality of use, not in the abstraction of a Post-it wall.

Lean Startup Got There First

And yes—Lean Startup got there first. Not just philosophically, but methodologically.

Build-Measure-Learn. Validated Learning. The MVP. Eric Ries explicitly advocated for cheap validation: landing page tests, explainer videos, manual concierge services. Dropbox validated its entire concept with a three-minute video. Zappos tested whether people would buy shoes online by photographing local store inventory and fulfilling orders by hand. Ries wrote in 2009 that even a two-week build was “way too long” when a simple AdWords smoke test would reveal whether anyone cared.

The theory was sound. The problem was adoption.

In most corporate contexts, “MVP” drifted back toward Robinson’s original meaning: first version of the product, meant to sell. What Ries intended as a cheap experiment became a relabeled development phase, routinely costing €50,000+ and months of engineering. Steve Blank, Ries’s mentor, called this “innovation theater”: great press releases about how innovative the company is, no real change in how products get built.

The organizations didn’t corrupt the methodology. They chose the definition that felt safer.

AI doesn’t improve on Lean Startup’s philosophy. It removes the organizational barrier that kept it theoretical. What AI changes isn’t the MVP itself—it’s what comes before: the ability to run behavioral experiments in product form, without committing to a product.

Lean Startup told organizations to build less. Design Doing says: build throwaway prototypes to learn, then decide what’s worth building for real. When a non-technical person can build a behavioral prototype in days, the validation loop can finally be what Ries always meant it to be.

The Bottleneck Changes Shape

When testing is instant and cheap, what becomes the constraint?

It’s not ideas. Everyone has ideas. It’s not building. AI handles that. It’s not even money. Prototypes cost almost nothing.

The new and old bottleneck is attention.

Think about it. If every innovation team, every startup, every corporate lab can spin up testable prototypes in days, what happens to the people they’re trying to reach?

They drown in requests. Another survey. Another beta test. Another “we’d love your feedback.” Another app asking for 10 minutes of their time.

User attention becomes the scarcest resource in innovation.

This has real implications:

Ad spend stops scaling linearly. Today, €500–€1,000 per week gets you in front of 100 real users. But what happens when everyone is running prototype tests? The same eyeballs get more expensive. The same audiences get fatigued. The same channels get saturated.

Access to users becomes a competitive advantage. Companies with existing customer relationships, communities, or distribution channels can Deploy to Users faster and cheaper than those starting from zero. Having an audience isn’t just a marketing asset anymore. It’s innovation infrastructure.

The irony is clear. The more AI accelerates building, the more human relationships matter. Strategic empathy. Community. Trust. The ability to earn attention rather than buy it.

I think we’re entering a world where anyone can build anything. If that’s true, the winners won’t be those who build fastest. They’ll be those who’ve built the relationships that let them learn fastest.

A First Draft, Not a Finished Framework

I’m still learning where this breaks, where it holds, and where it needs to evolve. If you’re experimenting with similar approaches, running into limitations I haven’t considered, or think I’m getting something fundamentally wrong—I’d genuinely like to hear it.

Reach out, challenge the thinking, or share what’s working in your context. The best version of this won’t come from one person writing in isolation.

A note on how this was written: The ideas, opinions, and experiences in this article are mine. I used Claude as a writing partner to help structure my thinking, sharpen the arguments, and get from rough notes to a finished piece. The process felt a lot like what this article describes: AI handled the parts it’s good at, the human brought what it can’t.