Flywheels Don't Spin in Sandboxes
Why AI pilots stall, and why the firms pulling ahead skip the experiment entirely

Flywheels Don't Spin in Sandboxes
Why AI pilots stall, and why the firms pulling ahead don't experiment: they deploy methodically in production
By John Miniati, Co-Founder & CEO, intoMO
Somewhere in your firm right now, there is probably an AI pilot. It has a small budget, an enthusiastic sponsor, and a demo that impressed everyone who saw it. And if this year's data is any guide, it is quietly going nowhere.
The 2026 numbers are in, and they are blunt. In Deloitte's State of AI in the Enterprise survey, published in January 2026 and spanning 3,235 senior leaders across 24 countries, only 25 percent of organizations had managed to move even 40 percent of their AI experiments into production. KPMG's Global AI Pulse, fielded in early 2026 across 2,110 executives in 20 countries, found just 11 percent of organizations mature enough to be scaling AI agents across the business. Meanwhile BCG's AI Radar 2026 reports that companies plan to roughly double AI spending this year, and that 94 percent of executives will sustain the investment even if returns do not arrive in 2026. Read those together and the picture is uncomfortable: conviction is at an all-time high, money is pouring in, and most of it is still funding experiments.
Here is the part I find most interesting. The failure usually gets blamed on the technology, or the data, or the vendor. In my experience it is almost always something more basic. The failure is built into the word "pilot" itself, or rather, into what that word has quietly come to mean.
We broke the word "pilot." Experiments are not pilots.
In almost every serious discipline, a pilot is real. A pilot plant in manufacturing is a functioning production facility, just smaller: real feedstock in, real product out, real operating data collected. A television pilot is a finished episode made to be aired. And the original pilot, the one on the ship or in the cockpit, is not simulating anything. They are steering the actual vessel.
Somewhere along the way, the enterprise AI pilot became the opposite: a non-production experiment. A sandbox environment. Sample data, or cleansed data. A handful of test users clicking around outside their actual workflow. No connection to the systems where work really happens, and no decision anyone would stake a number on. This is part of why they are everywhere: really good experiments, using real but curated data, are simple to build now on the frontier models (Claude, GPT, Gemini). The problem is that these experiments are not production pilots: they rarely have the AI governance infrastructure in place to launch in production.
Deloitte's researchers have a name for where this leads: the proof-of-concept trap. "Experiment pilots" run in isolated environments on curated data, and then production arrives with everything the experiment pilot was designed to exclude. The report notes that use cases estimated at three months routinely stretch to eighteen once real integration begins, and that companies respond by funding more pilots, which are cheap and low risk, rather than doing the harder work of scaling what already succeeded. One healthcare AI leader quoted in the study put it plainly: executing a hundred pilots just leads to poor results and failed value creation. There is even a term for the malaise now. Pilot fatigue.
This is why we draw a clear distinction between experiment pilots (also called sandbox pilots) and production pilots.
An experiment pilot is deliberately built so that nothing real depends on it. That feels safe. It is also why it fails: an experiment like that cannot succeed, no matter how good the technology is. To see why, you need one mental model.
AI systems are flywheels
A working AI system is not a tool you install. It is a flywheel: PRODUCTION USAGE CREATES DATA > DATA MAKES THE SYSTEM SMARTER > A SMARTER SYSTEM DRIVES PRODUCTIVITY GAINS > INCREASED PRODUCTIVITY EARNS INCREASED USAGE > AND THE COMPOUNDING FLYWHEEL ADVANTAGE SPINS FASTER WITH EACH CYCLE. Every AI product you consider best in class works this way, and the loop is the whole game. The system you deploy on day one is the worst it will ever be. What matters is whether the loop that improves it is turning.
Now look at the sandbox pilot through that lens. No real usage, so no real data. No real data, so no learning. No learning, so the system on day 90 is the same system you saw on day one, minus the novelty. The flywheel never turns because the pilot was designed, on purpose, to keep it from touching anything that would make it turn.
This is why sandbox pilot results are typically unpersuasive even when they are positive. An experiment can only ever produce an opinion: people liked it, the demo was impressive, we see potential. It cannot produce the thing executives actually need, which is evidence: adoption curves, cycle times, error rates, outcomes attributed to the system in the P&L. So the sandbox pilot ends, the opinions get debated, the budget cycle moves on, and the initiative joins the three quarters that never meaningfully reach production.
Meanwhile the cost of all this is not zero. Six months spent in a sandbox is six months a competitor's flywheel may have been spinning on real data. Flywheels compound, which means the gap between a firm whose loop is turning and a firm still running experiments does not grow linearly. It accelerates. This is the real meaning of first-mover advantage in AI: the prize is not being early to a technology, it is being early to a compounding loop. That is what KPMG's 11 percent understates: those firms are not merely ahead, they are pulling away at an increasing rate.
Leaders learn in-market with a production pilot
The firms getting real returns from AI are not running better experiments. Mostly, they are not running experiments at all. They identify a business opportunity, then build rapidly and directly on a production platform, with a deliberately controlled rollout. This is a production pilot, in the older, truer sense of the word pilot: the real system, operating for real, at reduced scale.
In practice, the production-first path looks like this:
A 2-week Discovery & Design Sprint first. Best practices in Jobs-to-be-Done, Outcome-Driven Innovation, and Design Sprints define a high-value AI business opportunity. Not a 6-month AI strategy. Not a portfolio of random experiments.
The real system, from day one. Built on production-grade foundations, connected to real data, doing real work inside the actual workflow. Not a mockup that will someday be rebuilt "properly," because that rebuild is exactly where three-month estimates become eighteen-month projects.
Agile in the age of AI. We've each spent five years building production-grade AI systems, the last three as one team, refining methods and infrastructure so it is faster to build good software the first time. Days, not weeks. Weeks, not months. Every 2-week sprint opens with software you can see and closes with software your users are using.
A controlled group of real users. Not a test panel. A small set of people doing their actual jobs with the system, whose usage generates the first real data the flywheel needs.
Human judgment at the gates. Experienced people review and approve the system's output at the decision points that matter, with a full record of what the AI produced and what a human signed off. This is what makes going straight to production responsible rather than reckless: risk is managed by architecture, not avoided by staying in a sandbox.
Mistakes on purpose, at small scale. Errors will happen. In a production pilot they surface early, with a small group, get caught at the gates, and feed the next iteration. Every mistake makes the system smarter. In a sandbox, mistakes teach you nothing, because the conditions that produced them were never real.
The contrast is worth stating plainly. A sandbox pilot generates opinions. A production pilot generates a head start.
Why the compounding belongs to whoever owns the loop
There is a second-order point here that I think is underappreciated. Once a flywheel is spinning on your firm's usage, your data, and your experts' judgment, the advantage it builds accrues to you specifically. It cannot be bought by a competitor, because it is made of things a competitor does not have: your accumulated interaction data, your encoded expertise, your refined workflows. Generic tools get better for everyone at once, including the firm across the street. A flywheel you own gets better for you alone.
And the economics run in the owner's favor over time. Underlying model improvements flow into the system. The marginal cost of the next user approaches zero. The falling cost of AI lands on your ledger, not a vendor's.
The honest objections
Two pushbacks deserve a straight answer. First: "going straight to production sounds risky." The 2026 data says the opposite. In KPMG's survey, only 20 percent of organizations still in the experimenting phase felt confident managing AI risk, versus 49 percent of the firms operating AI at scale. Sitting in the sandbox does not build risk muscle; operating real systems behind human gates does. Done right, the system also sits beside your infrastructure rather than inside it: reading what it needs, writing nothing without human approval, so that switching it off breaks nothing. What is actually risky is spending two quarters learning nothing while calling it prudence.
Second: "we are not ready." Readiness is mostly a question of aim. The most expensive mistake in AI is not building badly, it is building the wrong thing well. That is solved by a short, disciplined design phase: pick one business area that matters, find the problems worth solving, size them by value, and prototype before committing. Weeks of aiming, not months of strategy.
The firms I see winning treat this whole arc, from choosing the opportunity to a production pilot with real users, as roughly a one-quarter exercise. Not a year. A quarter to get the flywheel spinning; every quarter after that, it compounds.
The question worth asking about your current AI initiative is simple: is it an experiment, or is it a flywheel? If nothing real depends on it, you already know the answer, and 2026 has already published what happens next.
John Miniati is the Co-Founder and CEO of intoMO, a boutique AI systems firm that designs and builds Superminds: production-grade AI systems for judgment-driven businesses. If you are weighing where AI genuinely fits in your firm, he is happy to compare notes:



