Every week someone asks us the same question: "We know we should do something with AI. Where do we start?" Almost everyone expects the answer to be a tool, but it never is.
Flash summary
Don't start with technology. Start with a recurring problem or decision your business makes badly, often and expensively. Score your options on Profit × Planet × Feasibility. Give the winner an owner who knows the problem well, dig into it with your own team, build a small working version on your real data, and let the people who do the job every day tell you whether it helps. Thirty to ninety days from problem to answer. Then scale, adapt or drop.
Why most AI initiatives quietly die
The pattern is always the same, and it is rarely anyone's fault. Pressure arrives, licences get bought or a chatbot gets commissioned, a pilot runs, and six months later nobody can point to a number that moved the needle. A widely quoted MIT study from 2025 claimed roughly 95% of corporate GenAI pilots showed no measurable P&L impact. And nothing is wrong with GenAI technology which keeps advancing at an unprecedented pace. What was missing was a problem worth solving and a number to check it against. Nobody decided, at the start, which business decision the AI was supposed to improve.
AI doesn't fail at the technology stage. It fails at the question stage.
And this is not a small-company problem. McKinsey surveyed more than 10,000 senior managers across 15 countries. 88% of their organisations deploy AI somewhere. Roughly the same number report no significant bottom line impact, and fewer than one in five of those attempting an AI-first operating model have seen real EBIT impact. The conclusion is not to spend less on AI. It is to be clearer about what each piece of it is for.
First, a reminder: "AI" is not one single thing
AI is more than a chatbot. We explain the full picture in Beyond GenAI: AI's big picture, worth six minutes.
In short: several kinds of machine learning sit under the umbrella and do very different jobs. Supervised learning predicts a known outcome: which machine will fail, which invoice will be late. Unsupervised learning finds patterns nobody labelled, behind customer segmentation and anomaly detection. Reinforcement learning optimises a series of choices, like when to reorder or how to price. Generative AI, the part that got all the press, writes text and code and makes images. Agents are models that can act.
So the problem should pick the technique, not the other way round. A quality-control camera and a customer-service assistant are both "AI", and building them has almost nothing in common. Generative AI is what everyone is already buying, and rightly so. The quieter techniques, forecasting, clustering, computer vision, optimisation, are where operational money is still sitting untouched.
The order is the strategy
How it usually goes and how to sequence your AI discovery and innovation for true adoption:
- Buy the tool, or order the chatbot
- Hunt for a use case that fits it
- Ask who owns this. Get silence
- Discover the data isn't there
- Run a demo, call it a pilot, quietly stop mentioning it
- Pick the problem: scored on Profit × Planet × Feasibility
- Name the owner: one person, reachable inside a quarter
- Deep dive: the internal problem experts in the room, validate feasibility
- Prototype on real data: build fast, real users test it
- Scale, adapt or drop: data-driven verdict
The tool isn't the villain. Its position in the queue is. Move it from first to last and everything downstream gets easier, including explaining it to the board.
Step 1
Choose the problem first
This is the step everyone skips, and skipping it is the best predictor that the project will die. So be concrete. List the decisions your operation makes badly, often, and expensively. Which orders to rush. How much material to order. Which SKUs to stock. What price to quote. When to service a machine. Which appointments to take. How to schedule routes. Each of these is made hundreds of times a year on experience and gut feel, and each costs money when it goes wrong.
Then score the list, a rough estimate, before anyone builds anything. Most companies use one axis. We use three:
worth
Profit
It happens often, and costs real money when it's wrong.
burns
Planet
It wastes material, energy or emissions.
do it
Feasibility
Can this be done here, and at what cost?
Plot them on all three and the loudest idea is rarely the winner. Someone's favourite chatbot drops; an unglamorous quality-control camera rises.
Feasibility is where the data question lives: does data about this decision already exist somewhere digital, even a spreadsheet someone guards with their life? And would anyone act differently if a prediction landed on their desk? A model nobody trusts changes nothing, and the second question kills more projects than the first.
That is what an AI Value Deep Dive from AmPhi Labs, an AI innovation studio in Amsterdam, is for: we interview your people, explore your real data, and hand back a prioritised AI Opportunity Matrix with a readiness score and a clear "build and test this first". Sometimes the honest answer is that your data isn't ready, and it is better to hear that in week two than month nine.
And "feasible" is rarely a technology question. Managers name worries about AI itself, legal risk and change management ahead of everything else; infrastructure comes fourth. For every €1 on the technology, expect €5 on the people.
Step 2
Name an owner
With a ranked shortlist in hand, nominate someone for the problem at the top.
The problem owner is whoever owns the decision you picked in Step 1. The quote desk lead. The maintenance planner. The scheduler who has done it for eleven years. This does not have to be a manager, and often shouldn't be. What matters is that they own the decision, feel it when it goes wrong, and can tell you when the model is talking nonsense.
The sponsor sits high enough to unblock. Data access, IT time and permission are what actually kill pilots, and none of the three are in an operator's area of influence.
A 2026 survey of 500 organisations priced the gap: naming one accountable person is worth close to a full point of AI maturity. That costs no budget and no platform. It takes a decision.
Step 3
Deep dive, with the team in the room
Now, and only now, on the problem that won, go deep. Sit with the people who make the decision. What do they look at? What do they ignore? What's the rule of thumb they'd never write down? How often are they wrong, and how do they find out?
This part is unglamorous and it is where the value hides. One large company found it was making 35% of its decisions twice, in two different departments, sometimes with different answers. No model fixes that. Finding it was worth more than any model.
Who is involved here matters. Ideas your own people shape get adopted, while plans handed down from outside get resistance. That's why our deep dives put your own experts in the room: a 4 hour AI-Powered Idea Lab with an Innovation Catalyst, or a Hack Pack™ shipped to your office. Both end the same way: a decision made in the room and a build-ready specification.
Step 4
Build a small working version
Now build the smallest thing that can move the number and test its effectiveness. A working prototype or MVP (Minimum Viable Product) running on your real data, sitting next to your systems rather than inside them, so testing risks nothing live.
Agree the success measure before you start: increase leads by 2 points, bring quote time down 40%, forecast error halved. Build the measurement in from day one. Then hand it to real users: the people who will live with it. Theirs is the only reaction that tells you anything.
That is the point of rapid prototyping. You are not testing whether the model is clever. You are testing problem-fit: whether it solves the problem you picked, and whether anyone changes what they do because of it. Agile, in short sprints, with the scope small enough that being wrong stays cheap.
Step 5
Thirty to ninety days, then scale, adapt or drop
How long should this take? Thirty days at the fast end, ninety at the slow end. What sets the number is how often the decision happens, how long before you know it was right, and how small a change you are trying to prove. Size that window before the build starts, so you neither scale on noise nor burn the quarter waiting.
Two clocks run at once. Building is the fast one: days to weeks, and if it takes much longer the scope is too big. Proof is the slow one. A quote desk pricing 200 times a day answers within a month; a monthly planning cycle gives you three data points in ninety days. Small improvements hide inside a normal good week until you have enough weeks to compare.
Then take the answer honestly: scale, adapt or drop. All three are good outcomes. If it is scale, that is the moment for production deployment and real budget, earned with evidence instead of a business case.
Scaling up
A verdict of scale is not the same as "it worked". It worked at pilot size. Before you commit, price the gap between the two. Four questions do most of the work.
Cost at real volume
200 documents is a different product from 200,000 a month. With generative AI you pay per use, so the bill grows with the volume.
Integration
Production means going inside your systems, and replacing the Friday spreadsheet export with a real data feed. Most of the remaining budget goes here.
Infrastructure
Who refreshes the data, and how often? Does where it runs satisfy GDPR? Who gets called when it breaks at 03:00?
Ownership on day 400
Models drift, because the world moves and the model doesn't. If nobody is named to watch the numbers, you are scaling a future problem.
Which is why the prototype should be built with the scale-up already in view. A throwaway interface is fine; throwaway plumbing is expensive. Clean data connections, the success measure built in from day one, nothing hard-wired to a single vendor. It barely slows a two-week build, and it decides whether scaling later takes six weeks or six months.
None of this is really about technology. It is about the order.
Five common wrong turns
Most failed starts are one of these. Don't take them.
- Platform first. Tools chosen before problems become shelfware with a subscription. Know your top three use cases before you buy anything.
- The hardest problem. Start where the data already exists and someone cares about the outcome.
- The GenAI reflex. GenAI can superpower your workforce, but without use cases and training nobody sees it. Match it to your problem list.
- Twenty deep dives. Score them all, go deep on three at most. Keep focus, then expand.
- Outsourced thinking. Ideas your own people shape get adopted. Plans imposed from outside get resistance.
Find your starting point in days
The AI Value Deep Dive maps and scores your highest-value AI opportunities on profit and planet, before you commit a euro to building. It is the lowest-risk way to start: you find out where your data really stands, and where the return actually is. Take a look, or book twenty minutes and we'll help you find it.
Dynamic pricing: the same maths can serve profit and planet →
Sources: McKinsey, "The State of Organizations 2026" (10,018 executives, 15 countries, June–September 2025) for AI ownership, readiness, adoption barriers and the €1:€5 ratio; "State of AI trust in 2026" (~500 organisations, March 2026) for accountability and agent barriers. The ~95% figure comes from "The GenAI Divide", MIT Media Lab Project NANDA (July 2025), a non-peer-reviewed preprint, contested on methodology. Figures current as of August 2026.
