Enterprises have spent a fortune on AI and productivity has not measurably increased. Two things are in the way: a process nobody redesigned, and business definitions nobody ever wrote down. Here is what each one looks like, and what to do about it.
“the end result is just making garbage faster”Vas, CEO of Varick Agents, on what enterprises get when they apply AI to the processes they already run · Sep 2026
Two walls, one behind the other. This article is about both, in the order you will meet them.
Most of these programmes can show you a successful pilot. The tools were deployed, people used them, and the closing deck reported hours saved per user per week. Far fewer can show what those hours did to cycle time, headcount or margin, because the saving never arrived anywhere it could be counted.
Before LazyFox, our team spent more than fifteen years advising enterprise companies on digital transformation and running the change programmes inside them. That work ran across most industries: Audi, Bertelsmann, BNP Paribas, Coca-Cola, Daimler, Deutsche Bank, Deutsche Post, ERGO, Gruner+Jahr, Sanofi, SWISS. We built that business and sold it. Since then we have spent five years building AI systems for regulated and sensitive environments, financial services above all, where a wrong number carries regulatory consequences.
Fifteen years of transformation work taught us a lesson that keeps repeating: the technology is almost never why a transformation fails. It fails because the work around the technology stays exactly as it was. That was true of ERP rollouts, it was true of cloud migrations, it was true of every digital transformation programme we sat in, and it is true again now with considerably more money attached.
Vas at Varick Agents made this case sharply this month from the AI side, drawing on conversations with more than 300 CEOs, CIOs and CFOs at very large companies. His summary is hard to improve on: companies apply AI on top of the processes they already run, and “the end result is just making garbage faster.” He is right. The first half of this article covers that same ground, with the published numbers behind it.
There is a second wall behind it, though, and it is the one we run into every day at LazyFox. Redesign the process properly and you will get a real result. Then your agents start asking your systems for numbers, and your systems disagree about what those numbers mean, because nobody ever wrote the definitions down. A broken process announces itself: everything is slow and everyone complains about it. A missing definition does not. It produces confident answers that are quietly wrong, which is a considerably more expensive way to fail.
So this covers both walls in the order you will hit them. First, why speeding up the work changes nothing, and what to do to the process instead. Then the layer underneath it: what to define before you automate anything, and what it looked like when one of our customers went through both. If you are running or funding an AI programme right now, the aim is that you finish this able to name which of the two walls you are currently standing in front of, and what the next move against it is.
A thousand Copilot licences, three months, and no productivity anyone could measure.
In late 2024 the UK Department for Business and Trade ran one of the more honest experiments in enterprise AI. It gave 1,000 staff Microsoft 365 Copilot for three months, measured what happened, and then published the result rather than a press release.
People used it 1.14 times per working day. When researchers watched them perform real tasks, slide decks came out roughly seven minutes faster and scored worse on quality and accuracy. Data analysis in Excel took longer with the tool than without it, and was less accurate. Email drafting saved a little time. And 72% of users said they were satisfied or very satisfied, with a net promoter score of 31.
The evaluation's own conclusion was that it found no robust evidence the time savings were translating into improved productivity.
That pattern repeats across most AI programmes running right now: high satisfaction, genuinely faster individual tasks, and nothing at the other end. When the tool works and the users are happy and the productivity never shows up, the problem is not the tool and it is not the user. It is everything sitting between them and the outcome.
Vas, who runs the AI transformation firm Varick, put the mechanism bluntly this month: most companies apply AI on top of the processes they already have, and “the end result is just making garbage faster.” He is right, and the observation is a great deal older than the current wave.
In 1990 Michael Hammer published an article in Harvard Business Review with the instruction in its title: don't automate, obliterate. His complaint was that heavy investment in information technology kept returning very little, because companies left their existing processes intact and used computers to run them faster. His phrase for it has aged unreasonably well: “It is time to stop paving the cow paths.”
Hammer also had the number that explains everything else. He cites an insurer that estimated an application spent 22 days in process and was actually worked on for 17 minutes.
Read that again with an AI budget in mind. Deploy the best model in the world against that process and make every hands-on step twice as fast. You have saved eight and a half minutes out of 22 days. You will report it as a 50% reduction in processing time, you will be telling the truth, and the customer will not notice anything at all.
The 22 days were never the work. They were the queues: the handoff to the next team, the wait for a review, the second approval, the message to someone in another timezone who answers on Thursday. AI does not touch any of that by default, because nobody bought it to.
ERP and cloud punished laggards slowly. This one compounds against you while you deliberate.
There is a version of this argument that says none of it is urgent. You are large, you are profitable, the last three technology waves did not kill you, and this one can wait for a budget cycle. That reasoning held for ERP and it held for cloud. It does not hold here, and the difference is worth being precise about.
Cloud and ERP mostly changed your own cost base. AI changes what a competitor can do to you. A company that gets this right can undercut you on price, because the same work costs them a fraction of what it costs you. It can take on more volume without adding the headcount that processes it. It can reach every buyer in your market before your team has finished assembling the list. None of that was available to whoever adopted cloud faster than you did.
And the alternative to moving is not standing still. It is growing the expensive way. Demand rises, so three analysts become six. Six people need a manager. Two teams need a handoff that did not exist when it was one team. The handoff needs a status meeting, and the status meeting needs somebody to run it. That is how headcount charts bend upward while revenue per employee stays flat, and it is the good outcome, not the bad one.
Meanwhile the standard enterprise response is to launch a programme. The Financial Times reported that Bayer is midway through a six-year overhaul of an SAP-based system, with 30 AI agents now supporting coding and testing, and the goal of ending up with “significantly fewer consultants in the programme doing a deployment.”
Look at the timeline rather than the AI part. A six-year programme starting today finishes in 2032. Six years ago ChatGPT did not exist. Nobody in 2020 could have written a requirements document that would still be right today, and nobody writing one today can tell you what will still be right in 2032. Any plan that only pays off in six years is a bet on a technology forecast nobody is currently able to make.
Which is the strongest argument for doing the unglamorous work first. Process design and business definitions are the two things that do not expire when the models change. A redesigned approval flow is still the right flow whichever model runs it. A definition of revenue that finance, sales and delivery have agreed on is still correct whether it is read by a model that exists today or one that ships in three years. Nearly everything else in your AI stack is rented, and the rental terms change every few months.
Start with where the time goes, not with where the tools fit. And do not begin by consolidating your systems.
The first move is not a tool selection. It is finding out where your own 22 days go. Two things make that harder than it sounds.
The first is authority. Whoever signs the AI contract owns tooling and vendor selection. They do not own the process, and they certainly cannot walk into finance and announce that a fourteen-step approval chain is now five steps. So the process survives untouched and the AI gets applied on top of it, which is exactly the failure above. Redesigning anything that spans four departments requires someone senior enough to tell four departments that it is changing, and there are only a handful of those people in any company. If none of them is sponsoring the work, you will get a faster version of what you already have.
The second is the instinct to fix the systems first. One ERP, one CRM, one data platform, and then AI on top. This is the most expensive available way to postpone the actual work, and the odds on it are worse than most boards are shown. McKinsey, working with the University of Oxford, studied more than 5,400 IT projects and found that large ones run 45% over budget and 7% over schedule while delivering 56% less value than promised. Seventeen percent of them go so badly that they threaten the existence of the company running them.
It is also usually unnecessary. Deterministic integrations used to break the moment two systems disagreed about what a customer record was, which is what made consolidation look mandatory in the first place. Agents tolerate that kind of disagreement. An orchestration layer across the systems you already run does the job a multi-year migration was supposed to do, without the multi-year migration.
What does need standardising is the process itself. If you have acquired nine companies over twenty years, you have one process running nine different ways, and applying AI to that means building and then maintaining nine sets of agents in perpetuity. Define the process once, roll it out, then automate it.
Then map how the work actually happens, department by department, rather than how it is documented. For each workflow, establish the happy path, which is usually the only part anybody wrote down. Then the exceptions, which is where your time really goes: what share of volume leaves the happy path, where it goes, who gets pulled in, and what an error costs. Then what sits upstream and downstream, which system is the system of record and which one wins when two disagree, how the whole thing differs by region and entity, and the gap between touch time and elapsed time at every single step.
There are two ways to get that, and you need both. Process mining against the systems of record gives you timestamps, throughput and the reality that does not match the documentation. Interviews give you the part that lives only in people's heads. Sending an engineer to run those interviews does not work, and neither does sending an AI, because half the job is earning enough trust that an operator tells you about the workaround they have quietly been running for six years.
Three buckets, and the discipline is to use the cheapest one that actually works.
Once you know how the work really runs, redesign becomes a sorting exercise. Every step belongs in one of three buckets.
Deterministic. If X then Y, no judgment involved. If the invoice is under the threshold and matches the purchase order and the goods receipt, pay it. These belong in plain software. Regular code is cheap, fast, auditable and incapable of hallucinating. Do not spend a model call on something an if-statement handles.
Agentic. Steps where a human exercises judgment, you hold thousands of recorded examples of that judgment, and being occasionally wrong is survivable. Coding a transaction to the right account, routing a deal-desk approval, matching or not matching, flagging or not flagging. Ship those as agents and measure every output. The test is the difficulty of the action, not the difficulty of the thinking.
Human in the loop. Steps too risky or too novel to hand over. Here the agent's job is not to decide, it is to prepare: assemble every piece of evidence the decision needs, put it in front of the person, and propose a next step. The human still decides, but in thirty seconds rather than after forty minutes of hunting through email.
Done properly, a twenty-five step process collapses into a handful of agents with deterministic steps around them and one or two human checkpoints. That is where the return comes from. Not from the model.
Two things to settle before anyone builds anything. Baseline the numbers first: what this process costs today in cycle time, headcount and error rate, and where those figures are measured. If you cannot state that now, you will not be able to prove anything in twelve months and the programme will be judged on anecdote. And sequence by handoffs rather than by volume. The workflow with the most handoffs is where the queues are, and the queues are where the time is. It is rarely the workflow with the most transactions.
A redesigned process still runs on numbers, and your systems do not agree about what those numbers mean.
Do all of that and you will get a real result. You will also, at some point, hit the next wall. This is the one we spend our days on.
A redesigned process still runs on numbers. The agent deciding whether to pay an invoice needs to know what counts as a match. The agent flagging a deal needs to know what your company means by pipeline. The human at the checkpoint needs the evidence in front of them to be correct. Every one of those steps rests on a definition, and in most enterprises those definitions do not exist in any form a machine can read.
There is a clean experiment that shows the size of the gap. Spider is the standard benchmark for turning a question in plain English into a database query. On Spider 1.0, whose databases are small and tidy, o1-preview solves 91.2% of the tasks. On BIRD, which is harder, the same model solves 73.0%. Then researchers built Spider 2.0 from real enterprise workloads: databases taken from actual companies, often carrying more than a thousand columns, with answers that regularly run past a hundred lines of SQL. The same model solves 21.3%.
Nothing about the model changed across those three numbers. What changed is that the schema stopped explaining itself. In a textbook database the column is called cancelled. In yours it is called cncld, it holds a 6, and the 6 meant something specific to a contractor in 2019. A model can read your schema. It cannot read the meeting where finance decided that a cancelled order still counts toward the quarter if the cancellation came after the invoice date.
This is the second half of the reason those pilots return nothing. The demo runs on the part of your data that explains itself. Production runs on everything else.
And this failure is quiet, which makes it worse than the first one. A badly designed process announces itself: things are slow, everyone complains. An undefined metric does not. Your pilot did not fall over. It produced numbers. Some of them were wrong, and the wrong ones looked exactly like the right ones.
The semantic layer is exactly as old as the reengineering argument. Most companies skipped it, because until now they could afford to.
Hammer was not the only person to publish an answer in 1990.
That same year, a company founded in Paris called Business Objects built its product around an idea it named the universe: a layer sitting between people and the database that held what the data meant, so that someone asking for revenue received the company's definition of revenue rather than a column called rev. The industry ended up calling this the semantic layer, and in one form or another it has existed ever since.
Two answers, published in the same year, to two halves of one problem. Redesign the process, and define the meaning underneath it. Most enterprises did neither, and for 36 years the second omission was survivable, because a human sat between the data and the decision. When an analyst pulled a revenue figure, they knew to drop the intercompany entries, because they had been doing the job for nine years. The definition was missing from the system and present in the room.
Take the human out of that loop and the missing definition stops being an inconvenience and becomes the output. An agent does not know to drop the intercompany entries. Nobody told it, because nobody ever wrote it down, because until now nobody had to.
That is the real reason this is surfacing in 2026. Your data did not get worse. The last reader who could quietly compensate for it left the loop.
Gartner now expects universal semantic layers to be treated as critical infrastructure by 2030, on a level with data platforms and cybersecurity. That is a fair read of where this ends up. It is also 40 years after the idea first shipped.
Seven things to settle per metric, the two ways to find them, and where each one belongs afterwards.
Same shape as the process work. Start where people already argue, and settle seven things for each metric in dispute.
You find these the same two ways you found the process. Mine the queries, because whatever the wiki claims, the query is what actually ran. Then talk to the people, because the rest lives in ten or twenty heads, and an operator will not volunteer a rule they stopped noticing they apply years ago. Skip either half and you get a semantic layer that is confidently wrong, which is worse than none at all.
Sort what you find into three layers, and put each item in the lowest one that holds it. Structural is the shape of the data: tables, columns, relationships, derivable by reading the schema once at indexing time. Logical is the business rules above, which have to be authored, reviewed by whoever owns the metric, and versioned like code. Contextual is what never reaches a schema: the contract exception, the reason this one customer is invoiced differently.
One rule governs all of it. Resolve meaning once, at indexing time, into code the query engine executes deterministically, rather than per query inside a model. That is what makes an answer reproducible instead of probable, and why token cost stops scaling with the number of questions people ask. Then baseline it the way you baselined the process: ask every system holding your five most disputed metrics the same question on the same day, and write down how many different answers come back.
A fintech where adding one report cost weeks, and writing the query was never the reason.
Both walls show up together. One of our customers is a spend management platform whose own customers are finance teams, so reporting is not a feature of the product, it is what people open it for. Adding one report meant backend work, frontend work, testing and a release: a couple of weeks when it went smoothly, and every release risked breaking something unrelated. That is the process wall, in the exact shape of Hammer's insurance file. A faster way to write the query would have saved them almost nothing.
We connected read-only to their MongoDB and generated the structural layer automatically, because their tables carried no descriptions. They named the first five reports they wanted and sent a short brief for each. The output had large gaps against what they expected, and the gaps were the point: the brief described what they wanted, not the rules underneath it. So they sent the queries behind their existing reports, we reverse-engineered the logic out of them, and then sat with their team for a week. Mining, then interviews.
That week surfaced a rule nobody had ever written down: how a performance period should be assigned when an amount spans several months. It had never been specified, and their own system had been applying it differently from one report to the next. The same week surfaced two legacy reports, built three years apart by different engineers, returning different numbers for the same figure, with a condition hardcoded into one of them and no explanation attached. Neither engineer had been careless. Nobody had ever decided which one was right.
Once the logic existed as definitions rather than as code inside individual queries, four things changed.
The lesson their team took away is not a lesson about AI. The problem was never the AI and it was never the queries. The business logic underneath had never been made explicit, and until it was, nothing built on top of it could be trusted.
Nobody owns the process, nobody owns meaning, the last attempt at fixing both rotted, and the current pitch says none of it is needed.
Three reasons, and none of them is that the people involved are careless.
Nobody owns it. A company has an owner for the CRM, an owner for the warehouse and an owner for the AI budget. It has no owner for the end-to-end process, and no owner for what revenue means, because both answers span finance, sales and delivery, and none of the three can be told they have been holding it wrong. Our version of this is smaller than the process version and just as decisive. The customer above went in expecting to point us at a database and walk away, and came out knowing that somebody on their side has to own the definition layer permanently. Not a full-time job. A named person. Every deployment that works has one. Every one that stalls does not.
The last attempt rotted, and people remember. Most enterprises already have a data catalog or a data dictionary from a governance programme a few years back. It is out of date, everyone knows it, and nobody is enthusiastic about funding a second one. The same thing happened to reengineering in the 1990s, and for the same reason: the artefact was separate from the work. A document describing what a metric means starts dying the moment somebody edits a query, because nothing forces the two to agree and nothing breaks when they diverge.
What is different now is the only honest reason to expect a different outcome. Definitions that are executed do not rot. If the layer holding the meaning is the same layer the query engine runs against, the documentation is not a description of the system, it is the system, and drift becomes something you detect rather than something you discover in a board meeting.
And the pitch you are getting says none of this is necessary. Buy the licences, connect the model to your database, and it will work your business out. It will not. It will produce something plausible, which is the worst available failure mode for a number, because a wrong number that looks wrong gets caught and a wrong number that looks right gets presented. There is a related trap in where the definitions end up living. If the only place your business logic exists is inside a model provider's context window or memory service, your institutional knowledge is a rental. Definitions belong in versioned code you own, so that changing model is a procurement decision rather than an amnesia event.
Nine lines for anyone who scrolled.
Enterprise AI programmes keep making individual tasks faster without improving organisational productivity. A UK government trial of Microsoft 365 Copilot across 1,000 staff found 1.14 uses per person per day, worse quality output on observed tasks, and no robust evidence that time savings reached productivity. Michael Hammer diagnosed the same pattern in 1990: the time sits in queues and handoffs, not in the work itself.
The post covers how to redesign a process before automating it, why consolidating systems first is unnecessary, how to sort every step into deterministic software, an agent, or a human in the loop, and then the second wall behind the process: the definitions a redesigned workflow depends on were never written down, which is why the same model scores 91.2% on textbook database benchmarks and 21.3% on real enterprise ones.
LazyFox holds your business definitions as versioned code your company owns, and serves them to whichever model or agent is asking. It connects read-only. Nothing migrates.