If you asked a hundred industrial executives whether their AI projects are delivering value, somewhere between sixty and seventy of them would tell you no. They would not phrase it that bluntly. They would say the work is "in pilot," or "showing promise," or "in the next phase of deployment." Press a little further, and a quieter answer emerges: the model works in the lab, the prediction is accurate, the dashboard is operational, and yet the business has not changed. The maintenance schedule has not shifted. The reorder point has not moved. The quote a customer receives next week will not look meaningfully different from the one they received last year.

This is the structural failure of industrial AI, and it is so consistent that it deserves to be treated as a category problem rather than a string of unlucky implementations. Bain, BCG, MIT Sloan, and Boston Consulting have all published numbers in the same range over the past three years. Roughly seven in ten AI initiatives in industrial settings do not produce measurable business outcomes. The technology is functional. The integration into how work actually gets done is not.

What is going on here is not a problem of model accuracy. The models, in many cases, are excellent. The problem is that almost nobody on the implementation team is responsible for the operational change that has to follow the model. The data scientist's job ends when the prediction is correct. The IT team's job ends when the model is deployed to production. The procurement or maintenance or sales team that is supposed to act on the model differently has not been resourced, retrained, or restructured to do so. The model sits there, producing recommendations that nobody is required to use.

  • Roughly 70 percent of industrial AI projects fail to scale beyond pilot, per BCG's 2024 Industrial AI Survey.
  • Only 11 percent of manufacturers report capturing meaningful value from AI deployments at enterprise scale (MIT Sloan / BCG, 2024).
  • Companies that achieve scaled AI value invest 10x more in change management and process redesign than in the model itself (McKinsey Global Institute).
Figure 1
Where the value disappears
Industrial AI project funnel, from proof-of-concept to durable business outcome
Proof of concept (100%) 100 Model in production (62%) −38 Used by operators (34%) −45 Changed workflow (18%) −47 Measured P&L impact (11%) −39 Largest drop: between "model in production" and "used by operators." The technology arrives. The behavior change does not.
Source: BCG Industrial AI Survey (2024); MIT Sloan Management Review AI Survey (2024); composite of 320 manufacturing AI initiatives.

The funnel above tells the actual story. The first big drop, from proof-of-concept to production, is the one most executives focus on. It is the visible failure. The second drop, from production to operator usage, is larger and almost never discussed. It is the place where the company has spent the money but has not yet captured anything. The third drop, from usage to workflow change, is where most of the remaining value evaporates. By the time you reach measured P&L impact, you have lost roughly 90 percent of the projects that started.

Why the operator does not use the model

The reasons are not mysterious, but they require talking to operators rather than to executives. The maintenance technician who has been running a refinery for fifteen years has a mental model of the equipment that has been earned through hundreds of failure events. The AI model gives them a prediction. The prediction is occasionally wrong in a particular way, which costs the operator their reputation if they act on it and it is wrong. The prediction is also right in ways that the operator could have figured out themselves, which means acting on it gives them no credit. The asymmetry of incentives is severe: the operator owns the downside, the model owns nothing.

Now layer on the question of accountability. If the model recommends pulling a pump for early maintenance and the pump turns out to be fine, who explains the unplanned downtime? The operator, almost always. They are the human in the loop. The model cannot be reprimanded; the operator can be. Faced with this calculus, the rational behavior is to ignore the model when it disagrees with experience, and to use it only when it confirms what the operator already believed. This is not malice. This is the predictable behavior of a system that has placed a new tool into a workflow without restructuring the accountability around it.

The literature on this is consistent. In their study of 130 industrial AI deployments, MIT and BCG found that the single largest predictor of scaled value was not data quality, model performance, or executive sponsorship. It was the redesign of operator workflows and incentives to make using the model the path of least resistance. The companies that did this work captured value at six times the rate of those that did not. The companies that skipped it deployed the same models and got nothing.

Figure 2
Where the budget goes versus where the value comes from
Allocation of typical industrial AI program spend, contrasted with attribution of realized value
WHERE BUDGET GOES WHERE VALUE COMES FROM Data & infrastructure (38%) Model dev (31%) Deployment (20%) 6% Change mgmt 5% Workflow redesign 8%  Data & infrastructure 12%  Model accuracy 15%  Deployment quality 31%  Change management 34%  Workflow redesign 11% of spend 65% of attributable value
Source: McKinsey Global Institute "The State of AI" (2024); BCG Industrial AI value attribution analysis; based on 84 scaled deployments.

The mismatch in Figure 2 is the entire argument. Eleven percent of program spend goes to the categories that produce sixty-five percent of the value. The other eighty-nine percent goes to the categories that produce one-third. No financial discipline would tolerate this allocation if it were visible, but it is not visible, because change management and workflow redesign do not have line items in most AI program budgets. They are absorbed informally into operations or HR, where they are starved by competing priorities.

The model is the easy part. The business that uses the model is the hard part. Almost nobody budgets for the hard part.

The pattern of the projects that work

The industrial AI projects that produce durable value share a small number of structural features. None of them are technological.

The first is that they have a named business owner, not a named technology owner. The plant manager, the procurement lead, the maintenance director takes accountability for the outcome the AI is supposed to produce. The data science team is in service of that owner, not parallel to them. When the model produces a recommendation that the owner has to act on, the owner has been part of designing the recommendation and has authority to act on it. There is no hand-off across functional boundaries that strips the human-in-the-loop of agency.

The second is that the workflow has been redesigned before the model goes live, not after. The question of who does what differently the day the model is in production has been answered, documented, and rehearsed. The operators have been trained not on the model but on the new procedure. The procedure assumes the model. If the model fails, there is a fallback. If the model succeeds, there is a clear action.

The third is that the value is measured against a baseline that existed before the project. The most common failure mode in industrial AI evaluation is that the success metric is invented after the deployment. "We reduced unplanned downtime by 12 percent" is not a useful claim if the baseline was never measured, or if the comparison period had different operating conditions, or if the 12 percent reduction is within the noise of monthly variation. Companies that capture value rigorously also measure rigorously. They define the counterfactual before the work starts.

The fourth is that the project has a kill criterion. A real one. Most industrial AI initiatives have no defined failure condition, which means they cannot be killed, which means they accumulate. The portfolio of "pilots in progress" grows year over year, none ever cancelled, none ever scaled. Companies that produce value run their AI portfolios like a venture fund: explicit hypotheses, explicit failure conditions, explicit decisions to double down or shut down. The discipline of killing failures is the discipline that frees capital to scale successes.

What this means for the operator

If you run a business that has multiple AI initiatives in flight, three diagnostics will tell you which ones are going to deliver and which ones are going to quietly die.

First, ask each project owner what their operators do differently the day the model is in production. If the answer is "use the dashboard" or "see the prediction," the project is not going to produce value. The answer needs to be a specific change in action: this stage of the process is now automated, this approval is now conditional, this report now triggers a workflow. Vague answers indicate that the operational design has not been done.

Second, look at the budget. If less than fifteen percent of the program is allocated to change management and workflow redesign, the program is structurally underfunded for impact. The money has gone to the wrong activities, and the project is already on the path to the funnel above.

Third, look at how value is being measured. If the measurement plan is being written alongside the deployment, or after, it is too late. The baseline has been lost. The project will report success at the level of the model and silence at the level of the business, and the report will be roughly equally believable to nobody.

The point of these diagnostics is not to kill projects. It is to make visible the gap between what the technology can do and what the business has been resourced to absorb. Once that gap is visible, it becomes a planning question rather than an outcome.

The model is a recommendation. The business is what acts on recommendations. A company that has built one and not the other has not built an AI project. It has built an experiment, and experiments that nobody acts on do not have outcomes.