Blog

Why Most AI Projects Fail Before They Start

The numbers on this are worse than most people assume, and they come from real research, not just anecdote. RAND Corporation’s 2024 study of AI projects found that more than 80% fail to deliver business value, roughly twice the failure rate of non-AI IT projects. MIT’s 2025 “State of AI in Business” report went further: 95% of generative AI pilots showed zero measurable return. Different institutions, different methodologies, same story: a business buys a tool, points it at a task, and six months later can’t say what it saved them, if anything.

That’s not a model problem. Claude, GPT, and the rest are good enough for the vast majority of repeated business tasks today. The failures are almost always upstream of the model, in the decision about what to build in the first place.

The pattern behind the failures

It usually goes one of three ways.

The tool automates the wrong task. Someone sees a demo of an AI agent doing something impressive and buys it, without checking whether that specific task is actually expensive for their business. It might cost the team twenty minutes a week. The build cost more than it will ever save.

The tool automates a task that needed a human judgment call. This is the quieter failure mode. On the surface, a task looks mechanical, drafting a client update, say. Three steps in, there’s a decision that depends on context nobody wrote down: which clients get the soft version of bad news, which invoices need a second look before they go out. Automate past that step and the system makes wrong calls confidently, which is worse than making no calls at all.

Nobody measured the baseline. Without knowing how long a task actually took before automation, there’s no way to know afterward whether it worked. Six months in, the honest answer to “did this pay for itself” is “we’re not sure,” which is functionally the same as no.

What an audit actually fixes

An AI Audit exists to catch all three before any money goes into a build. It’s a structured review of the key decision makers on a team, an hour each, walking through what they actually do in a working week rather than what their job title suggests they do.

For each task that comes up, the questions are the same every time: what triggers it, how often, how long it takes, and, most importantly, how much of it is mechanical versus a real decision. That last question is the one most businesses skip, and it’s the one that predicts whether an automation will hold up in the real world or quietly make wrong calls nobody notices for a month.

The output isn’t a vague strategy deck. It’s a ranked list, each item scored on time cost, how automatable it actually is, and how ready the underlying data is. The items that score high are the ones worth building. The items that score low usually mean the process itself is broken, and automating a broken process just moves the mess faster.

The part that’s easy to miss

A properly run audit almost always turns up at least one “quick win”: something small enough to fix in the first week or two, cheap enough that it pays for the audit on its own, before the bigger recommendation is even scoped. That’s not a coincidence. Once you’ve actually watched someone do their job for an hour, the obvious fixes tend to surface fast. The businesses that skip the audit and go straight to a build almost never find these, because nobody was looking closely enough to see them.

None of this is exotic. It’s closer to how a good operations review has always worked, applied to the specific question of what a machine can safely take off someone’s plate. The businesses that get real value out of AI aren’t the ones that moved fastest. They’re the ones that spent a week and a half finding the actual expensive task before they touched a single tool.

If you want to see what that looks like against your own team, the AI Audit runs ten working days: interviews with the key decision makers on your team, a scored list, and an honest answer on whether there’s anything worth building yet.