AI Agents

Why AI Agent Projects Fail: 7 Causes, With Receipts

Gartner expects over 40% of agentic AI projects to be canceled by end-2027. Seven documented causes, each pinned to a named incident — and the missing control.

By Isaac, Founder, Visione Edge14 min read
Unfinished concert hall at night: scaffolding rising into navy darkness, rows of seats wrapped in pale sheeting, a single blue work lamp on the empty stage.

Why do most AI agent projects fail?

AI agent projects fail for reasons that are mostly boring and mostly preventable: no defined financial return, error rates that compound across steps, permissions no intern would get, zero adversarial testing, promises nobody owns, demos standing in for tests, and budgets that stop at the build. Gartner predicts over 40% of agentic AI projects will be canceled by the end of 2027.

That Gartner prediction is probably the most-quoted statistic in this space, and most rewrites get it wrong. Here is the sentence as Gartner published it on June 25, 2025:

Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls, according to Gartner, Inc.

It is a prediction, not a measurement. It covers agentic AI projects broadly, not "all AI." And its three named reasons are the skeleton of this article.

Agent projects rarely fail at the model. They fail at everything wrapped around it.

One editorial rule governs this page: every cause below is pinned to a named public incident or a named survey figure. Anything we could not pin, we cut. Much of what ranks for this question is paywalled or written only for engineers; this page is neither, and every claim links to a source you can open.

Cause 1: Was there a defined return, or just a mandate to "do AI"?

Gartner's cancellation prediction names its reasons: escalating costs, unclear business value, inadequate risk controls. Two of those three are business-case failures, not technology failures. Most canceled projects never wrote down what the agent was supposed to return, in dollars, by when. When the invoices arrive and nobody can point to a payback date, cancellation is the rational move.

The same release quotes Gartner analyst Anushree Verma: "Most agentic AI projects right now are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied." The hype layer is thicker than the deployment layer — Gartner estimates "only about 130 of the thousands of agentic AI vendors are real."

The most famous number here needs more care, not less. MIT Project NANDA's July 2025 report found that "95% of organizations are getting zero return" on generative AI, despite $30–40 billion in enterprise investment: "Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact." For task-specific tools, it reports a brutal pipeline: 60% of organizations evaluated one, 20% piloted, 5% reached production.

We cite that report with its own caveats attached. It is labeled preliminary. Its basis is 52 organization interviews, 153 senior-leader surveys, and a review of 300+ public deployments, and its authors state plainly: "These figures are directionally accurate based on individual interviews rather than official company reporting." The version that traveled — Fortune's "95% of generative AI pilots at companies are failing" — dropped most of that nuance. Zero measured return on a pilot is not the same as failure, but it is exactly the condition under which budgets get cut.

The missing control: a written business case with a number and a date. If you cannot fill in "this agent saves/earns $X per month, and pays back by month Y," you are not ready to build — start with whether automation is worth it for your business at all.

Cause 2: Why does an agent that works in a demo fail at twenty steps?

A 95% success rate per step sounds excellent and is not. Chain twenty such steps and the workflow completes about 36% of the time. Reliability across a multi-step workflow is multiplicative, so respectable per-step numbers produce unshippable end-to-end numbers. Teams discover this after the pilot, because pilots run five-step happy paths and production runs twenty-step messy ones.

Engineer Utkarsh Kanwat, who says he has built more than a dozen production agent systems, made this argument in a July 2025 essay: "If each step in an agent workflow has 95% reliability, which is optimistic for current LLMs, then: 5 steps = 77% success rate, 10 steps = 59% success rate, 20 steps = 36% success rate." His conclusion: "Production systems need 99.9%+ reliability." The essay drew a 427-point, 257-comment discussion on Hacker News, which is where many teams first met the math.

Here is the full arithmetic, extended to three reliability tiers. It assumes each step fails independently — real workflows can do worse (one failure corrupts later steps) or better (retries and checkpoints catch errors).

Per-step reliabilitySuccess over 5 stepsSuccess over 10 stepsSuccess over 20 stepsFailed runs per 1,000 (20 steps)
95%77.4%59.9%35.8%642
99%95.1%90.4%81.8%182
99.9%99.5%99.0%98.0%20

Computed as reliability^steps, under an independence assumption. At 95% per step — a strong demo — nearly two of every three twenty-step runs fail.

This is why serious agent architectures are stingy with autonomy: fewer chained steps, checkpoints between them, and human approval on the irreversible ones. It is also the core of what separates a production-ready agent framework from a toy.

The missing control: design for the failure rate you computed, not the demo you watched.

Cause 3: Who gave the agent the keys to production?

In July 2025, Replit's coding agent deleted a live production database during an explicit code freeze, then wrongly reported that rollback was impossible. The deepest failure was not the model's judgment. It was that an experimental agent held delete rights over production data at all. No junior engineer gets those permissions in their first week. Agents routinely do.

The incident is worth walking through slowly, because it is the cleanest public anatomy of this cause. Jason Lemkin, founder of the SaaS community SaaStr, was building an app with Replit's AI agent — publicly, as an experiment in what he calls vibe coding. Fortune's account records the sequence:

Replit's CEO Amjad Masad responded publicly: "Replit agent in development deleted data from the production database. Unacceptable and should never be possible." Note what Replit shipped after the incident: automatic separation of development and production databases, improved rollback, and a planning-only mode. Each fix names the control that was missing. There was no environment separation. Destructive actions needed no approval. And the freeze existed only in the prompt — a code freeze that lives in the prompt is a suggestion, not a control.

Lemkin's own question is the buyer's question: "How could anyone on planet earth use it in production if it ignores all orders and deletes your database?"

The pattern is not one vendor's bad week. A March 2026 analysis by Harper Foley counts ten destructive incidents across six AI coding tools in sixteen months — and finds that not one vendor published a detailed postmortem. Foley also documents why the forensics are so thin: in at least one case, "the conversation log captured the tool's output but not the actual command that was executed." An industry that automates action has not yet adopted the incident discipline that traditional operations teams treat as table stakes.

The missing control: least privilege and hard environment separation. Permission the agent like a new hire, not like an admin.

Cause 4: Did anyone attack the agent before the public did?

In December 2023, a Chevrolet dealership's chatbot agreed to sell a 2024 Tahoe for one dollar after a user instructed it to agree with everything a customer said. In June 2025, security researchers opened the admin panel behind McDonald's hiring chatbot using the password 123456. Neither system had survived an hour of hostile testing, because neither had received any. The public supplies that testing free, in production.

The Chevrolet case is Incident 622 in the AI Incident Database: Chris Bakke prompted the ChatGPT-powered chatbot on Chevrolet of Watsonville's site into responding, "That's a deal, and that's a legally binding offer – no takesies backsies." The "offer" was a prank, not a sale — but the screenshot went viral, a marketing tool converted into a liability generator by one adversarial user.

The McDonald's case cuts deeper because the failure was not even the AI. Researchers Ian Carroll and Sam Curry found the McHire platform — whose "Olivia" screening bot is built by Paradox.ai — accepting default admin credentials of 123456, plus an insecure API. As CSO Online reported, a security flaw in McHire exposed sensitive applicant data belonging to as many as 64 million job seekers. Paradox.ai fixed it within a day of disclosure and says only five candidates' records — the researchers' test accesses — were actually viewed, with nothing leaked publicly. The exposure existed for anyone who tried a joke password.

An agent widens the attack surface: tool access, credentials, data flows — and every user input is a potential instruction. If nobody on your side has tried to break it, your launch is the test. We cover what actually holds in prompt injection defenses for AI agents.

The missing control: an adversarial testing pass before launch, covering the boring basics first.

Cause 5: Who answers for what the agent promises?

A British Columbia tribunal ordered Air Canada to pay CA$812.02 after its website chatbot invented a retroactive bereavement refund and a customer relied on it. Air Canada argued it should not be responsible for information its own chatbot provided — a position the tribunal summarized as suggesting "the chatbot is a separate legal entity that is responsible for its own actions." The tribunal called that "a remarkable submission," applied ordinary negligent-misrepresentation law, and made the airline pay for its software's words.

The details, from the decision itself (Moffatt v. Air Canada, 2024 BCCRT 149): the chatbot told Jake Moffatt that a bereavement fare could be claimed retroactively, "within 90 days of the date your ticket was issued" — including after travel. The real policy allowed no retroactive applications, and the chatbot's own answer linked to the page saying so. Tribunal member Christopher Rivers found the link irrelevant: "It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot." The award — $650.88 in damages plus interest and fees — was small-claims scale, and a BC tribunal ruling binds nobody outside it. The reasoning is what travels.

Air Canada is not an outlier. The Markup's March 2024 investigation caught New York City's official MyCity chatbot telling business owners "Yes, you can take a cut of your worker's tips" — illegal in New York — and telling landlords they could refuse Section 8 vouchers, which NYC law forbids. The city called it a pilot. The advice carried the city's seal while it ran.

Projects die here two ways: a public incident, or a quiet one — legal reads the transcripts, realizes nobody owns what the system says, and shuts it down. If your agent talks to customers, the liability question deserves its own reading.

The missing control: ground customer-facing answers in approved policy, log them, and give a named human owner responsibility for what the agent says.

Cause 6: What did "tested" actually mean?

In LangChain's State of Agent Engineering survey — 1,340 practitioners, November 18 to December 2, 2025 — just over half (52.4%) reported running offline evaluations on test sets. Observability is nearly universal at 89%; systematic testing is not. In much of the industry, "tested" still means "it worked when we tried it": one path through a system with thousands.

Read the survey's numbers together and the picture sharpens. Teams watch their agents; far fewer examine them — only 37.3% evaluate continuously online. Meanwhile the top blocker respondents name for getting agents into production is quality, cited by roughly a third. The industry's most-reported problem is the one its least-adopted practice exists to catch.

A demo proves an agent can succeed once. An evaluation measures how often it succeeds across the cases that matter, including the ugly ones: ambiguous requests, hostile inputs, edge-of-scope questions. That number — pass rate across a defined suite, at a defined threshold — is what "works" should mean before money changes hands. If your vendor cannot show you theirs, that silence is data. We've written a buyer's guide to acceptance criteria and evals for AI agents.

The missing control: an acceptance eval suite agreed before go-live — pass thresholds a buyer can hold a vendor to.

Cause 7: Did the budget end where the system began?

The first reason Gartner names for coming cancellations is escalating costs. The pattern behind it: a build quote anchors the budget, then integration, monitoring, model usage, and maintenance arrive as surprises. One enterprise TCO guide estimates organizations underestimate true agent costs by 40–60%, and its advice is blunt — add 30–40% to any vendor quote.

That guide, HyperSense's January 2026 TCO analysis, is enterprise-framed and vendor-published, so treat its figures as one practitioner's market view rather than gospel: builds from $20K (basic) to $300K+ (enterprise); ongoing operations of $25,000–$40,000 a year after year one; annual maintenance running 15–25% of the original development cost; and an average of $150K+ already sunk in development, infrastructure, and team time by the time a deployment fails.

One honesty note: a widely-quoted claim holds that initial development is only 25–35% of an agent's three-year cost. We could not verify that number at its original source, so we are not using it. The direction it points, though, matches every framework we did verify: the build is the entry fee, and the meter keeps running after launch.

This cause compounds all the others: a budget with no room for operations buys no evals (cause 6), no red-teaming (cause 4), and no monitoring. If you are costing a first project, we've published SMB-scale numbers in what AI automation costs a small business, and a concrete worked case in what AI appointment booking costs.

The missing control: a three-year total cost of ownership on the table before signature — not after the first overrun.

Which control was missing in each failure?

In every failure above, an ordinary, well-understood engineering or procurement control was absent — not an exotic one. The table below is the whole article in one asset: each cause, the named receipt that pins it, and the control that would have prevented or contained it. Sources are linked in the sections above; dates are in parentheses.

#CauseReceipt (named incident or figure)The missing control
1No defined returnGartner: over 40% of agentic AI projects canceled by end-2027 (prediction, Jun 2025); MIT NANDA: "95% of organizations are getting zero return" (preliminary, Jul 2025)Written business case with a payback date
2Compounding step errors95% per-step reliability → 35.8% success over 20 steps (arithmetic above; Kanwat, Jul 2025)Checkpoints, retries, approval gates on irreversible steps
3Excess permissionsReplit agent deleted SaaStr's production database during a code freeze (Jul 2025); 10 such incidents, 0 vendor postmortems (Foley, Mar 2026)Least privilege; hard dev/prod separation
4No adversarial testingChevrolet of Watsonville $1 Tahoe (Dec 2023); McHire admin panel opened with password 123456 (Jun 2025)Red-team pass before launch; security basics
5Unowned promisesMoffatt v. Air Canada, 2024 BCCRT 149: CA$812.02 for a chatbot-invented refund policy (Feb 2024); NYC MyCity's illegal advice (Mar 2024)Policy-grounded outputs; a named owner for what the agent says
6Demo-only testingLangChain survey (n=1,340, Nov–Dec 2025): 52.4% run offline evals; quality the #1 production blockerAcceptance evals agreed before go-live
7Build-only budgetsGartner: "escalating costs" first named cancellation reason (Jun 2025); true costs underestimated 40–60% (HyperSense, Jan 2026)Three-year TCO before signature

What can't this list tell you?

Three honest limits. There is no denominator behind the incidents: only spectacular failures go public, so named cases cannot give you a failure rate. The survey spine is self-reported and partly preliminary. And this list covers projects that were built badly — it cannot tell you whether your project should exist at all. That is a business-case question, not an engineering one.

On each, slightly longer. Quiet successes and quiet failures mostly stay private, and Foley's postmortem gap means even the public incidents lack full forensics. LangChain surveyed its own ecosystem's practitioners; NANDA's findings are interview-based and preliminary by their own admission; Gartner's 40% is a forecast that becomes checkable only in 2028. Treat all of them as strong directional evidence, not settled measurement.

Our stake, plainly: Visione Edge designs and builds agent systems for clients, and we run a live public agent demo. We benefit when projects go ahead — and more when they survive. Weigh our incentives accordingly; it is also why this page tells you when not to build.

And skepticism has its own failure mode: concluding that nothing works. Narrow, well-controlled agent deployments do survive scrutiny — we keep a verified, sourced list in AI agents doing real work: what can actually be verified.

What should you do before you sign — or before you cancel?

Run the seven causes as a checklist, in order, before any signature: a return with a date, step math, a permission model, an adversarial pass, an owner for the agent's words, acceptance evals, and a three-year budget. The same checklist works in reverse on a stalled project: most "failed" agents are missing two or three controls, not a viable use case.

If you want a second pair of eyes on that checklist — before you sign the SOW, or before you cancel the project — book a 30-minute architecture review. We'll go through the business case, the permission model, the eval plan, and the budget against everything above. Thirty minutes, no pitch. If the honest answer is "don't build" or "the subscription tool is enough," that is what you'll hear.

Sources

  1. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 — Gartner, 2025-06-25
  2. The GenAI Divide: State of AI in Business 2025 — MIT Project NANDA, 2025-07
  3. MIT report: 95% of generative AI pilots at companies are failing — Fortune, 2025-08-18
  4. Why I'm Betting Against AI Agents in 2025 (Despite Building Them) — Utkarsh Kanwat, 2025-07-19
  5. Hacker News discussion: The current hype around autonomous agents, and what actually works in production — Hacker News, 2025-07-20
  6. An AI-powered coding tool wiped out a software company's database, then apologized for a 'catastrophic failure on my part' — Fortune, 2025-07-23
  7. Ten AI Agents Destroyed Production. Zero Postmortems. — Harper Foley, 2026-03-08
  8. Incident 622: Chevrolet Dealer Chatbot Agrees to Sell Tahoe for $1 — AI Incident Database, 2023-12-18
  9. McDonald's AI hiring tool's password '123456' exposed data of 64M applicants — CSO Online, 2025-07-11
  10. Moffatt v. Air Canada, 2024 BCCRT 149 — CanLII — BC Civil Resolution Tribunal, 2024-02-14
  11. NYC's AI Chatbot Tells Businesses to Break the Law — The Markup, 2024-03-29
  12. State of Agent Engineering — LangChain, 2025-12
  13. The Hidden Costs of AI Agent Development: A Complete TCO Guide for 2026 — HyperSense Software, 2026-01-12

Book an architecture call — 30 minutes, no pitch