Most Small Business AI Projects Fail. Stop Running Projects.

Share
Cover banner: Most Small Business AI Projects Fail. Stop Running Projects.

Somebody demoed a tool in March. You signed up, two people used it for a month, and now it sits on the company card next to a subscription nobody can explain. When your bookkeeper asks what the AI spend bought you this quarter, the honest answer is a shrug. That is not a technology problem and it is not a discipline problem. It is what happens when a small business runs AI as a project.

The pages ranking for this question all tell you to fix the foundations before you touch AI. Forbes says sort out leadership and document your processes first. GroundWorks says run an operations audit first. FullStack says build proofs-of-concept and buy vertical AI tools. My position is different, and it is the whole argument of this piece: a small business should never run an "AI project" or an "AI pilot" at all. The pilot framing is the failure mode. You attach one AI task to one number the business already tracks, and if you cannot name that number in a sentence, you do not wire anything in. I run my own shop that way, and it is the only version I have seen hold up past month three.

95%
Share of enterprise GenAI pilots that produced no measurable profit-and-loss impact, out of $30 to $40 billion spent. Companies with real budgets and full-time staff got nothing back, which should lower your confidence that a busier, smaller version of the same approach will work for you.
MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (52 executive interviews, 153 leader surveys, 300 public deployments).

A pilot is designed to end, which is why yours did

You scope it, you run it for six or eight weeks, someone writes a summary, and then the thing either gets "rolled out" into a workflow nobody prepared or it quietly dies while the subscription renews. MIT's NANDA initiative put a number on how that goes: 95% of enterprise GenAI pilots delivered no measurable P&L impact despite $30 to $40 billion invested, and only about 5% reached rapid revenue acceleration. If you are paying every month for tools that never touch a number you report on, that is real money buying you a feeling.

MIT was clear about where the gap comes from, and it is not model quality. The tools get bought or built, then never wired into real workflows, never measured against the business, and never owned by a specific person in production. That is a plumbing failure and an accountability failure wearing the costume of a technology failure.

I am on the integrated side of that divide in my own operation, which is why I can describe it rather than theorize about it. Over 17 years in search, I have built a set of local AI tools I use every working day: a content command center that runs my draft checks and archive audits, a live meeting transcriber that keeps deadlines and action items on screen while a call is happening, a Spanish-to-English interview translator, a PDF desk that replaced the Acrobat subscription, and a video captioner. All of them were built in Claude Code sessions on my own machine. None of them started as a project. Each one started because a specific task was eating a specific slot in my week, and each one is attached to work I would still have to do if the tool vanished tomorrow. That is the difference between AI in a workflow and AI in a trial.

QuestionThe AI project or pilotOne task tied to one number
What it starts withA tool somebody demoed or a competitor mentionedA number already on your dashboard or your P&L
Who owns itA committee, an intern, or whoever had time that monthThe person who already owns that number and reports on it
How it endsA deadline, a summary, and a subscription that outlives bothIt does not end. It stays in the workflow or it gets removed on a set date
How you judge itSeats used, staff enthusiasm, hours people say they savedThe number moved, or it did not
Documented outcome95% of pilots showed no measurable P&L impactThe ~5% that accelerated revenue were the ones wired into real workflows and measured
Failure-mode column: MIT NANDA, "The GenAI Divide: State of AI in Business 2025." Other rows describe the working method in this article.

You were sold a project because projects have a budget line

No vendor has ever opened a call by asking which number you already report on. They open with transformation, because transformation carries a longer contract. Gartner watched the same pattern in the agentic AI wave and published its verdict on June 25, 2025: over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls. Cancelation is not free. You pay for the license, the setup, the staff hours, and then you pay again in the credibility you burn with your team the next time you propose something.

“Most agentic AI projects right now are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.”
Anushree Verma, Senior Director Analyst, Gartner, in Gartner's press release of June 25, 2025.

The supply side is worse than the demo suggests. In that same release, Gartner estimates only about 130 of the thousands of vendors calling themselves agentic AI companies are the real thing, a practice the firm calls agent washing. Run the arithmetic on your own shortlist and the odds that all three tabs open in your browser are genuine are not good. A January 2025 Gartner poll of 3,412 webinar attendees found 19% had made significant agentic AI investments and 42% conservative ones, so most buyers are dipping a toe, which is exactly the behavior a pilot rewards and a wired-in task does not.

Small businesses copy enterprise behavior here, and it is the wrong body to copy. A 400-person company can absorb a dead pilot inside a departmental budget. You cannot. When your marketing lead spends hours every week babysitting a tool that writes drafts nobody ships, that is a real quarter of a real person's output. Before you sign anything, the three questions in my short vetting checklist for AI agents will kill about half the shortlist on the spot, and killing them early is the cheapest work you will do all month.

The 5% difference is not smarter tools, it is a number someone owns

McKinsey's State of AI survey puts roughly 88% of organizations using AI in some form, while most report no meaningful enterprise-wide EBIT impact, and agentic adoption stays concentrated in large enterprises. Near-universal usage, near-invisible profit. That combination only makes sense if what most companies bought was activity rather than a change to how work gets done.

“Agents don't fail because they're too advanced, they fail because they're not engineered for reality.”
Kore.ai, "AI agents in 2026: from hype to enterprise reality," company blog.

Kore.ai lists what "reality" means in practice: security reviews, compliance, identity management, audit trails, integration, and ambiguous ROI. You will not face an enterprise identity management review. You will face the small-business versions of the same friction, and they bite just as hard. Who has the login. Whether the output goes anywhere a customer sees it. Whether the person doing the task actually trusts the draft or quietly rewrites it from scratch, which is the most common silent failure I find when I look at how a team really works.

Ethan Mollick has argued this point consistently: enterprise AI adoption is an organizational problem, not an IT one. Unclear incentives, fear of losing a job, no permission to experiment. Your receptionist will not tell you the AI intake tool is worse than her own note-taking if she suspects the tool exists to replace her. She will use it while you watch and stop when you leave.

MIT found one buying pattern that clearly outperformed: purchasing specialized vendor tools succeeded about 67% of the time, while custom internal builds succeeded roughly one-third as often. Read that carefully before you conclude that my own toolkit contradicts it. I build because software is my trade and the maintenance sits with the person who wrote it. If nobody in your building can fix the thing at 11pm on a Thursday, buy the specialized tool and spend your energy on where it plugs in. That plugging-in question is the whole game, and it has its own ladder, which I laid out in the four levels between owning AI tools and running an AI system. Today's article is the entry move. That one is the climb after it.

Wire one task to one number you already track

Pick the number before you pick the tool. Not a new number, not a dashboard you would have to build, and not "productivity." Something already sitting in your reporting: quotes sent per week, calls answered inside 60 seconds, days from invoice to payment, new patient bookings, average handling time on a support ticket. A London ADHD clinic I worked with lived and died by one number, booked appointments, and when we moved it the clinic was full for three straight months and had to hire more specialists and outsource the overflow. Nobody in that room ever asked how many tools were involved. They asked about the diary. The sequence runs: name the number, write down today's value, find the task closest to it, hand both to the same owner, and set a removal date.

The five-step wiring sequence
1Name the number. One number your business already reports. If you have to build a report to see it, pick a different number.
2Write today's value down. On paper, dated, before anything changes. Without a baseline you will argue about memory in six weeks.
3Find the one task nearest that number. The step immediately upstream of it, done weekly or more often. Not the most annoying task. The closest one.
4Give it to the person who owns the number. Same human, both jobs. Split them and the tool becomes somebody's hobby.
5Set a removal date, not a review date. Thirty days out, the default is off. It only stays on if the number moved or the owner argues to keep it.
A working sequence, not research findings. The removal-date default exists because MIT NANDA found the common failure is tools never wired in, measured, or owned in production.

Step five, the removal date, is where most owners flinch, and it is the step that saves the money. Defaulting to off costs you nothing when the tool works, because the owner will fight to keep it. When the owner shrugs, you learned something for the price of one month. That is the opposite of a pilot, which defaults to on and calls the summary deck a result.

If you use an agency or a contractor, the exact sentence to send them is this: which single number will this move, who on my side owns that number, and what does it read today? A good partner answers in two lines. A weak one sends you a capabilities deck, and now you know.

Measure the old number, and ignore the usage dashboard

Every AI tool ships with a dashboard designed to make you feel good: prompts run, documents processed, minutes transcribed, seats active. None of those are your business. A support tool that processed thousands of messages while your first-response time did not move did nothing except make your team look busy in someone else's software. Read your own number, on the same report you read before, on the same day of the month.

The one-number gate, run it before you buy anything
You can name the number in one sentence. Out loud, without opening a laptop. If it takes a paragraph, you are describing a feeling.
You already know today's value. A number nobody currently watches will not suddenly get watched because you spent money.
One named person owns it. Not a department. A person who would notice within a week if it got worse.
The task happens weekly or more. Monthly tasks never build the habit, and the tool goes cold between uses.
You can rip it out in 30 days. Month-to-month billing, your data exportable, no other process quietly depending on it.
You would keep it with the usage stats hidden. If the answer depends on the vendor's dashboard, you are buying reassurance.
Six checks. Fail any one and the answer is not yet, not a smaller pilot.

The loudest vanity metric in this category is hours saved, because it is self-reported and nobody audits it. A team that saves six hours a week and fills those hours with more of the same low-value work has changed nothing about the business, and I walked through how to actually verify a time claim in how to prove an AI time saving is real. Apply the same suspicion to output volume. Publishing more, sending more, and drafting more are all easy to buy and none of them are the number. If the number moves and something else changed in the same window, a price change, a new hire, a seasonal swing, run one more cycle before you credit the tool.

The line in my own shop is judgment. I do not let AI run client strategy calls, and nothing reaches a client without my judgment on it. The transcriber sits beside a live call and keeps deadlines and action items on screen while I do the talking. The command center flags rule breaks in a draft and I decide what to do about each flag. Every tool I built assists a step and reports back. None of them holds the judgment, and none of them talks to a client without me in the room. If a vendor is selling you the judgment layer, that is the part to be slowest about.

Give it 30 to 60 days against the baseline you wrote down, because most of the tasks worth wiring in are weekly, and a handful of cycles is the minimum that separates a real shift from a good fortnight. Then make the boring decision. Keep it, or take it off the card.

Frequently Asked Questions

How much should a small business spend on AI to get started?

Spend less per month than the number you are trying to move is worth to you in a month, and insist on month-to-month billing so you can stop. The size of the budget is not what decides the outcome. MIT NANDA studied 300 GenAI deployments against an industry that has put $30 to $40 billion into the technology, and found 95% of pilots delivered no measurable profit impact, so money is clearly not the missing ingredient. Start with one paid seat attached to one task and one number, and let the result decide whether there is a second seat.

Should I buy an AI tool or have something custom built for my business?

Buy, unless someone inside your business can maintain what gets built. MIT NANDA found that purchasing specialized vendor tools succeeded about 67% of the time, while custom internal builds succeeded roughly one-third as often. Custom work fails on maintenance, not on the first version, and a contractor who disappears after delivery leaves you owning software nobody can fix. I build my own tools because software is part of my trade, which is exactly the condition most businesses do not have.

How do I know if an AI tool is actually working for my business?

Compare the one number you wrote down before you started against the same number today, read from the same report you always used. Ignore the vendor's usage dashboard, seats active, and any hours-saved figure your team estimates from memory, because none of those appear on your P&L. Give it 30 to 60 days if the task runs weekly, which is enough cycles to tell a real shift from a good fortnight. If the number is flat and the person who owns it would not fight to keep the tool, remove it.

Most owners reading this already have the number. It is on a report they open every Monday, and it has not moved in a year of buying software that was supposed to move it. If you want a second pair of eyes on which task sits closest to that number, book a call and bring the report. The AI spend was never the problem. Every result in your business you can actually name had an owner, a baseline, and one number attached to it, and the software you bought had none of the three.

Read more

Free, No Commitment

Find out exactly where your AI visibility is leaking. In 30 minutes.

No pitch. No fluff. A straight diagnostic on your specific situation and the single highest-leverage fix to make right now.