Before You Buy an AI Agent, Ask These Three Questions
The demo will work. Vendor demos always work, because they run on data the vendor controls and a workflow the vendor picked, and your business is neither of those things. What I watch owners do is buy the agent first and decide what it was supposed to improve second, which is backwards. It is why a pilot that looked brilliant in March gets quietly switched off in July with nobody willing to say out loud that it failed.
The conversation you run before you sign is the project. Three questions decide whether you bought a tool or a twelve-month distraction.
You can find those three questions on a dozen listicles by lunchtime. None of them show the framework being executed inside an actual small business, with the workflow step named, the success number attached, and the off switch agreed before anyone flips it on. I have helped 300-plus businesses across the USA, Canada, the UK, Singapore, Australia and New Zealand, in medical, ecommerce, local services, fintech, education and real estate, and scoping and running AI automations is a standing part of that work. So I can walk the whole procedure instead of handing you the questions and wishing you luck. For background, what an AI agent actually is and why so many projects get scrapped is its own read, but this one stands alone. This is the buying procedure.
The bill for a bad agent purchase arrives long after the demo
The license fee is the small number. The expensive part is the operations person you quietly assign to babysit the rollout, the four months of status meetings about why the outputs still need checking, and the fact that whatever was actually costing you money is still sitting there untouched while everyone talks about the agent.
Gartner said in a June 2025 press release, reported by Robert J. Szczerba in Forbes in July 2026, that more than 40% of agentic AI projects will be canceled by the end of 2027. Read the stated reasons before you read the number: escalating costs, unclear business value, inadequate risk controls. Not one of those is a model problem. All three are decisions somebody either made badly or never made at all, in a room, before the contract was signed.
Szczerba put the diagnosis in one line.
The ones that fail rarely die because the models were too dumb to do the work. They die because companies turn agents loose without a success metric, without access to the right data, and without a plan for what happens when the thing goes sideways. The coming cancellation wave is a management problem wearing a technology costume.
Robert J. Szczerba, Forbes, July 7, 2026
Three failures, three fixes, all available to you for free before you pay anyone. That is the whole thesis. A success metric, the right access, and a plan for going sideways are not technical artifacts. They are answers a business owner gives, and if the owner cannot give them, no vendor on earth can supply them afterward.
Gartner counted thousands of vendors and found about 130 building the real thing
Same Gartner research, same Forbes writeup. Of the thousands of vendors claiming agentic capability, Gartner judged only around 130 to be building something that deserved the label. Everyone else was repackaging. The industry term is agent washing, which means putting the word agent on a chatbot, a scripted automation, or a wrapper around someone else's model.
You do not need to know what is under the hood to catch this. You need to know what a real agent does differently in practice, and every difference shows up in a question you can ask in plain English on a sales call.
| What you are checking | Something that earns the word agent | A chatbot in a costume |
|---|---|---|
| How work starts | Runs a multi-step job from a trigger in your business, then reports what it did | Waits for a human to type a prompt, one step at a time |
| Contact with your systems | Named read and write permissions in the exact systems the job touches, granted by you | Demoed on sample data; integration described as coming soon or handled in onboarding |
| Definition of working | One business number the vendor will be judged on, agreed in writing before go-live | Usage stats: messages handled, tasks completed, seats active |
| Behavior when it is wrong | Stops, logs, escalates to a named person, and can be switched off by your staff in minutes | Keeps going, and you find out from a customer |
| Where it breaks in production | Vendor names their own past production failures unprompted: a missing invoice field, a duplicated customer record, a permission wall the demo never hit | Vendor insists it just works; you discover the edge cases yourself after go-live |
The last row is the one that separates people who have shipped from people who have sold. Szczerba's examples of how demos die in production are mundane. A field that is empty in your data and full in theirs. A customer who exists twice because your front desk spells the name two ways. A permission wall the sandbox never hit. Ask a vendor which of those killed their last deployment and watch what happens to the room.
Three questions, one sentence each, or keep your money
I do not ask about model architecture, and neither should you. The vetting works because these three questions are unanswerable by anyone who has not actually thought about your business, and answerable in one sentence by anyone who has.
Question one flushes out the vendors selling activity. A good answer sounds like a sentence with a noun and a direction in it, something like quotes sent within one business day, up from six in ten to nine in ten. A bad answer is a dashboard. Then ask who inside your business agreed that this is the number, because a target nobody owns is a target nobody defends. If the vendor's answer includes the phrase time saved, ask whose time, doing what instead, and how you would see that in the P&L, because self-reported hours are the easiest number in business to feel and the hardest to bank. I have written separately about how to prove an AI time saving is real, and the short version is that a saving nobody redeployed is not a saving.
Question two is where most small business deployments actually die, and it dies for an unglamorous reason: the data the agent needs lives in three places, one of which is a spreadsheet on someone's desktop. Ask what systems it reads, what it writes into, and who has to approve that access. Then ask the follow-up that vendors hate, which is what the agent does when the field it expects is empty. Access is also a risk question, not just a plumbing one, because opening your systems to an autonomous tool is a different exposure than opening them to a person who can be fired. Agent-ready and agent-safe are not the same thing, and the gap between them is where the quiet incidents happen.
Question three is the one owners skip because it feels pessimistic on a sales call. Skip it and you have bought a system that fails silently. Silence is expensive because the cost accumulates in customer trust, which shows up in your numbers a quarter later with no obvious cause.
The first deployment should be one step, one number, one off switch
The listicles stop at the questions. What you do with the answers decides the outcome. When I scope an automation for a client, the sequence never changes. Most of the elapsed time goes to waiting for data rather than doing anything clever.
Picture your own business through it. A physio clinic with two therapists and one person on the front desk, which is a scenario rather than a client of mine. The circled step is not patient communication, it is the rebooking follow-up after a discharge visit, because that is one step with a start and an end. The number is the share of discharged patients who book their next appointment within seven days, pulled from the practice software for the two weeks before anything changes, and the clinic manager is the name attached to it. Access is read on the appointment calendar and write into a drafts queue only, with no access to payments and no ability to send anything without approval for the first two weeks. The rollback is the manager revoking one integration key, triggered by any clinically inaccurate message or any message going to a patient flagged as discharged from care.
Everything about that scope is boring, and boring is the point. If the number moves, you have a result with a baseline behind it and a manager who will defend it. If it does not move, you have lost two weeks and one key, not a quarter.
The reason I scope this narrowly comes from watching the other side of it. When I filled a London ADHD clinic's calendar for three straight months, their bottleneck stopped being demand and became capacity, and they had to hire more specialists and outsource the overflow. The constraint moves. If you automate a step that was never the constraint, you have bought a faster version of something that was already fine, and you will feel busy while nothing improves.
The numbers that will lie to you after go-live
MIT's NANDA research, reported by Fortune on August 18, 2025, found that about 95% of enterprise generative AI pilots produced little to no measurable impact on profit and loss, with only around 5% seeing rapid revenue acceleration. That study rests on 150 leader interviews, a survey of 350 employees, and 300 public deployments, so treat it as a serious signal rather than a headline. The decision it should drive for you is simple: assume by default that your pilot will land in the 95% unless you can point at a number that moved.
Aditya Challapally, the lead author, described what the winners did differently.
Some large companies' pilots and younger startups are really excelling with generative AI. It's because they pick one pain point, execute well, and partner smartly with companies who use their tools.
Aditya Challapally, MIT NANDA lead author, quoted in Fortune, August 18, 2025
One pain point is step one of the scoping sequence, arrived at independently by researchers looking at hundreds of deployments. A 2026 Writer.com adoption survey found 97% of executives had deployed agents in the past year while only 29% saw significant return, which points the same direction, though it is one vendor's survey of its own market and I would treat it as directional rather than settled.
Your measurement plan needs three dates and one number. Day 30, does the single agreed number show any movement against the pre-agent baseline. Day 60, is the defect log shrinking or is your team still correcting the same category of mistake. Day 90, keep, rescope, or kill, decided against criteria you wrote down before go-live so the decision is not a debate about sunk cost.
The traps are all metrics that measure the agent rather than the business. Tasks completed tells you the thing ran. Messages handled tells you volume, which goes up when quality goes down. Self-reported hours saved is the worst of them, because everyone reports a saving and nobody can find it in a bank statement. Adoption rate measures whether your staff logged in. None of these are the number a clinic manager would defend in a meeting, and that is the test: if nobody would argue about it, it is not a metric, it is decoration.
Frequently Asked Questions
How do I tell if an AI agent vendor is selling a real agent?
Ask what it reads, what it writes into, and what it does when a field it expects is empty. Gartner reviewed thousands of vendors claiming agentic capability and judged only around 130 to be building something that deserved the label, so the word itself carries almost no information. A vendor who has shipped will describe a specific production failure from a past deployment without being pushed, usually something small like a duplicated customer record or a permission wall. A vendor who cannot name one has only ever run demos.
What should the first AI agent in a small business actually do?
One step in one workflow, with a clear start, a clear end, and a number already being tracked. Not customer service, but the follow-up message after a visit. Not marketing, but drafting the weekly quote chase for a person to approve. MIT's NANDA research found the pilots that succeeded picked a single pain point and executed it well. Pick the step that costs you money when it gets skipped, and confirm it is your actual bottleneck before automating it.
How long should I run an AI agent before I decide it failed?
Ninety days, with checkpoints at 30 and 60, and the kill criteria written before go-live. Spend the first two weeks with the agent drafting and a human approving every output, logging each correction, because that log tells you whether the errors are shrinking or repeating. At day 30 you are looking for any movement in your one agreed number against the baseline you measured before switching it on. If day 90 arrives with no movement and the same defects, switch it off and keep the learning, since Gartner attributes most cancellations to cost and unclear value rather than capability, and dragging a dead pilot into year two is how the cost part happens.
Every vendor selling you an agent is asking you to trust their technology. The vetting flips that, because all three questions are really asking whether your own business is legible enough to be automated, and most are not yet. That is not a reason to wait. It is a reason to run the questions on your next hire, your next agency, and your next software purchase, since a step nobody can measure and nobody owns was already broken before anyone mentioned AI. If you want a second pair of eyes on a scope before you sign something, book a call and bring the vendor's proposal with you.