Prompt Engineering Is Overrated. Verification Is the Real Skill.
Somebody forwarded your team a PDF of 200 power prompts last year. Maybe you paid for a course, or felt behind because you didn't. And the copy your AI produces still gets rewritten before it reaches a customer, which means the tool is billing you twice: once in subscription fees, once in the hours your people spend cleaning up after it.
The gap in your business is not prompt wording. It is verification, meaning whether every AI task in your operation has a built-in way to check the answer before anyone acts on it. Search this topic and you will find formulas, templates, and role-play tricks. None of those pages tell a non-technical owner that the skill has moved from writing the prompt to building the check, or how to install that check in a real marketing and content operation. I can, because I run one: I set the angle, every claim gets fact-checked against a real source before publish, drafts run through fixed rule checks, and nothing ships unverified.
Forty percent of the time AI saves you goes straight back out the door
Workday's research, reported by The Next Web, found that 85% of employees save between one and seven hours a week using AI, and that nearly 40% of the time saved is lost to rework, including fixing the AI's own output. Do that arithmetic on a person in the middle of that one-to-seven-hour range, saving five hours a week, and you get two hours a week spent correcting a machine. Six people on that pattern is twelve hours weekly, spent on work that never appears on anyone's task list.
The conservative baseline points the same direction. The Federal Reserve Bank of St. Louis, in its November 2024 "State of Generative AI Adoption" report, put average savings for gen-AI users at about 2.2 hours a week, or 5.4% of a 40-hour week. That was a 2024 snapshot, not a 2026 one, and it is worth holding onto anyway: the base is modest, so a 40% rework tax on top of it does not leave much.
Rework is invisible in a way that ordinary work is not. Nobody logs "spent 40 minutes checking whether the stat in paragraph four exists." It shows up as a marketing lead who feels busier than last year, a founder who reads every outgoing email again, a team that quietly stopped trusting the drafts. I wrote about this pattern in detail in AI isn't lightening your team's load, it's adding work, and the mechanism is the same every time: the tool moves labor from producing to checking, and nobody redesigns the job around the checking.
You almost bought a course for a problem that already got solved
In 2023 the models were brittle. Phrasing genuinely mattered, small changes flipped the output, and the people who figured out the incantations got better results than everyone else. That was real. It stopped being the main constraint as the models got better at inferring what you meant from an ordinary sentence.
Easier means less of a differentiator. If a skill gets easier every quarter, buying a course in it is buying a depreciating asset, and the depreciation is fast. Boris Cherny, who created and runs Claude Code at Anthropic, made a blunter version of the same point in an interview with Y Combinator's Diana Hu, reported by Roger Montti at Search Engine Journal on July 30, 2026. His advice, in paraphrase: stop chasing the "one weird trick" prompt hacks that circulate on LinkedIn and Twitter, and work empirically instead, meaning test what actually happens with your own tasks rather than copying someone's screenshot.
Prompt courses keep selling for a simple reason. Wording is teachable in an afternoon, it demos well, and it feels like a skill you can put on a slide. Verification is unglamorous operations work: source documents, sign-off, a log of what got corrected. Nobody builds a personal brand on that. It is also the part that decides whether the tool makes you money, and the purchase is always the easy half.
There is a second reason owners default to prompts. Prompting is something you do alone at a keyboard, and it requires no conversation with your team about who is accountable when the output is wrong. Verification forces that conversation. Most businesses would rather buy a template. None of the pages selling those templates will tell you the skill has moved, but the people who build the tools will, and they already have.
The people who build these tools talk about checking, not phrasing
Cherny's recommendation for getting real work out of a model has two halves, and most businesses do neither. Give the AI a task that seems slightly too hard, and give it a way to check its own work, not an easy task dressed in a beautiful prompt. A hard task with a test attached.
Read that as an operations instruction rather than a technical one. The man who built and runs Claude Code at Anthropic is not saying your team writes bad prompts. He is saying they hand the machine a job with no way to tell whether the job came back correct, and then act on the answer because it reads well. Confidence and accuracy come apart in these systems. Software has spent decades training your staff to treat them as one thing.
Addy Osmani, an engineering leader at Google Chrome, calls it the 70% problem in "The 70% Problem: Hard Truths About AI-Assisted Coding" on addyosmani.com: AI gets you most of the way to done fast, and the last stretch, the edge cases and the judgment calls, takes longer than everything before it. He wrote the framework for code, where the same verification tax lands on AI-written code, and it maps just as cleanly onto a service page, a patient email, or a monthly client report.
| The first 70% | The last 30% |
|---|---|
| Arrives in minutes and looks finished | Takes longer than the first 70% and needs a human |
| Structure, tone, the obvious points | Edge cases, exceptions, judgment calls |
| Better prompting improves it to roughly 80 to 85% | Better prompting does not close it at all |
| Where the demo happens | Where the wrong price, the wrong policy and the wrong claim live |
That bottom row is the money row. The 30% you cannot prompt your way out of is exactly the 30% that gets you a refund request, a bad review, or a compliance letter. Which is why the answer is a process, not a phrase.
Build the check before you write the prompt
The loop has five steps, and an owner can install all of them in a week without learning a single technical term. None of them require new software.
Picture your own business. A clinic uses AI to draft replies to patient intake enquiries, because the front desk is drowning and every reply is written from scratch. The drafts are good. They read warm, clear, and on brand. One of them tells a prospective patient that the clinic accepts her insurance, because the model inferred it from context rather than reading it anywhere. She books, attends, and finds out at reception. Now you have a refund, a wasted clinical hour, a one-star review that does not expire, and a front desk that no longer trusts the drafting tool at all.
The verification version of that same task: the AI drafts against the clinic's own intake policy document, and every reply comes with a four-line list of the factual claims it made, which insurers, what the price is, whether a referral is needed, what to bring. The front-desk lead scans that list against the policy sheet. That scan takes thirty seconds. Nothing containing a clinical instruction leaves without the clinician. Same tool, same speed gain, and the failure mode that costs real money is closed.
I booked a London ADHD clinic solid for three straight months, and the demand pushed them to hire more specialists and outsource the overflow. A jump like that is precisely when a front desk starts drafting replies with AI, because the alternative is a backlog of unanswered enquiries. Growth is the moment verification stops being optional, and it is also the moment nobody has time to design it. Build it before you need it.
If you use an agency, this is the question to put to them on the next call, word for word: "For each deliverable you produce with AI, what is the check, who performs it, and can I see last month's correction log?" A team doing this work will answer in under a minute. A team that has never thought about it will talk about their prompt library.
Track the correction rate, and ignore the numbers that flatter you
The metric that matters is the share of AI outputs that needed a factual fix before they went out. Count it for four weeks, by task type. A clinic drafting fifty intake replies a week with six corrections has a 12% correction rate, and knowing that is the difference between managing the tool and hoping. If the rate drops as you refine the source documents, the system is working. If it stays flat while output volume rises, you are scaling the rework, not the work.
Two more numbers support it. Time to verify, tracked separately from time to draft, because the verify column is the one that eats the savings. And escaped errors, meaning mistakes a customer or client caught rather than your team. The target for that one is zero, and a single escape justifies redesigning the task the same day.
Drafts in my own operation arrive AI-assisted and leave human-verified, and the fixed rule checks exist to catch exactly that failure, a confident statistic with no source behind it, before anything publishes. A number that cannot point to a source dies in the draft, not on a page carrying my name. A single caught fabrication pays for a year of thirty-second scans.
| Real signal | Vanity metric it replaces |
|---|---|
| Correction rate per task type, tracked weekly | Number of drafts, posts or emails produced |
| Minutes spent verifying, logged separately | Self-reported "hours saved" in a staff survey |
| Escaped errors a customer found | Seat licenses activated across the team |
| Tasks with a named human on the gate | Size of the shared prompt library |
Self-reported time savings deserve the most suspicion, because people report the drafting time they saved and forget the checking time they spent, which is exactly how a 40% rework tax hides inside a happy survey. If you want to test whether your own savings survive contact with reality, I laid out the method in how to check whether your AI time savings are real. Run it before you buy more seats.
Frequently Asked Questions
Do I still need to learn prompt engineering?
Learn the basics, which take an hour: say what you want, give the AI the actual document to work from, and tell it who the output is for. Beyond that, the returns fall off fast. Ethan Mollick, writing at oneusefulthing.org, has said that prompt engineering as a task has gotten easier, and Boris Cherny of Anthropic advises skipping the trick-of-the-week posts from social media influencers and testing on your own real tasks instead. Spend the course budget on building checks into your workflow, because that skill does not depreciate when the next model ships.
How do I check AI output when I'm not an expert in the subject?
You do not need subject expertise to verify the business output you send out every day, you need a source of truth to compare against. Ask the AI to list every factual claim it made alongside the draft, then match that list to a document you already trust: your price list, your policy sheet, your analytics export, or a live page you can click through to. Anything the AI cannot point to a source for gets deleted rather than published. For medical, legal or financial statements, no comparison document is enough on its own, and a qualified person signs off every time.
My team already uses AI with no checks. Where do I start?
Start with the one task where a wrong answer costs you the most money, usually anything that goes directly to a customer with a price, a date or a promise in it. Put a named person on that gate this week and log every correction they make for a month. That log will show you which tasks need a better source document and which need to come off AI entirely. Expanding from one gated task to five is straightforward once people see the corrections in writing.
The businesses getting real returns from AI right now are not the ones with the best prompts. They are the ones who decided in advance how they would know the answer was wrong. In 17 years across 300-plus businesses, the pattern has held for every tool that ever promised to do the work for you: the constraint is never the tool, it is whether anyone built a way to catch it failing. That is a management decision, not a technical one, and it is yours, not your IT provider's.
If you want a second pair of eyes on where your AI output is quietly costing you rework, and what a verification gate would look like inside your actual workflow, book a call and we will map it against how your team works today.