The AI Work You Only Do Because AI Made It Cheap
You are shipping more than you were a year ago. More posts, more decks, more dashboards, more versions of everything. The team is busier, the output folder is fuller, and the revenue line has not moved an inch. That gap is not a measurement lag and it is not your team lying on their timesheets: AI made a pile of tasks cheap, cheap tasks get chosen, and a large share of what you produced last month is work nobody would have paid for at last year's prices.
The fix is not a better tool, a tighter prompt, or another pilot. It is one pass over the task list with one question attached to every item on it. Stop auditing the software. Audit what the software talked you into doing.
Your output doubled because the price of output collapsed
Task substitution is the unglamorous name for what happened. When the cost of doing something falls far enough, you start doing things you would never have commissioned at the old price, and those things arrive in the week wearing the exact clothes of real work. Anthropic ran an internal survey, cited by METR in February 2026, that put a number on it: 27% of Claude-assisted work "wouldn't have been done otherwise." Read that as a budget line rather than a statistic. Better than a quarter of the output was work that would not have been commissioned at last year's prices. Some of it is worth something. None of it was worth enough for anyone to buy it before.
27%of AI-assisted work “wouldn’t have been done otherwise.” That share would not have been commissioned at last year’s prices.Anthropic internal survey, cited by METR, February 2026.
METR then went looking for the same effect inside its own building. In its February 2026 analysis of coding agent transcripts, a human-labeled validation sample of eight pull requests contained 3.5 that the staff member would not have done at all without AI. That was roughly 47% of that individual's estimated no-AI task time. Close to half the clock, spent on work that would never have started twelve months earlier. Nobody wasted an afternoon. Everybody was productive. The output simply had a different relationship to money than the output it replaced.
Three explanations get offered when the AI return refuses to show up. Measurement lag, meaning the gains are real but the reporting is slow. Verification overhead, meaning review now eats the time AI saved. Rework, meaning the output needed fixing. All three are real and I have written about the second one at length in how AI adds review work to the same tasks your team already had. The fourth explanation is the one that never makes it onto the slide, and it is the one that explains a flat revenue line better than the other three combined: the tasks themselves changed. Different work entered the week, not the same work done faster. I had to build two guardrails into my own publishing operation the moment production got cheap, for exactly this reason.
The cheap task gets chosen, and the choosing never gets logged
METR gave this a name. In "Task Substitution and Uplift," published May 8, 2026, Tom Cunningham and Parker Whitfill describe a category of work that exists only because the price fell.
“In the case where the uplift in value is much lower and the person is only doing the task because AI has made it cheap, we have been calling it a ‘Cadillac Task’.”Tom Cunningham and Parker Whitfill, METR, “Task Substitution and Uplift,” May 2026.
The mechanism is not laziness. It is arithmetic, and the same paper spells it out: "Therefore, we should expect people to substitute towards tasks that take advantage of AI. Their uplift on new tasks will be very high, but their uplift in value may be much lower." People move toward the work the tool is good at. Anyone would. A METR developer participant described the pull in their own words: "I found I am actually heavily biased sampling the issues … I avoid issues like AI can finish things in just 2 hours, but I have to spend 20 hours. I will feel so painful if the task is decided as AI-disallowed." That is a smart professional openly admitting they now pick work by what the tool handles well. Your marketing coordinator is doing the same thing on Tuesday morning and has never said so out loud, because nobody asked and there is no field for it in the project tracker.
Picture your own business. A physio clinic with four staff and one person handling marketing. Eighteen months ago that person wrote two social posts a week and one email a month, because that is what a human could do between reception shifts. Today the same person publishes ten posts, three emails, a monthly newsletter, a competitor teardown, and a live dashboard of ad metrics. Every one of those artifacts is decent. Not one of them was requested by a patient, a referring GP, or the owner. The clinic's phone rings the same number of times as before, the marketing person is visibly busier, and the owner cannot explain the discrepancy without accusing someone of slacking.
Cadillac tasks in small-business life✓The dashboard nobody opens. Built in an afternoon because it was finally possible, checked twice in the first week, never again.✓The second version of the brochure. The first one was fine. The second exists because a variant now costs ten minutes.✓The weekly report that used to be monthly. Nothing changed about the decisions it feeds. Only the cost of generating it changed.✓Ten social posts instead of two. Five times the volume, the same audience, the same booking form.✓The competitor teardown nobody acts on. Genuinely interesting. Sitting in a shared drive since March.Category named in METR, “Task Substitution and Uplift,” May 2026. None of these arrived because a customer asked.
A 3x speed gain and a 1.4x value gain are two different numbers
Cunningham and Whitfill put the relationship in one line: "uplift on old tasks ≤ uplift in value ≤ uplift on new tasks." Translated for a business owner, the number your team quotes you comes from the right-hand side of that chain, because new AI-friendly tasks are where the tool looks most impressive. The number that pays your staff sits in the middle. Their worked example shows +33% on the old work and +50% on the new, with actual value somewhere between the two. The paper's more extreme version runs the same shape at a bigger scale.
| Measure | What it actually counts | Worked example |
|---|---|---|
| Uplift on old tasks | Work you were already doing and paying for before AI | +33% |
| Uplift in value | What the business is actually better off by. Always between the other two | Between +33% and +50% |
| Uplift on new tasks | Work that only exists because AI made it cheap. The most flattering number in the room | +50% |
Source: METR, “Task Substitution and Uplift,” Cunningham & Whitfill, May 8, 2026.
The self-reports line up with the theory in an uncomfortable way. METR surveyed 349 technical workers between February and April 2026. Median self-reported speed gain: 3x. Median self-reported value gain: 1.4x to 2x. The same people, in the same survey, saying the work goes three times faster and is worth between 1.4 and 2 times more. Both answers are honest. They are answers to different questions, and every vendor pitch you have sat through quoted the first one.
| What was asked | What 349 technical workers said | What it means for an owner |
|---|---|---|
| Speed gain | Median 3x | Real, and the number your team will quote you |
| Value gain | Median 1.4x to 2x | The number that shows up in your accounts, if it shows up |
| Trajectory | 1.3x (Mar 2025), 2x (Mar 2026), forecast 2.5x (Mar 2027) | Steady climb, not a cliff. Plan in years, not quarters |
| Attachment to the tools | Median would give up 29% of one month’s salary to keep access | Nobody is faking the enthusiasm. Do not read the audit as an accusation |
| Big claims, checked | Of 7 people claiming 10x or more on two or more measures, METR reviewed public outputs and was confident the respondent overstated in 2 of 2 checkable cases | Treat any 10x claim as a feeling, not a finding |
Source: METR, “Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity,” Joel Becker, May 11, 2026. Survey run February to April 2026.
METR's own staff reported the lowest gains of any subgroup in that survey. The people who measure this for a living claim the least from it. Amy Deng, writing up METR's February 2026 transcript analysis, was blunt about why raw time savings mislead.
"People likely do not create 10x as much value with AI, even if we observe a 10x time savings factor on tasks that people do with AI."
Amy Deng, METR, "Analyzing coding agent transcripts to upper bound productivity gains from AI agents," February 17, 2026.
"The methodology overestimates the productivity multiplier on Claude Code-assisted tasks when people use AI to complete low-value but time-consuming work (we call them 'Cadillac tasks')."
Amy Deng, same paper. Her measured range across 5,305 Claude Code transcripts from seven METR technical staff in January 2026 was a time savings factor of about 1.5x to 13x, and she framed it as a soft upper bound rather than a productivity multiplier.
Anthropic's own aggregate estimate, from Tamkin and McCrory in 2025, lands near 17%: Claude covering roughly 11% of wage-weighted tasks in the O*NET occupational database, with an 80% time reduction where it gets used. METR argues even that number probably runs high, because substitution happens inside task groups as well as between them. So the most optimistic well-built estimate in the literature is a fraction of what your team feels, and the researchers think it is still generous.
Be careful with the headlines about this research. Secondary coverage has been reporting that METR's famous 19% developer slowdown result "reversed" into an 18 to 20% speedup. The February 2026 source does not say that. It reports a central estimate of -18% with a confidence interval running from -38% to +9% for returning developers, and -4% with an interval of -15% to +9% for new ones. Both intervals cross zero, and METR itself calls the central estimate "likely a bad proxy," which is why the organization published a note on February 24, 2026 about changing its developer productivity experiment design. If your agency or your software vendor is quoting the reversal at you, they read a summary of a summary. Fortune ran a piece by Sasha Rogelberg in February 2026 reviving the old Solow paradox comparison for exactly this reason, and Dr Philippa Hardman has been updating an analysis titled "The Illusion of AI Productivity Gains" through April 30, 2026. The doubt is not a fringe position.
| What the headlines say | What the METR data says |
|---|---|
| The 19% slowdown reversed into an 18 to 20% speedup | No reversal is reported in the source |
| Returning developers got faster | Central estimate -18%, confidence interval -38% to +9%, crossing zero |
| New developers got faster | Central estimate -4%, confidence interval -15% to +9%, crossing zero |
| The result is settled | METR calls the central estimate “likely a bad proxy” and announced a redesigned experiment on February 24, 2026 |
Source: METR, February 2026 results and the February 24, 2026 experiment-design note. Secondary-coverage claim as circulating in newsletters and aggregators, 2026.
Run the contractor test on last month's output
Sit down with everything your business produced in the last thirty days. Not the plan, the actual artifacts. Then apply a single filter to each one: would I have paid an outside contractor real money to make this, at last year's prices? Not "is it good." Not "did it take long." Would money have left the building for it. The share that survives is your AI return. The hours saved number is not, and never was.
The monthly Cadillac task audit1List the artifacts, not the activities. Every deliverable that got made last month, one line each. Reports, posts, decks, pages, dashboards, variants.2Apply the contractor test. Circle the items you would have paid an outside supplier to produce at last year’s rate. Circle fast, argue later.3Name who asked. For every uncircled item, write the customer, the deal, or the decision it served. No name means no requester, which means it chose itself.4Kill it in writing, and keep the list. A killed task that is not logged comes back in six weeks with a new title. The log is the whole mechanism.5Recount next month. Track the surviving share month over month. That percentage, not hours saved, is the number to put in front of your team.Built on the substitution finding in METR, “Task Substitution and Uplift,” May 2026.
I run two versions of this on my own publishing operation, and I did not build them because they sounded rigorous. I built them because AI made writing cheap enough that I could have filled a content calendar forever with pages that had no reason to exist. The first is an information-gain go/no-go log. No idea gets drafted until I write one sentence naming what a reader still would not know after reading the three pages currently ranking for that query, backed by something only I have: a real number from my own work, a client outcome, actual pricing. No sentence, no post, and the killed idea stays logged so it cannot quietly return. That is the same test I explain in the 10% rule for information gain. The second is a cannibalization check run against the live sitemap before I draft anything, so a new page cannot compete with a page I already published. Both guardrails exist for one reason: cheap production invents work that feels like progress, and I am not immune to it after seventeen years in search.
Your version of that is a standing question in the weekly meeting. Not "what did you get done," which rewards volume. Ask "what did you decide not to make this week, and why." The first time you ask it, expect silence. That silence is the finding.
If an agency is producing for you, the question to send them this week is short: which of last month's deliverables would you have quoted me for in 2024, and what did each one change in my pipeline. A good partner will answer with a shorter list than the invoice implies and tell you which items they kept because they compound. A weak one will send you a volume report. That distinction is the whole difference between a business that got value out of AI and a business that got activity, which is the pattern I unpack in why most AI pilots stall while a few companies convert speed into money.
The four numbers that will flatter you, and the three that will not
Four metrics will tell you everything is working while your bank balance disagrees. Hours saved, because it is self-reported and measures the wrong side of the substitution. Output volume, which now measures the price of production rather than the demand for it. Seats and licenses, which measure adoption. And any multiple above 5x: seven survey respondents claimed 10x or more on two or more measures, METR reviewed the public outputs of the ones it could check, and in both checkable cases the claim was overstated.
| Metric | Why it misleads, or what to do with it |
|---|---|
| Hours saved | Self-reported, and it measures the wrong side of the substitution. Ignore it |
| Output volume | Now measures the price of production, not the demand for it. Ignore it |
| Seats and licenses | Measures adoption, not return. Ignore it |
| Any multiple above 5x | Both 10x-or-more claims METR could check were overstated. Treat as a feeling, not a finding |
| Surviving share of the contractor test | Your real AI return. Track monthly |
| Revenue per person | Where substitution shows up first. Check quarterly |
| Length of the kill log | Direct evidence your team is choosing work. Review when it stops growing |
Metric traps and cadences from the audit above; the 10x check is from METR’s May 11, 2026 survey write-up.
Three numbers are worth the effort. The surviving share from your contractor test, tracked monthly, which should climb as the killed list grows. Revenue per person, checked quarterly, because substitution shows up there before it shows up anywhere else. And the length of the kill log, which is the only direct evidence that your team is choosing work instead of accepting whatever the tool made easy. Check the first monthly, the second quarterly, the third whenever it stops growing. Weekly measurement of any of them will produce noise and a bad argument.
The timeline in that survey is a slope, not a step change. METR's respondents put their retrospective gain at 1.3x in March 2025, 2x in March 2026, and forecast 2.5x for March 2027. A business that fixes its task selection this quarter compounds against one that keeps buying tools for three more years. The tools are not the variable anymore. The task list is.
Frequently Asked Questions
How do I tell which AI work is actually worth paying for?
Apply one filter to last month's output: would you have paid an outside contractor to produce this at 2024 prices? Anything that fails that test was chosen because it got cheap, not because someone needed it. METR named this category Cadillac tasks in its May 2026 paper, and an Anthropic internal survey cited by METR in February 2026 found 27% of AI-assisted work "wouldn't have been done otherwise." The surviving percentage of your output is your real AI return, and it is a far more useful number than hours saved.
My agency says AI lets them publish four times as much. Is that a good deal?
Only if the extra volume is work you would have commissioned anyway. Four times the pages at the same quality bar means the agency moved toward tasks their tools handle well, which is exactly the substitution METR documented. Ask them which of last month's deliverables they would have quoted you for two years ago and what each one changed in your pipeline. If the answer is a volume report rather than a shorter list with outcomes attached, you are buying activity.
Should I cancel the AI tools if the revenue has not moved?
No, because the speed is real and your team knows it. In METR's May 2026 survey of 349 technical workers, the median respondent said they would give up 29% of one month's salary to keep access. Cancelling punishes people for a selection problem you have not fixed yet. Run the monthly task audit first, kill the work nobody asked for, and see whether revenue per person moves over the next two quarters before you touch the subscriptions.
You did not get scammed. Your team is not padding their hours. AI genuinely made a category of work cheap, and cheap work gets chosen by good people with full calendars, which is why the busiest quarter of your business can also be its flattest. The audit takes one sitting and it will be uncomfortable, because most of what you cut will be work you were quietly proud of. If you want a second pair of eyes on last month's output before you start cutting, book a call and we will run the contractor test together. The question worth sitting with tonight is not whether AI made your team faster. It is which of the things on next week's plan would have existed at all if it had not.