research_paper
AI Marketing Platform Evaluation: Everyone Quotes The 19 Percent Study. The People In It Held Work Back.
September 16, 2026 · 10 min read · Scout7
An AI marketing platform evaluation often fails before it starts. The 19 percent study shows how held-back work decides the answer.

Introduction
You gave an AI tool some of your work. Maybe a week of it. Then you decided whether it was any good.
A big company would call that an AI marketing platform evaluation. It is really just one person trying a tool for a week and making their mind up.
Here is the problem with that. You chose which jobs to give it. You also chose which jobs to keep for yourself. That second choice is the one nobody writes down. And it can decide the answer before the tool does a single thing.
This is not a guess. It happened inside a proper study, run by researchers, with real money behind it. They found it, they published it, and almost nobody quotes that part.
What you get from this page:
- The story of the most quoted number in this whole argument, and what it actually measured.
- The sentence those same researchers published later, which changes how you should read it.
- One free habit that makes your own test worth something.
You do not need any background for this. If you have ever tried a tool for a week and made your mind up, you are the right reader.
The Number Everybody Quotes

On 10 July 2025 a research group called METR published a study. It went everywhere.
They took 16 experienced developers. Not students. People who had worked on the same code for years. Then they gave them 246 real jobs from that code. Bug fixes and small features. Real work, not puzzles.
On some of those jobs the developers were allowed to use AI. On the others they were not. A coin toss decided which. Then the researchers timed them.
Before starting, the developers guessed AI would make them 24 percent faster.
The clock said something else. With AI, the same kind of job took them 19 percent longer.
Now the strange bit. When it was over, the researchers asked them how it had gone. The developers said AI had made them about 20 percent faster. They had just been slower. They could not tell.
That is the finding worth keeping, and nobody argues with it. A person cannot feel their own speed.
One more thing, and it matters. In that same paper the authors wrote that their result does not show AI fails to speed up most developers. They also wrote that they were not claiming their 16 people stood for most software work.
So the study never claimed the thing people quote it for. Other people added that bit.
Your Test Is Decided Before The Tool Starts Work

Now the part that got left out.
The same researchers kept going. On 24 February 2026 they published an update. They had changed how the study worked, and they said why.
The new round was bigger. 57 developers, 143 different code projects, more than 800 jobs.
And this time they asked the people taking part a simple question. Are you handing in all your work?
Here is what they wrote:
"When surveyed, 30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI."
Read that slowly. Between a third and a half of the people in the test were keeping some jobs out of it. Not to cheat anyone. They just did not want to do those particular jobs the slow way.
So picture a running race. The fastest people in town decide to stay home that morning. Everyone else runs. The stopwatch works perfectly and every time on the board is true. But the race did not answer the question you asked it.
That is what happened here. The jobs that never went in are the ones that would have moved the number most.
There was a second hole, and the researchers named it themselves. People who were sure AI helped them would not sign up at all, because the study might take it away. In their words, the study is "systematically missing developers who have the most optimistic expectations about AI's value." Put plainly: the keenest users stayed out.
They even tried money. The first round paid 150 dollars an hour. The new round paid 50. That changed who volunteered, and it still did not fix this.
Your own week-long test has the same hole, and nothing is holding the list. You pick each job as it comes up, on a busy day, with nobody checking.
And it bends both ways, which is why it is so easy to miss.
- Give the tool the boring jobs and keep the rest, and it will look brilliant.
- Keep the jobs you care about most for yourself, and it will look useless.
Same hole. Opposite answers. Both of them feel like evidence at the time.
The Biggest Win Has The Same Hole

Now the fair part. This is not an argument that AI does not work.
The biggest real-world result goes the other way, and it is a big one.
Six researchers, He, Agarwal, Denisov-Blanch, Azaletskiy, Koyejo and Vasilescu, followed one mid-sized company for over two years. January 2024 to April 2026. 802 developers. 196,212 pieces of work handed in for checking.
From the middle of 2025 the company had told everyone to double how much they shipped. By April 2026 they had. The work coming out reached 2.09 times what it was before. The authors call it "among the largest gains reported from a field deployment of AI coding tools to our knowledge."
So where people really did use these tools, a lot more work came out. That is worth saying plainly and without any hedging.
What happened to the checking is interesting too. Each person checking work had roughly twice as much to get through. Machines ended up doing more of the checking than people did. And the amount of work that had to be undone afterwards did not change. So it was not quietly getting worse.
Now the part that makes this section belong here.
Nobody decided who used the tools heavily. The developers decided that for themselves. The authors say so themselves. How much people used the tools was "not randomly assigned". So they treat the result as pointing at a real effect, not as proof of one.
So part of that doubling could be something else. The people who were always going to get more done are the ones who used it most.
Look at the shape of that. The famous bad result and the biggest good result have the same hole in them. Somebody chose. And the choosing moved the number.
Then They Had To Go Back To Asking People

There is one more turn in this, and it is not a joke at anyone's expense.
Once the trial stopped giving a clean answer, the same group ran a survey instead. On 11 May 2026 they asked 349 people who work with these tools how much faster they were. Software engineers, researchers, students, and people running teams.
The middle answer was three times faster.
Then, in the same piece of writing, they warned the reader about their own survey. Their words: "survey results are not necessarily grounded in reality." They also pointed back at earlier work where people got their own time wrong by 40 percentage points.
So the group who showed that a feeling is not a measurement ended up having to collect feelings.
That is not a failure on their part. It is how hard this is to measure. These people had a budget, a plan, and a coin toss deciding who did what. And they still could not get a clean answer. Your quiet Tuesday trial will not manage it by accident.
Your AI Marketing Platform Evaluation Starts With A List

Here is the fix. It is free and it takes about ten minutes.
Write your list of jobs down before you know what is doing them.
That is the whole thing. The problem comes from choosing job by job while the test is already running. Take that choice away from yourself and the problem goes with it.
How to do it:
- Write down the real jobs you need done this week. Real ones, in the order they turned up.
- Include the jobs you want to keep for yourself. Especially those. They are the ones you would have quietly held back.
- Once you start, do not add anything and do not drop anything.
- Work down the list in order. Write down how long each job took, and how much you had to fix afterwards.
- At the end, read your notes. Do not go by how fast it felt.
Point 5 is the one people skip. It is also the entire finding of that first study.
The moment you choose job by job, you stop testing the tool. You start testing your own mood that afternoon.
Frequently asked questions
What is the 19 percent study?
A study by a research group called METR, published on 10 July 2025. They timed 16 experienced developers doing 246 real jobs from code they already knew. With AI, the same kind of job took them 19 percent longer.
Does that mean AI makes developers slower?
The authors say no. In that same paper they wrote that their result does not show AI fails to speed up most developers. They also wrote that their 16 people did not stand for most software work.
What did the same researchers change later?
On 24 February 2026 they published an update. They had changed the design and run a bigger round, with 57 developers and more than 800 jobs. They also found a problem with what was going into the test in the first place.
What does holding work back mean?
Some people in the study did not hand in every job they were given. Between a third and a half said they were keeping some back. They did not want to do those jobs without AI. So those jobs were never counted at all.
Does more work actually come out when people use these tools?
At one mid-sized company, yes, and by a lot. Researchers followed 802 developers and 196,212 pieces of work from January 2024 to April 2026. The work coming out reached 2.09 times what it was before.
How should I run an AI marketing platform evaluation on my own work?
Write your list of jobs down before you know what is doing them. Include the ones you want to keep for yourself. Then work down the list in order. Write down what actually happened. Do not go by how fast it felt.
What A Better Test Really Tells You
A test like that will not tell you whether these tools work in general. Nothing you can run in a week will do that.
It tells you something more useful. It tells you how the tool did on a fixed piece of your own real work. That includes the parts you did not want to hand over.
Somebody always chooses. In the famous bad result it was the people taking part, holding jobs back. In the biggest good result it was the people deciding how much to use it. In your own test it is you, on a busy afternoon, with nobody watching.
Writing the list down first does not remove every problem. It removes the one most likely to fool you.
When you tried an AI tool on your own work, which job did you not give it?