# AI Coding Agents Slowed Developers Down 19% — How Growth Companies Measure AI's Real Impact Before Scaling

> Experienced developers were 19% slower with AI tools (METR, 2025). Here's how growth companies measure real AI productivity gains before scaling up.

- Published: 2026-08-20
- Author: Master Mind
- Canonical: https://aimaster.fi/en/artikkelit/ai-koodausagentit-hidastivat-kehittajia-19-nain-kasvuyritys-mittaa-tekoalyn-hyod

Experienced software developers were 19% slower when using AI coding agents than without them. That's the finding from a randomized controlled trial run by AI research organization METR in July 2025 (METR, 2025). The result contradicts both the developers' own perception and expert forecasts — and it's exactly why growth companies shouldn't scale AI adoption before measuring its real impact.

This isn't a case against AI. It's a case for measurement. AI productivity gains vary by task, by user experience, and by the quality bar the work has to meet — and the only way to know what's happening in your own company is to measure it.

## What did the METR study actually show?

METR recruited 16 experienced open-source developers working on codebases they knew well. Researchers randomly assigned 246 real issues into two groups: some allowed AI tools (mainly Cursor Pro with Claude 3.5/3.7 Sonnet), others didn't. The result: tasks done with AI tools took 19% longer. Developers had expected a 24% speedup beforehand, and even after the experiment they still believed they'd been 20% faster — despite recorded time proving otherwise (METR, 2025).

Researchers tested 20 possible explanations for the slowdown and found five likely contributors: limited familiarity with the specific tool, time spent verifying AI-generated code, and high quality bars in the codebases that the agent didn't automatically meet. This doesn't mean AI is useless in software development generally — the study covered one specific setting: experienced developers on familiar, high-quality codebases.

## Why does this matter for your company?

Growth company leaders constantly hear pitches promising an automatic productivity leap from AI. The METR result is a useful reminder: the gain isn't automatic, and it can even be negative — without anyone noticing, because the experience feels faster than reality actually is. The same risk applies beyond coding: reporting, customer communication, analysis work.

The risk isn't using AI. The risk is rolling it out based on an assumed benefit calculated on slides, instead of a measured, real benefit. Slide-deck consulting sells a promise. A working system produces a measurable result — or reveals what needs fixing before you scale.

## How does a growth company measure AI's real impact?

The first step isn't buying a tool — it's mapping where AI produces measurable value in your specific business. At AIMASTER, this phase is productized as [Master Plan](https://aimaster.fi/en/tuotteet/master-plan), an AI strategy sprint that maps where AI creates the most value for your company, measured in euros. The mapping happens before tool rollout, not after.

In practice, measuring means three things: record the baseline (time, error rate, or throughput before AI), compare the same metric with AI on real work — not a demo — and trust recorded time over a felt sense of speed. METR's result shows exactly why: self-assessment and reality can diverge significantly.

| Measurement method | What it tells you | Risk |
| --- | --- | --- |
| Developer's own speed estimate | Subjective experience | Can be off by 20+ percentage points (METR, 2025) |
| Recorded time on a real task | Actual throughput | Requires a control group and real tasks |
| Demo or pilot in an isolated setting | Tool capability under ideal conditions | Doesn't reflect real production impact |

## What if the agent's output needs heavy review?

One of METR's findings was that part of the slowdown came from reviewing the agent's output — developers had to verify the code met style, testing, and documentation standards. This is a common pattern in other knowledge work too: AI produces a fast first draft, but review eats the time savings if the process isn't designed around it.

The fix isn't dropping AI — it's designing the process so review happens where it adds value, not where it erases the gain. [Master Mind](https://aimaster.fi/en/tuotteet/master-mind) agents are built into your business process with human checkpoints placed where they matter. That requires clean data and system connections, which is what [Master Layer](https://aimaster.fi/en/tuotteet/master-layer) — a data foundation layer that connects your existing systems securely for AI use — is built to solve.

## How does this connect to Anthropic's agent design patterns?

We've previously covered why the most complex agent setup isn't always the best choice — see [Workflow or Agent? Anthropic's Five Design Patterns](https://aimaster.fi/en/artikkelit/workflow-vai-agentti-anthropicin-viisi-suunnittelumallia-kertovat-miksi-monimutk). METR's finding reinforces the same principle from another angle: a simple, measured, bounded task where AI actually proves faster beats a broad, unmeasured rollout every time.

## What does AI adoption cost if you measure the benefit first?

Measurement isn't an extra cost — it's part of the rollout. The sprint model makes this predictable: development proceeds in 3-day cycles, and measurement is built into the first sprint. Custom AI solutions are delivered through an agile sprint model: one sprint is 3 development days. That way you see the result in euros and hours before deciding to scale.

## FAQ

Why were experienced developers slower with AI tools?

In METR's controlled trial (2025), 16 experienced open-source developers completed 246 real issues on codebases they knew well. Tasks done with AI tools took 19% longer, driven by review time and high quality bars the tool didn't automatically meet. Developers believed they'd been faster, even though recorded time showed otherwise.

Does this mean AI isn't worth adopting?

No. The study covered one narrow setting: experienced developers on familiar, high-quality codebases. In other contexts — less experienced users, unfamiliar systems, routine tasks — the gain can be clearly positive. The deciding factor is measurement: know what's happening in your specific situation.

How does a growth company reliably measure AI's benefit?

Record baseline time and quality before AI, compare the same metric on real production work with AI, and trust recorded data over self-assessment. A Master Plan mapping does this systematically before any broad rollout.

Can AI slow down other knowledge work besides coding?

Yes, by the same logic: if AI-produced work requires significant review or correction, total time can exceed the baseline. This applies to reporting, content production, and analysis work — not just code.

How do we know our AI project is actually delivering results?

The clearest approach is starting with a measurable pilot where both the baseline and the outcome are expressed in euros or hours. A [free Master Mind analysis](https://aimaster.fi/en/analyysi) gives you a concrete picture of where AI creates measurable value in your specific business.

## Frequently asked questions

### Why were experienced developers slower with AI tools?

In METR's controlled trial (2025), 16 experienced open-source developers completed 246 real issues on codebases they knew well. Tasks done with AI tools took 19% longer, driven by review time and high quality bars the tool didn't automatically meet. Developers believed they'd been faster, even though recorded time showed otherwise.

### Does this mean AI isn't worth adopting?

No. The study covered one narrow setting: experienced developers on familiar, high-quality codebases. In other contexts — less experienced users, unfamiliar systems, routine tasks — the gain can be clearly positive. The deciding factor is measurement: know what's happening in your specific situation.

### How does a growth company reliably measure AI's benefit?

Record baseline time and quality before AI, compare the same metric on real production work with AI, and trust recorded data over self-assessment. A Master Plan mapping does this systematically before any broad rollout.

### Can AI slow down other knowledge work besides coding?

Yes, by the same logic: if AI-produced work requires significant review or correction, total time can exceed the baseline. This applies to reporting, content production, and analysis work — not just code.

### How do we know our AI project is actually delivering results?

The clearest approach is starting with a measurable pilot where both the baseline and the outcome are expressed in euros or hours. A free Master Mind analysis gives you a concrete picture of where AI creates measurable value in your specific business.
