The substitution trap
Buy the assistant, measure the hours saved, report fifteen to twenty percent, collect the green check. The ceiling on that story is set the day you define success as doing the existing work faster, and the maximum upside is a fraction of your current labour cost.
A VP buys an AI coding assistant, and six months later the team reports a fifteen to twenty percent gain. Leadership builds a deck, the board is satisfied, and the initiative gets a green check.
Nothing in that sequence is false. The gain is real, the deck is honest, and the board is right to be pleased with what it was shown. But the ceiling was fixed the moment somebody defined success as hours saved on existing tasks.
The claim
Substitution is a fine starting point. Every technology begins there, and I'm not arguing anyone should skip it.
The trap is stopping there. Substitution takes the work as it is, the same steps in the same order with the same people handing it to each other, and makes each step a bit cheaper. Nobody questions the structure of the work, or the sequence of handoffs, or the organisational assumptions sitting underneath both. Price it out. The maximum upside you can capture is a fraction of your current labour cost. That is a cost optimisation story. Not a capabilities one, and the two get funded, measured and defended very differently.
Why it's so seductive
It produces a clean ROI narrative you can put in a quarterly report. Hours before, hours after, multiply by a loaded salary. Done. Nobody has to redraw anything and nobody's authority changes, because the tool slots into the existing process the way a new laptop does and the process gets to claim the improvement.
I had a version of this wrong for a while.
My first instinct was to optimise for deployment: more tools, more coverage, more seats. Coverage is the number that's easy to move and easy to report, and it may not be the number that matters. Adoption isn't a result is the longer argument. The short one is that seats bought measures how much substitution you paid for.
Where the time actually goes
Coding is only fifteen to twenty five percent of shipping software, and the other three quarters is code review, testing, deployment, integration and monitoring. That's where the real friction lives. It's also why individual practitioners report three to ten times gains on specific tasks while most engineering leaders report one to ten percent at the organisation level. It looks like the tools underdelivered. Not quite. They sped up the minority of the work, and leaders who haven't built anything with them don't know which parts got faster and which part is now the constraint. Then there's the chain itself: ticket, queue, planning meeting, sprint, build, review, QA, staging, deploy, measure. At every handoff context bleeds, and the original intent degrades a little more each time the work changes hands or waits in a queue.
A developer picks up a ticket three weeks after a PM wrote it. By then the PM's mental model has moved and the codebase has changed, so the developer fills the gaps with assumptions. The reviewer catches some and misses others. QA tests against a spec that is already stale, and what reaches production is an approximation of the intent, passed through six handoffs and three weeks of decay.
Six handoffs. Three weeks.
Substitution AI speeds up the individual stations and leaves that chain intact, which is why the gains stay modest. You can make every station twice as fast and still lose most of the intent in the gaps between them, because the gaps were never a station.
Invoked and ambient
There are two ways to have AI in the work, and the difference explains more than any benchmark.
In substitution mode you bring context to the AI: open a window, explain the problem, paste the code, describe the constraints. Every interaction starts cold, every single time. The whole cognitive tax falls on the human, who has to know the right question and then evaluate whether the answer fits. In ambient mode the system already holds the context and surfaces what needs your judgment. That split is why the gains so far have been so unequal. The people benefiting most from invoked tools are already high performers (they know which question to ask, and they can evaluate output fast), so the productivity lands with the people who need the least help. Ambient inverts that. The skill that matters moves from prompting well to setting intent clearly, which is a different distribution of who benefits. It's the skill intent architecture is about.
The question that gets you out
The move from treating AI as autocomplete to treating it as a delivery loop turns on one uncomfortable question. It isn't whether AI can do this step.
It's whether a person should be spending time on it at all.
Substitution asks the first one, station by station, and the answer is increasingly yes. The second is where the real cultural friction sits, because it touches identity: engineers who define themselves by their implementation craft are being asked to reframe their value around judgment, problem definition and validation design. That's a conversation, and one most orgs avoid.
Past that point the work stops being about whether engineers use AI and becomes structural, a matter of whether the pipeline itself can be trusted to catch what humans used to catch. The test shifts from "did a human review this" to "do our automated gates catch the failure modes that matter." That is a genuinely different engineering problem, and the one nobody approves the agent's merge picks up.
Productivity was always the wrong return
When Ada started tracking AI adoption, the obvious thing to measure was productivity: how much faster, how much more output. Those numbers are real. The more time I spend inside this, though, the less convinced I am that productivity is the right thing to optimise for.
What I keep noticing, and what we don't know how to measure yet, is that people who work alongside AI seem to think differently after a while. Engineers scope problems at a different altitude. Support teams pattern match across conversations and change the questions they ask.
This is a hypothesis, not a proven playbook. The bet is that where AI is embedded in the actual work rather than the periphery, people's cognitive habits genuinely shift and their bar for an acceptable answer rises. Whether that's real, whether it's durable, whether we're building the conditions for it on purpose, I don't know yet.
It's the bet I'd make.
Substitution can't place it. A tool you visit to do the same task a bit faster doesn't change how anyone thinks, and contact has to be frequent enough and real enough that the rewiring actually happens. People build new habits in the actual work, not by watching demos.
That makes this a leadership problem rather than a tooling one.
What to do with the green check
Keep it. The fifteen to twenty percent is real money and you earned it.
Just don't let it close the initiative. Take one feature and trace it from idea to production. Count the handoffs. For each one, ask what the step is for if the outcome could be reached without it.
The answers won't fit on the slide that got you the green check, which is roughly how you'll know you've left the trap.