4 min readai strategyagentic engineeringdelivery

Kill the AI pilot

A pilot is what you run when you can't roll back. If you can, skip the committee: hand the tool to one team, keep a real record, and decide in three weeks.

By JP LeBlanc

A branching pilot-committee structure with many gates dead-ending, beside a single direct line from one team through one decision gate to full rollout.

A pilot is what you run when you can't roll back the decision. Most AI-tooling rollouts aren't that kind of decision. You can cancel a seat in an afternoon, turn off an integration, or send a team back to its old workflow before lunch. If the decision is reversible, a pilot is the wrong instrument. Give the tool to one team, watch four numbers for three weeks, and decide.

Ask a VP of Engineering how they're rolling out a coding assistant and watch the reflex arrive. A steering committee. Representation from three teams so nobody feels left out. A success-criteria document circulated for comment. A readout to leadership sometime next quarter. Eight weeks of scaffolding, built to answer a question a rollback button already answers by itself.

What a pilot actually produces

Strip away the deck and a pilot produces one artifact: a report. Somebody writes it, somebody edits it for tone, and it lands on a shared drive where three people read it before the initiative gets folded into "ongoing evaluation."

The report is also, almost always, written by people who already knew which answer they were supposed to reach. Whoever championed the tool sits on the steering committee. Whoever was skeptical sits there too, having agreed to participate specifically so the "objective process" would carry their name. Neither one changes their mind because a slide deck told them to. The pilot doesn't produce information so much as cover: a documented process everyone can point to, regardless of which way the decision eventually goes.

The obvious fix is a faster pilot. Shorter timeline, smaller committee, a decision in two weeks instead of eight. That's still a pilot. The fix isn't speeding up the ritual. It's asking whether the ritual was ever the right instrument for a decision you can already undo.

The question underneath it

Reversibility decides the process, not seniority, not consensus, not how nervous the room feels. The same question runs one-way and two-way doors everywhere else in the org: how expensive is it to be wrong, and how fast can you undo it.

An AI coding tool rolled out to one team is not a one-way door. Nobody's data gets migrated. No customer sees it. The worst case is a few weeks of a team using a mediocre tool, which they were going to notice and complain about regardless of which committee approved it. Compare that to what's actually irreversible in most rollouts: a multi-year vendor contract, a workflow so embedded that ripping it out costs more than it saved. That's worth deliberation. Eight engineers running a coding agent for three weeks is not.

A pilot substitutes for a rollback plan that already exists, and it costs more than the rollback plan would have.

One team, three weeks, a real record

The alternative isn't reckless. It's smaller.

Pick a team that actually wants to try the tool, not the one that got volunteered because nobody else raised a hand. Pick a stretch of comparable work and write down the numbers evals-not-vibes already argues for: deploys to production, cycle time on a consistent slice of tickets, defect rate split by whether the change was agent-touched. Do it before anyone touches the tool. That one step, done ahead of rollout instead of retrofitted three months later, is the difference between an answer and an opinion.

The record-keeping is duller than it sounds, and it's the first thing to slip under deadline pressure, which is exactly when skipping it costs the most. Nobody remembers to log cycle time on week two of a real deadline. Somebody has to be told it's their job.

Then let the team actually use the tool, in production, on real work, for the full three weeks. Not a sandbox repo. Not a synthetic ticket built to flatter a demo. A pilot fails to produce a decision because it's designed to be safe to ignore, and a sandboxed evaluation is safe to ignore by construction.

Real work answers faster than any committee would believe. A designer on our team rebuilt two pages carrying ninety percent of the platform's traffic in three days with an agent, against a two-month estimate. That is what a real problem, a person, and a tool produce. Three days of production work answered a question a steering committee would have spent a quarter debating.

Deciding, on a date, with a name attached

At the end of three weeks, one person decides. Not the committee. The entire reason to skip the pilot process is to make this the reversible call one person can make and unmake, the same way they'd make any other decision that doesn't require a memo and six weeks of alignment.

Three honest outcomes, and all three are fine. Roll it out further, because the numbers moved. Kill it, because they didn't, and the team goes back to its old workflow having lost three weeks, which is cheap. Or run it again with a different team, because the first result was genuinely ambiguous, which is a real outcome and not a failure to reach one.

What doesn't happen is "ongoing evaluation." That phrase is where decisions go to avoid being made, and anyone who's sat in one of those meetings already knows it on contact.

Where a pilot is actually right

Not never. If the tool touches customer data in a way you can't unwind, if the vendor wants a multi-year commitment before you've used it for a week, if adopting it means retraining a workflow that took two years to build the first time, that's a one-way door. And the agentic prerequisite applies before anything else does anyway: check whether your test suite still tells the truth before you check whether the tool is any good, because a bad answer to the first question makes the second one unanswerable regardless of process.

For everything short of that, the pilot isn't caution. It's an unmade decision dressed as one, run by people who already agree on the answer, taking longer to reach it than the tool would have taken to prove itself wrong.

Questions this answers

What's wrong with running an AI pilot before rolling out a coding tool?
Nothing is wrong with the concept. What's wrong is reaching for it when you already have a rollback plan, which most teams do. A pilot is the process you run when a decision can't be undone. Rolling a reversible tool out to one team while you measure it is cheaper and faster than convening a committee to study the question.
How long should an AI tool pilot run before you decide?
Three weeks is usually enough, provided you picked real, comparable tickets and measured deploys, cycle time, and defect rate before and after. A longer pilot doesn't produce better data. It gives more people time to build a political position around the answer they already expect.
Who should decide whether to roll out an AI coding tool company-wide?
One named person, using a number instead of an adjective. The whole point of skipping the pilot ritual is to treat this as the reversible decision it actually is, one person can make and unmake, not a consensus exercise that outlasts the tool's own trial period.