You passed the gate. Every core workflow in every function in scope sits at rung 5, and its intent survived a scheduled review. Take the win. Then take the bad news: it isn't finished, and it never will be.
Intent drifts. A market moves, a vendor reprices, a law changes, a competitor ships the thing your roadmap said nobody wanted. The document the agents read still describes last quarter, and they act on it with complete confidence. Nothing throws an error. The machine keeps doing exactly what the page says.
That's why this phase has no gate. You don't leave it. It's the operating model now: drift reviews, escalation paths, a trust log. And your job changes for good. You write intent and you handle escalations. Everything else belongs to the machines and to the people who own them.
- What happens
- Drift reviews, escalation paths, the trust log. The leader's job becomes writing intent and handling escalations.
- The gate
- None. This phase is ongoing.
- What changes for the people doing the work
- New standing roles: intent owners, escalation handlers and drift reviewers.
The job now
Every leader says they want to work on the business rather than in it. Phase 4 takes the choice away.
Writing intent is the work. The outcome, why now, the trade-offs already decided, the decisions still open with a name and a date on them, and what would trigger a revisit. That page is the thing the whole company runs on, and at this point it's the only place your judgment reaches the work without you in the room. Write it vaguely and you've built a machine for executing your ambiguity at a scale no team ever could.
Handling escalations is the other half. Some land on you because they're one-way doors. Most shouldn't. Delegating an outcome means handing over the decisions that come with it, and keeping only the definition of done and how often you check. At rung 5 those two things have names: the intent document is the definition of done, and the drift review is how often you check.
I learned the cost of getting this wrong before any agent was involved. I blocked a colleague by accident, because I didn't have all the details and didn't trust them to move until I'd caught up. The work sat waiting for my understanding. The rule I took from it holds harder now: I cannot be the bottleneck of thinking. Trust and verify. In a company where machines do the hands-on work, a leader who queues every call behind their own attention has rebuilt the interpreter at the top of the org chart.
None of this lets you stop touching the work. You can't write honest intent for work you've never done with these tools, and leaders who stay out of the details lose the calibration that made their intent worth reading.
Drift reviews
Monthly for operational functions, quarterly for strategic ones. Each intent document gets a review on that schedule, and the guide warns you with a date when one is overdue.
A drift review isn't a standing meeting. It's a verdict with a date on it. The owner reads the document against what actually happened since the last review and gives one of three answers: still true, amended, or rewritten. Then it gets stamped. If it takes an hour, the intent was worse than anyone admitted.
The review is the backstop, not the maintenance. The rule that holds day to day is the one from context files: whoever changes the behaviour a document describes updates the document in the same change. A revisit trigger that fires on a Tuesday gets handled on Tuesday. The scheduled review exists for everything that slipped past both.
What the reviewer reads:
- Every escalation since the last review. Each one marks a place where the intent ran out. One is an incident. The second time the same case escalates, it's a pattern, and the document is missing a sentence. Write it.
- The revisit triggers. Did any fire without anyone noticing?
- The open decisions. Any past their date belong to someone by name, today.
- A sample of the output itself. Evaluation practitioners recommend reviewing 10 to 20 traces weekly and repeating a full error analysis every two to four weeks. You don't need that exact rhythm. You do need to look at real work, not just at the document.
Anthropic's guidance for agent context files puts the posture well: review it when things go wrong, prune it regularly, and test a change by observing whether the behaviour actually shifts. An intent document is the same kind of object with a bigger blast radius.
Pick the cadence you want the guide to hold you to on the check-in page. It records the rung changes, flags documents whose review date has passed, and notices when a function stalls.
Escalation paths
Every core workflow already names what escalates. Phase 4 names who it escalates to: a person, and a backup. An escalation path that ends in a team is a queue nobody owns.
Then watch the volume, because it's an instrument. Too many escalations and the intent is too thin, or the agent's autonomy is set more cautiously than the evidence supports. Zero escalations isn't calm. It means the agent isn't flagging what it's unsure of, or the intent is so loose that nothing ever looks out of bounds. The behaviour you want is visible in the data: on complex tasks, Claude pauses to ask for clarification more than twice as often as humans interrupt it, and the same research recommends against per-action approval in favour of making sure people can monitor and step in. That's rung 5 in one sentence. OpenAI's guide to building agents draws the lines the same way: escalate when retries pass a limit, and before irreversible or high-value actions.
Some escalations never go away. Hiring decisions, performance ratings and pay stay with people permanently. Don't treat those as unfinished work.
When something gets through that shouldn't have, resist the reflex to pull the autonomy back. At CircleCI we never walked a permission back. Every incident went into tightening the gate that missed it, never into shrinking the grant. Taking a permission away is an afternoon that looks decisive. The defect is still there afterwards, waiting for the next agent or the next person to find it.
The trust log
The trust log records the calls you argued with, and how they turned out.
You've been writing it since phase 0, whether you noticed or not. If you overruled the guide's pick for your second function, the reason went in. If you appealed a Court ruling because you were sure "security won't approve it" was real at your company, that's an entry. If a timeline verdict said "won't hold" and you told the board it would, that's the biggest entry of all. Each of those is a bet with your name on it.
Each entry needs five things: the call, what you believed instead, the date, what evidence would settle it, and, when it arrives, how it turned out. The last field is the one people skip, and it's the only one that teaches anything.
The point is learning when to push harder. If you were right every time you argued, either you argued too rarely or the log is lying. If you were wrong every time, the next time a ruling comes back false, believe it faster. Either way, you can only see the pattern once it's written down, with dates, where someone else can read it.
A trust log you only fill in when you win is a trophy shelf.
Share it with your exec team. It's the most honest document you'll produce this year, because it's the one where you keep score against yourself.
What changes for the people doing the work
Three standing roles, named in every function: intent owners, escalation handlers, and drift reviewers.
"Standing" means it's in the job description with hours attached, and not a favour someone does on top of their real work. The intent owner is usually the function's best operator, the person who became its owner in phase 2 or 3. The escalation handlers are mostly the people who used to do the hands-on work, and the escalations are where their judgment is worth the most. Handle an escalation well and you've made a call only a person could make. Handle it twice and you owe the intent owner a sentence.
Make the drift reviewer someone other than the author. People can't see what they didn't write. A reviewer from the next function over will ask why the refund threshold is what it is, and the author will realise the answer lives only in their head.
What I don't know
There's a worry I haven't solved, and I'd rather name it than pretend.
Teams are how people make sense of work. "I'm on the platform team, we own billing" tells someone who they are and what they're responsible for. Phase 4, done well, dissolves a lot of that. Fewer handoffs, smaller teams, people working alone with agents and meeting at the escalation. Every step defensible. Each one justified by a metric that went the right way.
The loneliest version of this future is one where everyone is productive and nobody is connected.
I don't know how to prevent that. The cross-function drift reviewer is one small attempt, because it makes two people read each other's work on a schedule. It's a thread, not a fix. The companies that get this right won't be the ones with the best tooling. They'll be the ones that find a reason for people to build together that has nothing to do with efficiency. Watch for the signs in your own org: attrition, or a quiet that's easy to mistake for focus.
The challenge
Your move this week
- For every function in scope, write down three names (the intent owner, the escalation handler with a backup, and a drift reviewer who didn't write the document) plus the date of that document's next review, and write them into that function's intent document by Friday.
- Pull every escalation since the last review in one function, mark each case that escalated twice, and write the missing sentence into that function's intent document before its next review date.
- Open your trust log with every call you've argued with so far (the overruled pick, the appealed ruling, the timeline verdict you rejected), each with its date and the evidence that would settle it, and send it to your CEO.