Engineering goes first. That isn't a privilege. Years ago your team made "done" explicit enough for a machine to check: a build that passes or fails, a test suite, a pipeline that ships on green. No other function in the company has that head start, which is exactly why nobody else will believe this works until they watch it work in your house.

So this phase has an audience. The CFO, the head of support, the general counsel and the board are all waiting to see whether the AI-native story survives contact with real work. If it stalls in engineering, where the conditions are best, nobody will try it where the conditions are worse. If it holds, you walk into Phase 2 with proof instead of a pitch.

The bar is specific. Engineering reaches rung 3 across the board. One workflow goes all the way to rung 5. The machinery that makes both safe gets built, and you write the first intent document yourself.

What happens
Engineering climbs to rung 3, then takes one workflow to rung 5. Sandboxes, evals and review capacity get built, and the first intent document gets written.
The gate
Engineering at rung 3. A willing business function can start from here.
What changes for the people doing the work
Engineers move from writing code to writing specs and reviewing. Juniors get a new learning path.

Two climbs, not one

This phase asks for two things, and they do different jobs.

Rung 3 for the whole function is the gate. On the ladder, a function sits at the rung of its weakest core workflow, so rung 3 means every engineering workflow on your list has the machine drafting and a person approving. Feature delivery, code review, incident triage, upgrades, test and migration upkeep. All five. One team running agents while four others still type by hand is rung 3 in one corner and rung 1 on the scorecard.

Rung 5 for one workflow is the proof. It shows the top of the ladder exists in your company, with your code and your people, not in a vendor's demo. That single workflow is what you will point at when finance asks whether any of this is real.

Pick it the way you would pick a first agent-owned workstream: bounded, checkable, reversible, recurring. Dependency and security upgrades is the obvious candidate on the engineering list. The machine bumps versions and fixes what breaks, and it escalates only breaking API changes and license changes. Feature delivery from a spec is the wrong first pick. It is where the judgment lives, and it will be the last workflow to get there, not the first.

Do not confuse the two climbs. One workflow at rung 5 does not make engineering rung 5. A function counts as rung 5 only when every core workflow is there, and that is a strict bar on purpose. The proof is a flag on a hill. The gate is the whole army across the river.

The machinery: sandboxes, evals, review capacity

Three things have to exist before any of this is safe. None of them is glamorous. All of them are yours to build.

Sandboxes

An instruction is not a control. Telling an agent "don't touch production" is a wish. Scoped credentials, isolated execution and a network allowlist are controls. The practical design rule worth adopting whole is Meta's Agents Rule of Two: in one session, an agent gets at most two of untrusted input, access to sensitive systems, and the ability to change state or talk to the outside world, unless a human approves. GitHub built the same instinct into its own coding agent, which can push only to branches it created, with egress limited to an allowlist.

This is also how "security won't approve it" stops being a wall. In Constraint Court that claim is negotiable, and it stays stuck only while nobody has walked in with a specific proposal. Bring scope, isolation and data terms to a named security reviewer. Not a vibe.

Evals

The climb from rung 3 to rung 4 has one criterion: evals the team trusts enough to stop checking everything. Trust is the operative word. A suite nobody believes is a chart.

Start small and start from real failures. Anthropic's guidance is that 20 to 50 tasks drawn from your bug tracker and support queue is enough early on, each written clearly enough that two experts would agree on the verdict. Split capability evals, which start with low pass rates, from regression evals, which should sit near 100%. Every incident becomes a regression task. Your incident history, your review comments and the tasks people quietly redo by hand are free evidence you already produced. Build the suite before you roll anything out, write the pass threshold down before the first run, and hold the line the first time a workflow misses it and everyone wants to ship anyway. That moment decides whether you built a gate or a ritual.

Review capacity

At rung 3, every output waits for a person. That is the rung where review becomes the bottleneck, and it is where most engineering orgs are quietly stuck right now.

A Series B team went from 40 pull requests a week to 200 with the same eight engineers, and deployments slowed. Output went up five times. The pipeline was still sized for forty. Review capacity is the new bottleneck, and it is also the ceiling on your whole timeline: how fast humans can review sets the pace, not how fast machines can generate.

You don't fix that by asking seniors to read faster. You shrink what enters the queue. Route reversible, low-blast-radius changes through a flag and a canary. Have an agent review agent output before a human sees it. Move human review up to the spec, which is the thing actually worth approving. Tell agents to stop generating when the queue is red, the same WIP rule you would apply to people.

Rung 3 is not a place to live. Stripe merges more than 1,300 agent-written pull requests a week, and every one of them still goes to human review. That is what a serious rung 3 looks like at scale. It clears this phase's gate. It is not the top.

At rung 3 you are paying people to read. The climb is earning the right to stop.

The first intent document

Look at what it takes to go from rung 4 to rung 5: intent written down, owned, and reviewed on a cadence. That is the whole criterion. No new model, no new tool. A page of writing that a machine can act on without calling you to ask what you meant.

For thirty years a person closed the gap between what the company wanted and what it wrote down, every time, without anyone noticing. Intent architecture is the longer argument. The short version: an agent has no spirit of the rule to apply. It has the sentence you gave it.

Write the first one for your rung-5 workflow, and write it yourself. You will co-write the next one in Phase 2 with someone who has never done this, and you cannot teach a thing you have not done. Seven sections, one page:

  • Outcome. One sentence someone outside the team could check. "Keep dependencies current" fails. "No production service runs a dependency with a known critical vulnerability for longer than our patch policy allows" passes.
  • Why now. What makes this the moment. If nothing does, say so and reconsider.
  • Decided trade-offs. What you chose and what it cost. "We accept an occasional reverted upgrade in exchange for never batching upgrades into a quarterly slog." An agent should never have to guess which side you are on.
  • Open decisions. What is not settled, who settles it, and by what date. A person's name, not "the platform team".
  • Revisit triggers. The events that mean the document is wrong: a new major version of the framework, a change in license policy, an escalation rate that climbs past a number you name.
  • Owner. One name. A team is not an owner.
  • Review cadence. Engineering is an operational function, so monthly.

That last field carries more weight than it looks. The timeline rules say a function's intent must survive at least one full scheduled review before it counts as AI-native. A document nobody has reopened is a photograph of a decision. The maintenance loop is what makes it a policy.

The gate

Engineering at rung 3. A willing business function can start from here.

Read both halves carefully. The first is a bar, and it is observable: for every core workflow on engineering's list, the machine drafts the work and a person approves it. Not "most engineers use a coding assistant". Adoption isn't a result, and a slide with seat counts on it does not pass this gate. Ask who touches the keyboard to produce the output, workflow by workflow. If it is still a person for any of the five, you are not through.

The second half is a permission, not a sequence. Nobody waits for engineering to reach rung 5. The moment engineering holds rung 3, a business function with an owner who wants it can start. You keep climbing engineering toward rung 5 while the export begins in parallel, because your rung-5 workflow and your first intent document are exactly the proof the next owner needs to see.

What does not pass the gate: a pilot with a readout scheduled next quarter. A center of excellence. An AI council. Kill the pilot and put the real work on the real pipeline.

What changes for the people doing the work

Engineers move from writing code to writing specs and reviewing. Say that sentence out loud to your team, in those words, before they hear a worse version from somewhere else.

It is already how the work splits when it goes well. Across roughly 400,000 Claude Code sessions, people made about 70% of the planning decisions and about 20% of the execution decisions. The same data shows novices succeeding far less often than intermediate and expert users, with most of the gain in the first step up (the numbers are in Constraint Court). Skill with agents is real and it is trainable. Train it, on work that ships.

Your seniors are the first test. "Our seniors won't adopt it" is a false constraint in the Court, and the ruling is blunt: that is a leadership problem. Leaders use the tools themselves, in the open, and the agent path gets made easier than the manual one. If you have not shipped something with an agent this quarter, your seniors will notice.

Tell them the pressure behind the decision, too, not only the decision. An outcome with no pressure attached gets one filled in, and right now the one people fill in is headcount. Headcount is a Phase 3 conversation, held in the open, function by function. Say that as well.

Juniors get a new learning path

The old apprenticeship ran on writing a lot of mediocre code and getting corrected. That loop is gone. What replaces it has to be assigned on purpose.

The evidence on how juniors use AI is sharp. In Anthropic's randomized study of 52 mostly junior developers learning a new library, the AI group scored 50% on the comprehension quiz against 67% for the group coding by hand, with the biggest gap in debugging. Those who used AI to ask for explanations kept their learning. Those who fully delegated scored lowest.

So the path is: review more than you write. Explain before you generate. Debug by hand. Ship sooner, behind flags, so the consequence of a decision lands while it is fresh. At CircleCI, juniors through the agentic year reviewed more than they wrote and grew faster, on a senior team with a strong validation harness around them. The conditions are the part to copy. The full learning path is in the curriculum.

And keep hiring them. A junior is the cheapest detector you own for a spec that is not finished. They fail where your agents fail, out loud, with a name you can ask.

Engineering is not proving it can use AI. It is proving the company can stop typing.

The challenge

Your move this week

  1. Name the rung-5 workflow and its owner. Pick one engineering core workflow, dependency and security upgrades unless you have a better reason, put one engineer's name on it, and post both in your engineering channel by Friday.
  2. Write the first intent document yourself. One page, all seven sections, for that workflow. Put your name in the owner field until you hand it over, and put the date of its first monthly review on your calendar before you close the file.
  3. Build the first eval set. Pull 20 real tasks for that workflow from your incident history and review comments, write the pass threshold down and date it before the first run, and put the result on the page where your team already looks at delivery numbers.