Ask any executive team how far along they are with AI and you'll get a confident answer. Everyone has licences. Everyone has a channel where people post wins. The number of people using the tools every week is up and to the right.
None of that tells you where the company is. It tells you how much AI people are using, and usage was never the question.
The ladder asks a different one. For each function, where does the human sit? Typing the work? Pasting prompts? Reading every output? Reading a sample? Or writing the intent and taking the calls the machine hands back? Five answers, five rungs. Most companies will find their honest reading is lower than the one they've been telling the board, and that's the point of taking it.
Five rungs, scored per function
1Hands on keyboard
- Human
- Does the work by hand
- Machine
- Nothing, or autocomplete
- To climb
- Tools approved and in daily use
2Interpreter
- Human
- Prompts and pastes, pays the "prompt tax"
- Machine
- Answers when asked
- To climb
- Work moves out of chat and into pipelines
3Reviewer
- Human
- Reviews every output (review is the bottleneck)
- Machine
- Drafts the work, waits for approval
- To climb
- Evals the team trusts enough to stop checking everything
4Sampled review
- Human
- Reviews a sample plus every escalation
- Machine
- Does the work, flags what it's unsure of
- To climb
- Intent written down, owned, reviewed on a cadence
5Intent-setter
- Human
- Writes intent, makes the calls, handles escalations
- Machine
- Does the work end to end, escalates when intent runs out
- The top
- Nothing above it. Hold it.
Read the table by the human column, not the machine column. The machine gets more capable every quarter whether you do anything or not. The human's seat only moves when you move it.
The rung is set by where the human sits, not by how much AI they use.
That rule cuts both ways. A support team that pastes every ticket into a chat window and pastes the answer back is at rung 2, however fluent they are at it and however much time it saves. They are the interpreter. They pay the prompt tax on every exchange, and the moment the person who knows the right prompts goes on leave, the workflow degrades to rung 1. An engineering team that reviews every agent-written diff by hand is at rung 3, even if agents write most of the code. The machine drafts, the human approves, and the queue in front of the human sets the pace.
Rung 3 is where most engineering teams stall, and it's where the numbers start lying to you. Output goes up, the review seat doesn't move, and the queue grows. From the inside it looks like more output, more waiting, and a dashboard that says everything is fine. Review capacity is the new bottleneck, and no amount of generation fixes a bottleneck that sits downstream of it.
Why every function gets its own rung
A company average is a comfortable lie. "We're about a 3" usually means engineering is at 3, support is at 2, finance is at 1, and somebody did arithmetic to make the slide look even. That's a typical first reading, and a fine starting position. It's only useful if you refuse to blend it.
Functions sit at different rungs because they made "done" explicit at different times. Engineering did it years ago: tests, pipelines, a build that goes red. That's why engineering goes first and why it's the proving ground. Finance still closes the month with a person reconciling one spreadsheet against another. Averaging the two hides the fact that finance hasn't started, and hides it from the one person who would have to fix it.
Within a function, the rung is set by the weakest core workflow, not the strongest. If support's triage runs through a pipeline with sampled review but refunds are still a person checking policy by hand, support is at the rung refunds are at. Scoring by the best workflow is how a function declares victory with half its work still done by hand.
Climbing one rung at a time
A function can't skip a rung. That rule sounds conservative. It's the opposite. Every rung builds the thing the next one stands on, and a team that tries to jump from pasting prompts to sampled review has nothing to sample against. It has no pipeline producing consistent output, no evals, no record of what good looks like. It has a hope.
Here's what each climb takes. None of it is a tooling purchase.
From 1 to 2: tools approved and in daily use
The first climb is permission and habit. The company approves the tools, and the people doing the work use them every day, not once a week in a demo.
Two things block this and both are leadership problems. The first is that asking feels risky, so people use unsanctioned tools quietly and you find out during an incident. The second is that leaders talk about the tools and don't use them. Leaders stay hands on now, in the open, where people can see them do it. Training matters more than enthusiasm here: in BCG's 2025 survey of more than 10,600 workers, only 36% of frontline workers felt adequately trained, and people with five or more hours of training were far more likely to become regular users.
Don't measure this rung by adoption and then stop. Adoption isn't a result. It's the ticket to the next climb.
From 2 to 3: work moves out of chat and into pipelines
This is the climb that removes the interpreter. The work stops starting in a chat window and starts in a pipeline: a ticket, a queue, a trigger, a schedule. The machine drafts the output where the output belongs and waits for approval. Nobody copies anything anywhere.
It's the most underrated climb on the ladder, because rung 2 feels productive. People are faster. They like it. But every one of those faster exchanges is a person translating, reviewing, pasting and remembering to run the thing, and the total bill doesn't shrink as you get better at it. It grows with volume. Moving to rung 3 is the first moment the work stops depending on one person's Monday morning.
The test is simple. If the person who runs the workflow is out for a week, does the output still get drafted? If not, you're at rung 2.
From 3 to 4: evals the team trusts enough to stop checking everything
This is the hardest climb, and it's where the real engineering lives. You're asking a team to stop reading every output. They'll only do that when something they trust is reading it first.
That something is an eval suite, and it doesn't have to be big. Anthropic's guidance is that 20 to 50 tasks drawn from real failures is enough early on, written clearly enough that two experts would agree on the verdict. Split capability evals, which start low and tell you what the agent can learn, from regression evals, which should sit near 100% and tell you what it must never break again. Measure consistency, not a lucky best run. Hamel Husain and Shreya Shankar add the part most teams skip: one domain expert as the arbiter of quality, binary pass or fail, and error analysis repeated on a schedule instead of once. Build these before you roll anything out, not after the first incident. An eval written after the incident only encodes the failure you already caught.
Evals tell you what's safe to stop reading. Blast radius tells you what's safe to stop reading first. Sort work by how expensive it is to be wrong, not by how hard it looks. At CircleCI, agents merged to main with no human approval, and the ramp ran in a set order: dependency bumps first, then documentation, then front-end code, with database schema dead last. The gate at every step was a validation suite, not a person. Set those tiers before you need them, and when something gets through, tighten the gate. Don't shrink the grant. Taking a permission back looks decisive and fixes nothing.
What rung 4 changes is how oversight works. The human stops approving each action and starts watching and interrupting. That's what experienced users of agents do anyway: Anthropic's data shows full auto-approve rising from about 20% of sessions for new users to more than 40% after roughly 750 sessions, while interrupt rates rise too. Trust and attention go up together. That's the rung-4 posture. The human reviews a sample and every escalation, and the machine flags what it's unsure of.
From 4 to 5: intent written down, owned, reviewed on a cadence
At rung 4 the machine does the work and a human still decides, case by case, what "good" means. At rung 5 that judgment is written down once, so the machine can act on it and escalate only when it runs out.
Written intent has five parts: the outcome, why it matters now, the trade-offs already decided, the decisions still open, and what would trigger a revisit. It has an owner, a name, and it gets reviewed on a schedule, because intent that nobody maintains drifts, and drifted intent executes at volume. Intent architecture is the longer argument. The short one: for thirty years a person closed the gap between what the company wanted and what it wrote down. When a machine is doing the reading, that gap stays open unless someone closes it on purpose.
This climb isn't done when the document exists. It's done when the document has survived its first scheduled review without the work falling apart.
When a function counts as rung 5
A function counts as rung 5 only when every one of its core workflows is at rung 5. Not most. Not the important ones. Every one.
The bar is strict on purpose. The easiest way to fake an AI-native function is to automate its three clean workflows and quietly leave the messy one to a person. The messy one is usually where the judgment lives, which means it's also where the interpreter hides. Declare victory on the average and that person becomes the permanent translator between the machine and the one workflow nobody wanted to write down.
A function with one workflow still done by hand has an interpreter. It isn't rung 5.
tinycto publishes a starter list of core workflows for each function. You edit it at intake, because your support team isn't everyone's support team. But taking a workflow off the list is a claim, and every claim goes to Constraint Court. "Refunds are too sensitive for agents" gets the same hearing as "our codebase is too legacy." Who owns that belief? What would have to be true for it to go away? What does it cost to keep it?
Two things make the bar less brutal than it sounds.
First, rung 5 doesn't mean the machine decides everything. Every workflow has cases it hands to a person: architecture calls, customer-facing outages, amounts above policy, anything off the legal playbook. A few escalations are permanent by design. Hiring decisions, performance ratings and pay stay human forever. A workflow is at rung 5 when the machine does all the hands-on work and escalates exactly those cases, not when the escalations disappear.
Second, rung 5 needs time you can't compress. A function's intent has to survive at least one full scheduled review, monthly for operational functions and quarterly for strategic ones, before the function counts as AI-native. That rule lives in the timeline check, and it's the one boards find least welcome. Better they hear it from you now.
The challenge
Your move this week
- Take the placement for every function in scope at intake, workflow by workflow, answering yes or no with no "partly". Write the function's rung as its weakest workflow's rung, and put the list on one page with today's date on it.
- Pick the one engineering workflow closest to rung 4, name the domain expert who will judge its output, and have that person write 20 pass or fail eval tasks from real failures in your bug tracker by Friday.
- For every workflow you want to take off a function's core list, file it as a claim in the Court before the end of the week, with the name of the person who holds it. No name, no removal.