Ask any company whether it is AI-native and it will show you a number. Seats bought. Weekly active users. The share of code a model wrote last month. Every one of those numbers is going up, everywhere, and almost none of them tell you anything.
Use is universal now. Google's 2025 DORA report found 90% of the technologists it surveyed use AI at work, for a median of two hours a day. If using AI made a company AI-native, nearly every company would qualify, and the word would be useless.
So this guide uses a definition you can check by watching, not by asking. It is strict on purpose. Most companies that call themselves AI-native fail it, and the ones that pass it look different from the inside.
The definition
An AI-native company is one where machines do most of the hands-on-keyboard work. Humans think, make judgment calls and handle escalations. Machines do the work.
Read it slowly, because every clause excludes something.
"Most of the hands-on-keyboard work" excludes the company where AI helps people type faster. Help is not the work. If a person is still producing the output, with a model offering suggestions, the person is doing the work.
"Humans think, make judgment calls and handle escalations" is not a consolation prize. It is the job, and it is harder than the one it replaces. Deciding what should be true, choosing between three options that all look reasonable, and catching the case the rules never imagined: none of that got cheap. Implementation got cheap. Judgment didn't.
"Machines do the work" means end to end, on a pipeline, without a person carrying the output from one tool to the next. Not a chat window. A loop that plans, acts and checks its own work, inside limits someone set.
Notice what the definition does not mention: headcount, a model, a vendor, a percentage of anything. It describes where people sit relative to the work. That is the only thing that changes when a company actually crosses over.
The test anyone can observe
For any workflow, ask one question. Who touches the keyboard to produce the output?
If it is a person, the workflow is not AI-native yet, however much AI they use along the way.
Run it on a few workflows and it gets uncomfortable quickly. A support agent asks a model for a draft reply, edits it, and pastes it into the ticket. Person. An engineer describes a function in a chat window, copies the answer into the editor, fixes the imports and opens the PR. Person. A controller asks an assistant to explain a variance and then types the commentary into the monthly pack. Person, again, and these are all teams that would describe themselves as heavy AI users.
Now the other side. A ticket arrives, an agent reads the spec, writes the change and the tests, runs them, and opens a PR that a human reviews. The machine produced the output. The human is still checking every one, which puts that workflow on rung 3 of the ladder, not at the top. But the keyboard moved.
Three rules make the test useful.
Watch, don't ask. People are bad at reporting this, and the evidence is specific. In METR's randomized study, experienced open-source developers using AI tools were 19% slower on their own repositories, while believing they had been about 20% faster. Self-reports measure how the work felt. Sit next to someone doing the workflow and see whose hands are on the keys.
Test workflows, not functions. "Support is AI-native" is not a claim anyone can check. "Refunds under the policy limit are issued without a person typing anything" is. A function is only as far along as its weakest core workflow, which is why the ladder scores each one separately.
Count the carrying. A person who moves output from one tool into another is touching the keyboard, even if they wrote none of it. Copying is work. So is remembering to run the thing.
Substitution or redesign
The keyboard test has a consequence most AI programs are built to avoid. You cannot pass it by buying a better tool.
Most companies start with substitution: take the work as it is, the same steps in the same order with the same people handing it to each other, and make each step a little cheaper. That is a fine place to start. It is a trap to stop there, because the ceiling is set the day you define success as hours saved. The most you can ever capture is a fraction of your current labour cost.
The ceiling is low for a structural reason. Coding is fifteen to twenty-five percent of shipping software. Review, testing, deployment, integration and monitoring are the rest. Speed up the station you can see and the work piles up at the next one. One Series B team went from forty pull requests a week to two hundred with the same eight engineers, and its deployments slowed down. Output went up five times. The review step was still sized for forty.
When electricity replaced steam, factories bolted the new motor into the old engine's spot and kept the shafts and belts. It took decades before anyone redesigned the floor around the flow of work. Most companies are in that gap now: a faster motor, the same belts.
The keyboard test forces the redesign, because the only way to take the person's hands off the keyboard is to change the workflow itself. Somebody has to decide what the step was for, whether it still needs to exist, and what the machine should do when it is unsure. Those are questions about structure and authority, not about tools. That is why they get avoided, and why this guide keeps asking them.
You don't become AI-native by making the old workflow faster. You become AI-native by changing who does it.
What stalls most companies
Almost every company climbs the first rung without trouble. Tools get approved, people start using them, the adoption chart goes up. Then it stops, at rung 2, and stays there.
Rung 2 is the interpreter. A person sits in the middle of the workflow, turning a messy situation into a prompt, reading what comes back, fixing it, and pasting it wherever it needs to go. The model answers when asked. The person does everything else.
It feels like progress, and in a narrow sense it is. The individual step got faster. But look at what the interpreter pays on every exchange, and keep paying, for as long as the workflow runs. The prompt tax has six line items: translating the situation into words a model can act on, reviewing the output, handing it off to the next tool, remembering to run it at all, being the only one who knows which steps need a close look, and absorbing the decay when the world the workflow assumed quietly changes. Most teams price the first two. The same person pays all six.
Anthropic's own study of its engineers shows how deep the interpreter habit runs, even among the people building the models. They reported being about 50% more productive with Claude, and said they could fully delegate only 0 to 20% of their work. The rest still passed through a person.
Rung 2 is sticky for a second reason. The people who get the most out of it are your best people. They know which question to ask and they can judge an answer fast. So the interpreter role is held by exactly the people you would least like to lose to copy-and-paste, and they are the people least likely to complain, because the tool made them look even better.
The interpreter is what happens when the only place intent lives is inside someone's head. The model cannot ask what you meant, so a person stands next to it and answers on its behalf, every time.
Written intent
Removing the interpreter takes one thing above all: intent written down, precisely enough that a machine can act on it without asking.
For most of the history of organisations, nobody needed this. A leader stated a direction loosely, people underneath filled in the specifics, got it slightly wrong, got corrected, and converged. It worked because whoever received the instruction could always ask what you meant. An agent cannot. It has the sentence you gave it, and it will do exactly what the sentence says.
An intent document answers five questions.
The outcome. What is different when this works, stated so that two reasonable people would agree on whether it happened. "Faster support" fails that test. A named measure, on a named queue, passes it.
Why it matters now. Conditions change. A document that doesn't say why it exists gets treated as timeless, and a timeless document is a stale one nobody dares touch.
The trade-offs already decided. What you chose and what you gave up for it. This is the section people skip, because naming what you are giving up is a small, specific decision someone can later point at. It is also the section that stops an agent from guessing.
The decisions still open. What is not settled, who settles it and by when. Leave these unwritten and whoever hits the gap first invents an answer, which becomes policy by accident.
What would trigger a revisit. The events that mean the document is wrong: a metric crossing a line, a vendor change, a new law. "When things change" is not a trigger.
Add an owner, one person's name, and a review cadence, and that is the page. The guide gives you a template for each function in your takeaways.
It is one page for a reason. Anthropic's guidance on context engineering is to find the smallest set of high-signal tokens that gets the outcome. A fourteen-page style guide is not intent. It is noise with a table of contents.
Writing it will be harder than it looks, and not for lack of discipline. Vagueness is often load-bearing. "We're prioritising reliability and velocity" lets two factions who disagree both feel represented. Writing down the trade-off ends that truce. Expect the argument, and have it on paper, where it is checkable, rather than in production, where it is a postmortem.
Then keep it alive. Intent written once drifts, quietly, until the machine is executing last quarter's priorities with complete confidence. So the written intent needs a maintenance loop, with an owner who reviews it on a schedule and a record of when they did. Anthropic's advice for its own coding agent's instructions applies to every intent document you write: review it when things go wrong, prune it regularly, and check whether behaviour actually changes. The guide makes this a hard rule. A function doesn't count as rung 5 until its intent has survived at least one scheduled review.
The two jobs
Everything in this guide serves two jobs, and you need both.
Job one: move each function's hands-on work to machines. Workflow by workflow, rung by rung, until the keyboard test comes back "machine" for every core workflow in every function in scope.
Job two: build the intent and escalation plumbing that makes job one safe. Written intent with an owner. Evals the team trusts. A defined point where the machine stops and hands the case to a person. A trust log of the calls you argued with and how they turned out. At rung 5, the machine does the work end to end and escalates when intent runs out. Job two is what makes "when intent runs out" something the machine can detect.
Do job one without job two and you get speed without judgment. Klarna moved its support front line to AI, and in 2025 its CEO said publicly that the company went too far, focused too much on cost, and got lower quality. It started hiring human agents again, with AI still on the front line.
Do job two without job one and you get a governance program with nothing to govern: councils, policies, review boards, and the work still done by hand. That version is more common, because it feels responsible and it threatens nobody.
The two jobs move together, and they share a ceiling. Review capacity, how fast people can check what machines produce, sets the pace of both. Generation is never the constraint for long. That is why the ladder is built around where the human sits, and why the phases that follow ration review before they ration anything else.
The challenge
Your move this week
- Run the keyboard test by watching. Pick three workflows before Friday: one in engineering and two outside it. Sit with the person who does each one while they do it. Write down, for each, who touched the keyboard to produce the output, and every point where output was copied from one tool to another.
- Price one interpreter. Take the recurring AI workflow your team runs most often, list the six line items of the prompt tax against it, and put a name next to each one. If the same name appears six times, that person is your interpreter. Tell them you noticed.
- Write one intent document. Choose one engineering workflow. Write its outcome, why now, decided trade-offs, open decisions with an owner and a date, and revisit triggers, on one page, with one owner and a review date. Commit it to the repository where the agent will read it, and put the review date in the owner's calendar.