Put a security agent on every service you own
A standing agent that hunts for vulnerabilities in your own systems, every night and on every risky change, with findings going to the engineer who owns the code rather than a security backlog. The setup takes a month. The hard part is making the findings stick to an owner and a clock.
The first run is the one people remember. You point an agent at a service that's been in production for three years, ask it to find the ways a logged-in customer could read somebody else's data, and go and make coffee. When you come back there's a list. Most of it is noise. One item is an endpoint that takes an invoice ID from the URL and never checks who's asking, and it has been sitting there since the service shipped.
Nobody wrote that bug on purpose, and nobody would have found it on the current schedule, which is the annual pen test that covers two services and was scoped by someone who has since left.
I've argued elsewhere that your vulnerabilities are yours: finding them is now the owning engineer's job, not the security team's, because an agent made looking nearly free. This is the mechanical version. What to set up, in what order, and the parts that decide whether it keeps working after the first exciting week.
What you're building
One agent per owned system, running on a schedule and on risky changes, with a standing brief that says what to look for. Its findings land in the owning team's queue with a severity, a line number and evidence. A confirmed finding gets an owner by name, a deadline set by severity, and a regression test when it's fixed.
That's all of it. The agent is the easy part. The failure mode I'd expect, and have watched happen with ordinary scanners for a decade, is a tool that runs faithfully every night into a dashboard nobody opens.
So the one number to organise around is time from finding to fix, per team. Not how many findings. A count mostly measures how hard you're looking.
Ground rules before the first run
Two of them, and they aren't optional.
Point it only at systems you own and are authorised to test. Your code, your infrastructure, your staging environment. Not a vendor's API, not a partner's endpoint, not the SaaS tool your team happens to use, even when the agent would be happy to try. If a finding leads outside your boundary, it becomes a report to that vendor, written by a person.
Then give it the least access that does the job. Read access to the repository and the infrastructure config. Test accounts in a staging environment that holds no real customer data. No production credentials, no write access to anything but its own findings, no ability to open network paths the service itself couldn't. A security agent is an agent like any other, and its blast radius is a design decision you make up front. The version where your vulnerability scanner becomes the breach is not hypothetical, it's what happens when someone gives it an admin token to save an afternoon.
Write both rules into the brief itself, where the agent reads them, and into the one-page note where your team reads them.
Time: an hour, mostly getting the staging test accounts set up.
Step one: list what you own
Before anything runs, get the list of services, jobs, buckets and infrastructure stacks, each with an owner's name next to it. A team's name doesn't count. A person.
You'll find two things. Some systems have nobody, usually the ones built by someone who left, or the scheduled job that "just runs." And some have two owners who each think it's the other one's. Both are findings, and they're the most important ones you'll get this month, because an agent finding a hole in an unowned system produces a ticket that nobody picks up. Ownership is the atomic unit, and it's the thing this whole setup runs on.
Don't start step two until every system on the list has a name.
Time: half a day, longer if the answer to "who owns this" is a meeting.
Step two: write the standing brief
The brief is a file in the repository that tells the agent what to hunt for, what's out of scope, and how to report. Keep it short and specific. A vague brief ("find security issues") produces vague findings, and vague findings get closed as noise by week two.
Start with the classes of problem that actually get companies your size breached, and nothing clever:
Broken access control. For every endpoint that takes an identifier, does it check that the caller is allowed to see that record? This is the single most valuable question on the list, and the one a human reviewer misses most often, because the code looks fine and the check is simply absent.
Injection. Follow every piece of user input to wherever it ends up: a query, a shell command, a template, a file path, a URL the server then fetches.
Authentication gaps. Routes that skip the middleware, tokens that don't expire, password reset flows that leak whether an account exists.
Secrets. Credentials in the code, in config, in the logs, in the history of the repository.
Dependencies. Packages with a published, exploitable vulnerability that your code actually calls. The reachability part matters. A vulnerable function you never invoke is a low, and pretending otherwise is how the list fills with noise.
Infrastructure. Public buckets, security groups open to the world, roles with wildcard permissions, anything that assumes the network is the boundary.
Then the part most briefs leave out, which is what the agent should ignore. Style findings. Theoretical issues with no path from outside. The internal admin tool that's only reachable over the VPN (say so, with the reason, so a future reader can check whether it's still true). Every false positive you reject in step four should end up here as a line.
The brief is context, and context is maintenance, not a build. Give it an owner, and look at its git log once a quarter. A brief nobody has touched in six months is describing a system that no longer exists.
Time: two hours for the first version. It'll be wrong in useful ways.
Step three: three triggers
Nightly, a full pass over the service. This is the baseline and it catches drift: the dependency that got a new advisory, the config that changed without a code change.
On every change that touches the sensitive paths: authentication, authorisation, payments, anything that reads or writes personal data. Mark those paths once (a list of directories in the brief is enough) and run the agent on the change before it merges. It's a reviewer that only asks security questions and never gets tired of asking them. If you already run a machine first pass before any human review, this is the same slot with a different brief.
On a new advisory for something you depend on. When a vulnerability is published, the question is whether your code reaches it, and an agent can answer that in minutes where a human spends an afternoon grepping.
Leave out the fourth trigger people always want to add, which is "on demand, whenever someone's worried." It's fine to run it that way, but it can't be the plan. Worry is not a schedule.
Time: an afternoon to wire into CI and a scheduler.
Step four: findings an owner can act on in five minutes
Every finding carries the same five things. A severity. The file and line. Evidence: the request that demonstrates it, or the path from input to sink, or the policy that grants too much. A proposed fix. And the name of the owner from step one.
Findings go to the owning team's normal queue, next to their other work. Not a separate security dashboard, not a spreadsheet, not a channel. A separate place is a place nobody visits.
The owner triages within a working day: confirmed, or rejected with a reason. That reason is the most valuable output of the whole system in the first month, because it goes straight back into the brief. Every rejection that says "this endpoint is internal only" or "the ORM already constrains that field" is a line that stops the same noise tomorrow night.
Watch the confirmation rate for the first few weeks. If fewer than one finding in five is real, engineers stop reading them, and they're right to. That's not a reason to switch the agent off. It's a sign the brief is too loose, and tightening it is the owner's job, the same way tuning a noisy alert is the owner's job.
Time: ongoing, and front-loaded. Expect the first two weeks to take real attention.
Step five: the fix clock
A confirmed finding gets a deadline set by its severity, and the deadline is public inside the team. The numbers I'd start with:
Critical, meaning exploitable from outside with real customer impact: treat it as an incident. Same day. Somebody runs it like any other incident, and it gets a postmortem.
High: a week.
Medium: this sprint, or thirty days if you don't work in sprints.
Low: the backlog, with the honest expectation that some of these will never be fixed and the brief should stop reporting them.
Tune the numbers to your risk, but have numbers. A finding with no deadline is a finding with a deadline of never.
Every confirmed finding, when fixed, gets a regression test. Not just for the specific endpoint but for the class where you can manage it: a test that walks every route and asserts each one checks ownership is worth more than fifty individual fixes. It's the same principle as post-incident evals being regression tests. A hole you closed without a test is a hole you'll reopen during the next refactor.
Time: whatever the fixes cost. That's the actual work, and it was always there. You just couldn't see it before.
Step six: the two-tenant test
Once the static passes are settled, add one dynamic check, because it finds the class of bug that does the most damage.
Create two customer accounts in staging, in two separate tenants. Give the agent both sets of credentials and one instruction: using only what account A is allowed to do, try to read, change or delete anything belonging to account B. Let it walk the API, the web app and the export paths.
This single test covers most of what a breach at a B2B company actually looks like. It's also a test almost nobody runs, because it's tedious for a person, and tedium is exactly what an agent doesn't mind.
Run it nightly against staging. When it finds something, it's almost always critical.
Time: a day to set up the accounts and seed data, then nothing.
The four numbers
Time from finding to fix, by severity, per team. The main one. It tells you whether findings stick to owners.
Confirmation rate. Confirmed findings over total. Too low means the brief is loose. Very high in a mature service can mean the agent's only finding the easy things, which is worth a look.
Coverage. The share of owned systems that had a full pass in the last seven days. Anything under all of them means something fell off the list, usually a new service nobody added.
Recurrence. Findings in a class you've already fixed once. A class that keeps coming back is a missing regression test or a missing paved-road default, and it's a better use of the security team's time than any individual finding.
Put all four on one page, per team, and look at it monthly. Not weekly. Weekly turns it into a leaderboard, and a leaderboard for security findings gets gamed by looking less hard.
What the security team does now
They stop being the people who find the holes. They become the people who make finding them cheap and fixing them easy.
Concretely: they own the base brief every team starts from, and review team briefs that have drifted. They set the severity definitions. They build the paved-road defaults that make whole classes of finding disappear (an authorisation middleware that denies by default, secrets injected at runtime, infrastructure templates that are closed unless opened). They go deep, by hand, on the few areas where specialist judgment still earns its cost: payments, identity, cryptography, anything regulated. And they look at the recurrence number, because that.s where one change can retire a whole class of finding.
What they shouldn't do is become the triage layer. The moment findings route through the security team on their way to the owner, you've rebuilt the queue this was supposed to remove, and a gate built to feel in control isn't the same thing as a control.
What this won't catch
Business logic that's wrong in ways only a person who understands the product would notice. A refund flow that can be run twice isn't a code smell, it's a product bug with a security consequence, and the agent won't know refunds are supposed to happen once unless you tell it.
Chains. The attack that combines a low in one service with a medium in another and a misconfiguration in a third. Agents are getting better at this, but it's still where a skilled human tester earns the fee. Keep an external pen test for your most sensitive systems, once a year, and give the testers your findings history so they don't spend a week rediscovering what you already fixed.
People. Phishing, a contractor account nobody disabled, a laptop left in a taxi. The short list of controls (identity, secrets, patching, restored backups, access review) still comes first, and an agent reading your code does nothing for any of them.
And compliance. This will make the audit easier, because you'll have evidence of continuous testing. It won't make you compliant and it isn't a substitute for doing the audit if a deal needs it.
A thirty-day rollout
Week one: one service. Pick the one that handles the most customer data and has the most confident owner. Do steps one through four for that service alone. Expect noise, and spend the week turning rejections into lines in the brief.
Week two: add the fix clock and the two-tenant test to that same service. By the end of the week you should have at least one confirmed finding fixed, with a test. If you don't, look hard at whether the brief is looking in the right places before you conclude the service is clean.
Week three: every service owned by that team. The brief should now be good enough to copy. Watch the confirmation rate on services the brief wasn't written for.
Week four: the team presents what it found, what it fixed and how long it took to the rest of engineering, in the open, and the other teams start. Resist mandating it. A team that watched a peer find a cross-tenant leak in their first week needs no mandate, and a team that was told to do it will write a brief designed to find nothing.
Where to start
Tonight. Pick the service you know best, give an agent read access to the repository, and ask it one question: which endpoints take an identifier and don't check that the caller is allowed to see it?
Read the answer tomorrow morning, before standup. Whatever's on that list is yours, and it has been for a while.