Last month I wrote about why the model should run on hardware you own. That solved the personal side of the privacy equation. But it leaves the harder question for anyone working in IT — and in government IT in particular: fine, the agent runs in-house, but what is it allowed to touch?

I work in hosting and middleware support at Shared Services Canada. We build and run server estates on behalf of other departments, which means the environments are regulated, change-controlled, and firmly off-limits to anything that isn't a credentialed human. That constraint is the job, not an obstacle to it: agentic AI is fast, and it is also capable of exactly the kind of confident mistakes you never want near a production system. So I've spent the last few weeks running an agentic AI where it's allowed to live — in our lab, where it can make every mistake it needs to make and none of them cost anyone anything. The thing that surprised me isn't what the agent can do there. It's that the lab turns out to be the whole point, not the waiting room.

The obvious way: work backwards from production

When people imagine "AI in operations," they picture plugging the agent into the live estate. The slightly more cautious version keeps prod untouched but asks the AI to model what's already there — give it a live server, have it figure out the build. Either way, you're working backwards from a manual setup, and that reproduction is less accurate than it feels. The build decisions live in human muscle memory and half-remembered ticket threads. Every hand-applied fix since go-live is drift the agent inherits without context. And there's the risk side: an agent pointed at a live server is an agent that can touch a live server. Agentic tools make confident, plausible mistakes — wrong assumption, wrong command, wrong host — and in a regulated estate a mistake there isn't a lesson learned, it's an incident report.

Here's what that costs in practice. I've watched a four-server estate get stood up by hand — four identical build sessions — and then watched the same changes get applied to all four servers by hand again, one at a time, months later. The knowledge of how that estate works lives in people, not artifacts. An agent reverse-engineering it from the outside starts from a lossy copy of a thing nobody fully wrote down.

The way that compounds: make the agent build the world first

Flip the order. Hand the agent the requirements, not the running servers, and have it build the entire environment as a mock-up in the lab: an identical Rocky Linux VM estate to your exact spec — packages, config, services, the whole stack, end to end. Then have it document its own decisions and decide, from its own testing, the best way to stand the stack up.

For a human, building and testing a full mock estate is a hell of a lot of work — it's the part everyone skips. For an agent, it's just work. It's also somewhere for the agent to make every mistake it's going to make anyway, at the price of a re-run. And everything downstream gets better because of it:

  • You can try it before prod exists. The mock serves pages. Open them in your browser, click through them, iterate in the lab while iterating is still free.
  • The deployment falls out of the testing. When the mock behaves the way you want, the byproduct is a documented deployment artifact that produced that behavior — not a wish list you have to hand-translate into servers.
  • Replication stops being an event. Four servers, forty servers — it's the same artifact applied N times, instead of a build session you have to live through again.
WORKING BACKWARDS the agent models a manual prod build built by hand, ×4 same fixes, applied ×4 agent guesses the build ! Build knowledge lives in human heads ! Every change repeated by hand, per server ! Reverse-engineered mock drifts from prod ! The agent never watched anything get built ! Nothing is replicable; it all has to be redone BUILD THE MOCK FIRST the agent builds it, tests it, then ships it requirements agent builds the full mock in lab end-to-end tests pass √ The agent wrote the build — it knows every layer √ Try it in your browser while it’s still just a lab √ Future issue? Reproduce it in an identical mock √ The deployment artifact is the byproduct √ Apply it to 4 servers or 40 — same artifact human gate PRODUCTION — a human moves it across
The same agent, two workflows. One inherits drift; one produces it on purpose, in the lab, where it's cheap.

When it breaks, debugging flips inside-out

This is the part I didn't expect. Later issues are easier to diagnose and fix precisely because the agent built the thing — it is fully aware of how the stack is put together, even though it can't go into the live server. It can't touch prod and it doesn't need to. When something misbehaves, it spins up an identical mock to test against, reproduces the problem, tries the fix where mistakes are free, and hands over the diff. I review it. The lab gets the fix first; production only ever sees finished work.

The human gate is where the responsibility lives

Agentic AI makes mistakes — not occasionally, constantly; that's how it learns, same as us. The honest question isn't how to stop that, it's where the mistakes are allowed to happen. In the lab, a mistake costs a re-run. In production, it costs an outage, an incident report, and a conversation with the department whose service you host. That's the actual reason the agent lives in the lab and never near production, with a human in the loop to move whatever it produced onto prod servers. That's the shape I run now, inside the SSC lab, and it doesn't feel like a leash. The gate isn't slowing the work down, because the mock-first workflow is fast on this side of it — all the risk-taking happens where there's nothing to break.

This is all new, so I'm not complaining — just sharing the observation. Maybe someday it will be responsible to let an agent fix your production server directly. Not yet. And the teams that get good at building and testing mocks now are the ones who'll be ready to let agents closer to production later — with receipts, not hope.

The takeaway

Using agentic AI on systems you're accountable for isn't about access at all. The question isn't "how do we give it production access?" and it isn't even "how do we stop it from touching anything?" It's: "can the agent rebuild this entire environment, from requirements alone, in the lab?" If yes, every mistake the agent can make has already happened somewhere cheap, and production becomes a copy operation with a human pressing enter — every future change, diagnosis, fix, replication across a fleet, is a lab exercise first. That's what protecting production from agentic AI actually looks like: not banning it, giving it somewhere to be wrong safely. If no, the agent doesn't know your environment. It knows a story about it. And stories drift.

— Kris