
AI Security, Made Simple: Defence in Depth for Agents
Treat an AI agent like an eager developer: capable, fast, well-meaning, and guaranteed to occasionally overreach. Give it its own machine, its own accounts, and the least privilege it needs.
Treat an AI agent like an eager developer: capable, fast, well-meaning, and guaranteed to occasionally overreach. Give it its own machine, its own accounts, and the least privilege it needs, and its mistakes stay contained. This post is the first in a practical series on AI security, and it starts where security should always start: protecting your own data, your machine, and your systems.
If you have read The Panic Dividend of AI Security, you know my position on the vendor noise: most of what changed is how visible the old problems became. This series is the other half of that argument. If I am telling you not to buy the panic, I owe you a description of what I actually do instead. Everything here is published as working configuration in my public ai-toolkit repository, in the security folder, so you can read the exact rules rather than take my word for it.
The Eager Developer
Here is the mental model that drives every control in this post.
When you run a coding agent on your machine, you have hired a developer who works at unreasonable speed, never gets tired, genuinely wants to help, and has read more code than any human alive. You have also hired a developer who will occasionally misread your intent, take a shortcut you did not ask for, delete something it believed was in the way, or follow instructions that arrived in a file it happened to read. Not because it is hostile. Because it is eager, and eagerness plus authority is how accidents happen.
You would not give a new developer, however talented, the keys to production on day one. You would not let them commit straight to main, run arbitrary commands on your personal laptop, or hold your AWS root credentials. Not as an insult. As standard practice, because everyone makes mistakes and good organisations design for that.
So the question is never “do I trust the AI?”. The question is the same one you ask about any new team member: what is the blast radius when they get something wrong?
Two failure modes matter:
- Overreach. The agent does more than you intended: edits files outside the task, reconfigures its own guard rails, reads credentials because they were readable.
- Manipulation. The agent follows instructions you never gave, planted in a README, a dependency, an issue comment, or a web page it fetched. This is prompt injection, and no model is immune to it.
Both have the same mitigation. You do not fix the developer. You shape the environment so that mistakes cannot reach anything you cannot afford to lose.
Layer 1: A Separate Context
The strongest control is the simplest one: run the agent somewhere that is not your daily machine.
A spare Mac Mini. A Docker container. A throwaway VM. A cloud development environment. Anything where the honest answer to “what happens if the agent trashes this?” is “I rebuild it in ten minutes and lose nothing”. On a disposable machine there are no personal files to read, no browser sessions to hijack, no SSH keys to exfiltrate. The shell can be as porous as it likes; there is nothing behind it.
This is the same reasoning as giving a contractor a locked-down laptop instead of your own. It is not about distrust. It is about making the failure cheap.
If a second machine sounds like a luxury, you already have this layer available for free: run Claude Code on the web. Each task gets a fresh, isolated virtual machine that Anthropic manages. Your repository is cloned in, your GitHub token stays outside the sandbox (the session receives scoped, short-lived credentials instead), and the whole environment is disposable by design. It is the simplest possible version of a separate context: nothing to buy, nothing to patch, and the honest answer to “what happens if the agent trashes this?” is “nothing, it was never my machine”. A dev container on your own hardware achieves the same isolation with more setup; the web sandbox is the low-effort default.
Isolation also solves the problem that permission rules cannot. A deny rule that blocks the agent from reading ~/.ssh can be sidestepped by a shell command the rule did not anticipate: cat, head, xxd, a one-line Python script. You can play whack-a-mole with deny patterns, or you can run the agent on a machine where ~/.ssh is empty. Containment beats enumeration, every time.
Layer 2: Dedicated Accounts
The second layer applies whether or not you achieved the first: the agent gets its own identity.
Its own GitHub account, or at minimum its own fine-grained access token scoped to the repositories it works on. Its own email address. Its own service account in each system it touches. Never your personal login, never a shared team credential.
This buys you three things:
- Containment. If the account is compromised or misused, you revoke one identity. Your own access, and everyone else’s, is untouched.
- Attribution. Every commit, ticket update, and API call made by the agent is visibly made by the agent. When something odd happens at 2 am, you know immediately whether it was you or the machine.
- Revocation as routine. Rotating or killing an agent’s account is a non-event. Rotating your own credentials, everywhere, is a bad week.
Again the analogy holds. A new developer gets their own accounts on day one. Nobody says “just use my login for now”, and the reason they do not is exactly the reason that applies here.
Layer 3: Scoped Credentials, Always
The first two layers are context-dependent. You may not have a spare machine; a client environment may not allow a dedicated account. The third layer applies regardless, because it costs almost nothing and works everywhere.
Whatever identity the agent uses, cut its permissions to the task in front of it:
- In AWS, an IAM role that can read the one bucket and deploy the one stack, not
AdministratorAccessbecause it was quicker. - In Azure, a separate token or service principal scoped to the resource group, not your own
az loginsession that spans every subscription you can touch. - In GitHub, a fine-grained token limited to the target repositories with the minimum scopes, not a classic token with
repoacross your whole account. - In Kubernetes, its own service account bound to its own namespace-scoped roles.
Least privilege is old advice, and that is the point. Nothing about AI changes the principle; what changes is the volume of actions. An eager developer with too much access makes one expensive mistake a quarter. An agent making hundreds of calls an hour with the same excess access just runs the same odds far more often. Scoped credentials cap the cost of any single mistake, and they keep capping it even when the outer layers fail: when the container has a hole, when the account was reused, when a deny rule missed a case. That is why this layer is not optional.
Defence in Depth, Made Simple
Put together, the model is three questions, asked in order:
- Can it run in its own context? Separate machine, container, or VM. If yes, most of the risk is gone before any rule is written.
- Can it have its own accounts? Dedicated identities in every system it touches. Containment, attribution, cheap revocation.
- What is the least access it needs today? Scoped roles and tokens, whatever the answer to the first two. This layer always applies.
No layer is sufficient alone. Deny lists are porous. Containers have escapes. Accounts get over-provisioned over time. Defence in depth does not require any layer to be perfect; it requires the failure of one layer to land on the next one instead of on you.
The ai-toolkit security folder holds the concrete versions of all of this for Claude Code: a machine-level baseline that denies the agent read and write access to credential stores and stops it editing its own settings, and a repo-level fragment for disposable machines where the repo, not the host, is the thing worth protecting. Each README explains not just the rules but the gaps in them, because knowing where a control fails is part of the control.
What I Actually Do
Theory is cheap, so here is my own setup, layer by layer.
The context is a Mac Mini server. All agent work runs on a dedicated Mac Mini, not on the laptop I live on. It is reachable only over Tailscale, a private mesh network between my devices, so nothing is exposed to the internet and I can reach the agents from anywhere. On that machine, every customer gets their own macOS user account: separate home directory, separate keychain, separate credentials. Client contexts never mix, and wiping one account touches nothing else.
The work is driven remotely, against written plans. I steer sessions through Claude Code’s remote control from wherever I am. Communication about project and planning work goes through a shared Obsidian file per project, a plain markdown document both I and the agents read and write. Every plan is a markdown document in the repository, so the full detail survives the session that produced it. And the work itself always comes from a project management system, Monday or Jira depending on the customer, so an agent is never acting on a verbal instruction that nobody can audit later.
The accounts implement separation of duties. The agents run under two identities. The first is the engineer: its goal is to complete the work. The second is QA: its goal is to verify that the goal plan was met. The QA account reviews pull requests, and it owns approval for specific change types only, such as version bumps and markdown changes. It cannot manipulate or change its own policies or rule systems. When CI fails, a security finding appears, or quality slips, QA does not fix it; it steers the engineer back to the plan.
The premise is deliberately simple: one agent’s goal is to complete the work, the other’s is to confirm the goal was met as planned. That only works because the plan is written down first. The pair of skills that produce and execute those plans, agree a Definition of Done as a Test Plan with a human, then drive it autonomously to a validated result, are now shared in the toolkit as goal plan and goal execute.
If that sounds familiar, it should. Maker and checker, four-eyes approval, a reviewer who cannot rewrite the rules they enforce: none of it is new. It is how we have always organised humans we trust but expect to make mistakes. The controls transfer to agents almost unchanged, which is the quiet good news of this whole series.
Where the Series Goes Next
This post stayed deliberately at the first level: you, your machine, your accounts. The further we get, the deeper we dive. Coming up: how an agent can quietly weaken its own guard rails and how to make every bypass require a human; what changes when agents run unattended in CI; and what agent-driven development means for the software supply chain. Same principle every time, applied one layer further out: expect mistakes, contain them, and make the failure of one control the problem of the next control rather than your problem.

25 years building secure digital solutions for regulated and modern teams.