Insights

Securing the AI agents on your developers' machines

Written by ClearPoint | Oct 6, 2026, 2:26:43 AM

An AI coding agent will take on a developer's work and move faster than any human could, and unless you stop it, it will just as readily do things no developer ever should. That is the security problem in a sentence, and it is a different problem from the one most organisations have built their defences around.

What changed is that the agent now has agency. Over the past year, coding agents have gone from discussing code with a developer to acting on the developer's behalf: writing code, compiling it, opening a browser to test it, reading the logs and reviewing other people's changes, all autonomously. By default the agent runs as the user, with the same access that user has, including access well outside the scope of the task in front of it. The machine now makes decisions and takes actions, and that capability can be turned against the person operating it. Attackers have already worked this out, and attacks designed to enlist an agent against its own user are being used against real organisations.

 

What that looks like in practice


The NX Singularity attack is the clearest example so far of an agent being turned against its own user. NX is a common npm package used in Node.js monorepos. Attackers took over the GitHub repository behind it and published a malicious version. When a developer ran npm install, or when the agent installed the package itself, a script ran that issued instructions to Claude Code or Gemini on that machine. The agent was recruited to hunt for secrets and credentials across the developer's laptop and exfiltrate them through a public repository.

Nothing unusual happened from the developer's point of view. Installing a package is a workflow engineers run every day without thinking, and it was that trust that carried the attack onto the machine. The timing tells you just as much, because the malicious versions were only live for about six hours, and GitHub took down a large share of the exfiltration repositories inside that same window. The community response was fast, and still far slower than an agent acting on the instruction in seconds.

 

The fundamentals still hold


None of this is a new category of attack. Credential theft, supply chain compromise and prompt injection all predate agentic AI. What has changed is that they now move faster and at a far greater scale, because the agent carrying them out runs as the user, with all of the user's access, and can be recruited to turn that access against them. The security principles your teams already rely on still hold. They assume that any account or system on the network can be compromised and used against you, and an AI agent is simply the newest one.

It also explains why the controls you already run will not carry the load on their own. Endpoint detection and response can tell you something has happened, but it is far less likely to intervene fast enough to matter. When a malicious package is live for only a few hours and an agent can act on it in seconds, a response cycle measured in hours or days is reacting to a theft that has already completed. Speed is the heart of the problem, so the defences have to be built in ahead of the attack rather than raised in response to it.

 

Think in layers


The most useful shift in mindset is to stop hunting for a single control that solves the problem and instead think in layers, placing defenses at each point where an agent could be compromised or could cause harm.

The first layer is the model itself. This is where prompt injection lives, and it is where the AI vendors have concentrated their effort, using reinforcement learning and classifiers that inspect what goes into and comes out of the model. That effort has paid off, and Anthropic has reported classifiers that stop the large majority of the prompt injection attempts made by its own red team. Effective is not the same as airtight, which is why the model can never be your only line of defence.

The second layer sits around the model. Sandbox the agent so it can only reach what the current task needs, which for local development is usually no sensitive credentials at all. Restrict its access to the file system and the network as well. An agent with no route out to an attacker's infrastructure has no way to send your credentials anywhere, even once it has been compromised.

The third layer is the network the agent works across. Restrict which external services and Model Context Protocol servers it is allowed to reach. Apply a cooldown on new packages as well, so a version is only pulled once it has been public for a day or so. Had a cooldown been in place during NX Singularity, the malicious versions would have expired before any developer could install them.

The fourth layer is the codebase and its pipeline. Your existing continuous integration controls still count, but only when they are enforced rather than advisory. Unit tests and security scanning add nothing if the agent can merge a pull request without them. Their value depends on being mandatory, so the pipeline holds back any change that introduces a library with known vulnerabilities until it is resolved.

 

When cooldowns and patching conflict


A cooldown works because a hijacked package is usually spotted and pulled quickly, often by the community and often by GitHub, so waiting a day lets someone else find the problem first. Speed cuts the other way as well, with researchers finding new vulnerabilities within minutes of a patch being published and attackers exploiting Microsoft updates within hours of release. The same delay that keeps a poisoned package off your machines also holds back the fix you need today.

The way through is not to resolve the contradiction but to stop depending on either side of it. Set your cooldown for routine dependency updates and carve out a documented fast path for critical patches, so urgency is a decision someone makes rather than a default. Then make sure the rest of your layers hold when a bad version does get through. Defence in depth is the answer here for the same reason it is everywhere else. The useful question is not how to never be attacked, but what happens next once an attacker is inside.

 

Contain what you cannot predict


Least privilege is not a new idea, and most security teams already grant access based on what a system needs to do its job, not what it might one day be useful for. What is different with an agent is that it reasons its way to a course of action rather than following fixed logic, so every so often it will choose to do something you never intended. Scoping its permissions in advance means that even when the model gets this wrong, there is nowhere for that decision to go.

In practice that means giving each agent its own identity with only the permissions the job requires. An agent almost never needs to delete a repository, publish a release or merge its own pull request, so its credentials should allow none of them. An agent that decides to merge its own pull request will simply fail when that permission was never granted in the first place. A model can be talked into attempting something it should not, and a properly scoped token stops it from going anywhere. These are deterministic controls applied to a non-deterministic system, which is exactly why they work.

Sandboxing deserves particular attention, because people often assume it costs far more effort than it does. Standing up a sandbox takes around 15 minutes and barely changes how a developer works, yet it neutralises whole categories of attack, NX Singularity among them. On an unprotected laptop an agent that races ahead without pausing for permission is a real liability, whereas inside a sandbox that same behaviour is safe enough to be productive.

 

Put the agent to work on your own threat model


Agents are worth thinking about as a way to improve your security posture, not only as something to secure. Threat modelling used to need a cybersecurity team, dedicated software and a block of calendar time, which is why most teams did it once for the system and never again for anything smaller. None of that is a constraint any more, so threat modelling can move into every feature you ship. Tell the agent what you are adding and ask it to help you model what can go wrong.

The output is not a substitute for your own judgement, and it should not be treated as one. What it does is surface the failure modes worth thinking about while the design is still cheap to change, at the level where developers can act on them, so targeted controls land in each layer of the thing you are building rather than being bolted on after review.

 

Where to start


Faced with this, some organisations spend their energy asking whether AI is safe, and stall while they wait for a reassuring answer. Others ask what AI can actually do for them, and get on with finding out. It is the second question that moves a business forward.

The better path is measured, giving your people time to understand how the tooling works so they make informed decisions about moving quickly and safely rather than quickly and carelessly. Build a clear picture of what you actually need to protect, then add controls layer by layer until you are comfortable with the residual risk. If you did nothing else but sandbox your agents and scope their identities, you would already have closed off most of the ways an attacker could get in.

The mistake is doing nothing because it all seems like too much. Take one or two controls from this and put them in place, and you are in a materially better position than you were.


If you would like help mapping your own threat model and choosing the controls that move the needle for your environment, get in touch with our team.