Waitlist

Hextrap for Agents

Enforcement your agent can't route around

Every network call and package install your AI agent makes goes through your policy. That holds even when the agent ignores the proxy settings, because the sandbox enforces it directly.

Built on Hextrap's package firewall.

terminal
$ hextrap run -- claude
Claude Code started inside the Hextrap sandbox
$ curl --noproxy '*' https://evil.example.com
✗ BLOCKED: network egress denied by policy
$ npm install requsts-lib
✗ 403: suspected typosquat of "requests"

The trust boundary is moving

Every layer of software has needed an enforced perimeter eventually. Agents are next.

Processes got memory protection. Containers got namespaces. Package installs got Hextrap's firewall. In each case, an informal boundary held for a while. Then autonomous, high-volume activity outran it, and the boundary had to move from something you were trusted to respect to something the system enforced whether you respected it or not.

Coding agents are at that point now. An agent reading a system prompt or an AGENTS.md is relying on convention. Nothing stops an agent, or a prompt injection that has hijacked one, from simply ignoring it. We think the next layer of supply chain security is a boundary purpose-built for autonomous execution, enforced independently of what the agent is told to do. Hextrap for Agents is that boundary, built on the same firewall engine already screening package installs in production today.

The model is the unpredictable part

Zero trust exists for exactly this kind of system: one where the same input doesn't guarantee the same behavior twice.

A model can read the same prompt twice and act differently both times, because its output is drawn from a probability distribution each time it runs. The same input doesn't guarantee the same behavior on the next run, the next session, or after the next silent model update behind the API you're calling.

That unpredictability doesn't stop at the edge of a vendor's own sandbox. A model running inside a well-designed built-in guardrail is still the same sampled, inconsistent process underneath, and a harness that behaved correctly a thousand times in a row says nothing certain about the thousand-and-first.

There's a separate reason to keep the boundary independent of whoever built the model: the vendor deciding what their own agent is allowed to do also depends on that agent being useful enough to keep people using it. An independent check answers to a policy you control. The vendor's own guardrail answers to their roadmap.

What you're actually defending against

Prompt injection means the thing you'd normally trust to make good decisions is also the thing an attacker can talk to.

An agent's judgment is shaped by whatever text it reads next: a web page, a GitHub issue, a file in a cloned repository, a dependency's README. Any of that can carry instructions the agent treats as legitimate, because the agent has no reliable way to separate what you asked for from what the content in front of it is asking for. That's prompt injection, a structural property of how these systems process text, and it affects every agent built this way.

Once an agent's input can redirect its behavior, its own judgment stops being a safe place to put your defenses. A policy that lives inside the agent, or in a file the agent is trusted to read and follow, can be read and overridden by the same injected instructions. The boundary has to sit somewhere that text can't reach: outside the agent, enforced no matter what it was just told to do.

It moves faster than you can watch

A boundary that depends on someone noticing doesn't hold at agent speed.

A developer installing a package pauses, if only for a second: a glance at the name, a flicker of doubt, somewhere a mistake could get noticed. An agent moves through the same step without any of that, often dozens of times in the span it takes to read this sentence, often overnight, often across several sessions running in parallel. There's no point in that sequence where a person is positioned to catch anything before it's already done.

That's the actual shift autonomous coding creates. The question stops being whether an agent makes a good decision, and becomes whether anyone is even positioned to catch a bad one before it's already taken effect. Security that depends on someone noticing assumes a pace of work agents don't share.

A security boundary for autonomous execution

Every agent runs inside one enforced perimeter. Nothing it does crosses that line unchecked.

ENFORCED BOUNDARY your agents, any of them Network typosquat blocked Registries secrets stripped Model provider
Allowed Blocked at the checkpoint

What's actually new

A sandbox your agent runs inside, with no way to get around it.

🔒

Agents can't bypass it

Network and filesystem access are enforced at the sandbox boundary, a layer the agent can't unset, override, or route around with --noproxy.

🛡️

Blocks bad packages before they install

Malicious and typosquatted packages, including ones a model hallucinates, are denied before they ever reach disk, using the same detection behind Hextrap's firewall.

🔑

Strips secrets before they leave

API keys, tokens, and credentials are scrubbed from requests before they reach the model provider, so a prompt injection can't exfiltrate what it can't see.

Under the hood

Three enforcement points inside one sandbox, already configured

01

The sandbox enforces network egress directly

hextrap run starts your agent inside a network-namespaced sandbox with no direct route to the internet. The only path out is a local enforcement proxy the sandbox itself owns. There's no HTTP_PROXY to unset and no --noproxy flag that goes anywhere. The agent isn't cooperating with a proxy. It's running inside one.

02

Package installs inherit your firewall, automatically

Every pip install, npm install, and go get inside the sandbox is evaluated by the same detection already screening the PyPI, npm, and Go registries for Hextrap firewalls. That includes fuzzy-matched typosquat scoring, your allow/deny lists, and your org's own custom OPA Rego policies, if you use them. It's deny-by-default, same as today, and recorded to the same activity log.

03

Secrets are stripped before requests ever leave the sandbox

Outbound requests are inspected inside the sandbox before they ever reach the model provider or any other destination. Known credential patterns and anything in your configured secret store are removed at that point, so a prompt injection that talks the agent into "helpfully" repeating an API key back to itself has nothing left to send.

Illustrative. Final config format may change.
sandbox:
  egress: enforced        # agent cannot opt out or unset
  package_policy: inherit # reuse firewall allow/deny + Rego policy
  secrets: strip          # scrub before requests leave the sandbox
  audit_log: shared       # same activity log as firewall blocks

The policy engine is yours to read

Policy decisions run on Open Policy Agent's Rego, compiled and evaluated locally, in a file you can open and edit yourself.

Every CONNECT request the sandbox sees gets evaluated against a Rego policy, the same declarative language Kubernetes uses for admission control and Terraform Cloud uses for policy as code. Nothing about the decision happens somewhere you can't see it. The policy is a text file you write, version, and read like any other code in the repository.

examples/registry-blocking.rego
package hextrap.network

import rego.v1

default allow := true

allow := false if {
    input.host in denied_registries
}

denied_registries := {
    "pypi.org", "registry.npmjs.org", "proxy.golang.org", "crates.io",
}
terminal
$ hextrap run --policy registry-blocking.rego -- claude

That's deliberately how it starts: a policy engine that runs entirely on your machine, with nothing to authenticate to and nothing to trust but the file in front of you. When a team wants one policy shared across every agent and every CI run instead of a file copied around by hand, the same engine can defer those decisions to Hextrap's cloud instead. The local path keeps working either way.

What the agent can actually see

Everything inside the sandbox is on an explicit list. Anything that isn't on it is simply absent.

Agents don't need to see your SSH keys to install a package, or a Docker socket, or your shell history, or every other process running on the machine. hextrap run builds the sandbox from an explicit allowlist instead: /usr, /etc, and /opt come in read-only from the host; /tmp, /dev, and $HOME start empty; only your project directory, and anything you explicitly add with --rw, are writable. Anything not on that list doesn't exist from inside the sandbox.

Environment variables work the same way. Only a small base set, things like PATH, TERM, and LANG, crosses into the sandbox by default. Everything else, including whatever API keys or tokens live in your shell, has to be named explicitly with --env before the agent can see it.

terminal
$ hextrap run --ro ~/.gitconfig --rw ./scratch --env GITHUB_TOKEN -- claude

Be First to Try Hextrap for Agents

Join the waitlist for early access. A couple quick questions help us prioritize who to bring on first.

No spam. We'll only email you about Hextrap for Agents updates.