Your Secrets Never Leave the Machine: Why We Put a Proxy in Front of Every AI Coding Agent
Every AI coding agent you run needs API keys to do its job — a provider key for the model itself, and usually a handful more for whatever it touches on your behalf: a database connection string, a GitHub token, a cloud credential. All of that ends up sitting in shell history, config files, and — this is the part people don't think about — the request bodies your agent sends upstream, because a prompt that includes "here's my .env, debug this" carries the real secret value straight into a third party's logs.
We made AgentProxy because that risk shouldn't require trusting every tool in your pipeline to behave. It's a local proxy — one Go binary, no cloud dependency for the core loop — that every AI coding agent's traffic routes through before it reaches OpenAI, Anthropic, Google, Azure, or Bedrock.
The problem: your agent is only as careful as everything downstream of it
Point a coding agent at a real codebase and it will, sooner or later, see a real secret — pasted into a prompt for context, read out of a config file it's debugging, echoed back in a stack trace. Once that value is in the request payload, it has left your machine. It's sitting in the provider's logs, in any observability tooling in between, in whatever the agent framework itself persists for replay or caching. You can tell people not to paste secrets into prompts. They will anyway, because the whole point of an agent is that it reads your files for you.
How AgentProxy actually works
AgentProxy sits at 127.0.0.1 between your agent and its upstream provider — an OpenAI-compatible or Anthropic-compatible endpoint the agent already knows how to talk to, so no SDK changes, no new client library, just a different base_url.
Every outbound request is scanned before it leaves the machine. Recognizable secret shapes — API keys, tokens, database connection strings, JWTs, private keys — get swapped for shape-preserving surrogates: values that look and behave like the real thing (same length, same character class, same prefix pattern) so the model can still reason about "an API key is here" without ever seeing the actual bytes. The response comes back through the same proxy, and the surrogate is restored to the original value before it reaches the agent. Nothing about the interaction changes from the agent's point of view — the model just never had the real secret to begin with. It also unwraps encoded secrets nested up to four levels deep (base64, hex, URL-encoded), because agents pass configuration around in encoded blobs constantly, and a masking layer that only catches plaintext misses most of the real traffic. Originals stay in memory for the length of the request and are never written to disk.
That's the masking layer. It sits inside a broader proxy that also does the things a team actually needs once more than one person or more than one provider is involved:
- Multi-provider routing — one endpoint, config-driven upstreams for OpenAI, Anthropic, Google, Azure, and Bedrock, so switching models or providers is a config edit, not a re-integration.
- License enforcement — trial gates, seat limits, and machine binding for teams that need to control who's actually pointed at the proxy, with an Ed25519-signed license format and three binding modes (none for enterprise self-host, first-run bind for a SaaS trial, strict for a locked seat).
- MCP hosting — AgentProxy discovers and proxies Model Context Protocol tools (stdio and streamable-HTTP) to any upstream provider, so an agent gets the same tool access no matter which model is answering.
- Agent skills — it globs for
SKILL.mdfiles and exposes a singleskilltool that loads instructions on demand, the same progressive-disclosure pattern Claude Code itself uses for skills, available to any agent talking through the proxy. - Usage metering — every request and token count is logged per license, so a team knows what it's actually spending before the provider's own bill shows up.
The dashboard is where this becomes visible instead of theoretical: every request line shows the original, what actually left the machine masked, what came back from upstream, and what was restored to the agent — a live audit trail of exactly what a third party saw versus what your agent saw.

Why this over routing everything through a hosted gateway
Hosted LLM gateways (the OneGate/Portkey category) solve routing and metering well, but the traffic still leaves your infrastructure to reach them before it reaches the model provider — which means the gateway itself is now a place your secrets transit and get logged. AgentProxy runs on 127.0.0.1. There's no telemetry call, no third leg in the request path for the core proxying and masking behavior. What you get is closer to what a careful engineer would build for themselves if they had the time: a proxy that assumes prompts will eventually contain secrets, because they always do, and handles it structurally instead of relying on discipline.
Plans, if you're evaluating it for a team
There are three tiers: Solo for a single seat, Team for many seats, and Enterprise, self-hosted on your own infrastructure with Azure AD SSO and custom signing keys. All three tiers get the same core: multi-provider routing, license enforcement, MCP hosting, agent skills, and usage metering — the difference is who hosts it and how many people are on it.
If you're running AI coding agents against a real codebase — which almost always means real secrets somewhere in the loop — this is the kind of infrastructure that should exist before the first leak, not after. Reach out and we'll walk you through setting it up against your own stack.
Found this useful?
Related posts
The Cheapest AI-Code Check You Can Ship Today
We keep [arguing](/insights/stop-reviewing-ai-code-harder-start-checking-it-cheaper) that the fix for AI-written code isn't more trust or more review — it's making the check cheap enough that skipping it makes no sense.…
Your agent graded its own homework. Use a second model as the reviewer.
A coding agent finishes a change and reports that it is done and the tests pass. Often both are true, and the change is still wrong: the tests check what the agent thought the code should do, not what you needed.
A small model checking our docs was a coin flip. Giving it the docs fixed most of that.
A Hacker News thread this week argued that agents don't need memory, they need documentation ([discussion](https://news.ycombinator.com/item?id=49945933)). We had just measured a small version of that question, so here i…
We scanned our own agents' Claude Code transcripts for secrets. What we found, and what we changed.
We scanned our coding agents' Claude Code transcripts for secrets, then changed how prompts reach the model. What we found, measured, and what still fails.