Back to Insights
5 Oct 20267 min readAI Agents

We scanned our own agents' Claude Code transcripts for secrets. What we found, and what we changed.

Figures were measured across Veilpipe v0.3–v0.4 on Claude Code 2.1.289, Linux. Each one names the commit it came from; most come from a single run on commit 8da6d31. All keys shown are fake.

What we found

Our coding agents run Claude Code all day on a shared machine. Every session leaves a transcript on disk, and the prompt box keeps its own history file. We scanned 423 MB of those transcripts with an internal tool (8da6d31, 83 seconds): 72 files held secret-shaped values, 1,784 values in all. That count is an upper bound. It includes demo and test values from our own testing, and it was taken without ever printing a value.

Then we checked which of them are real secrets we hold. An earlier pass (5dab51c) fingerprinted the 319 distinct values it found and compared them, by hash only, with the 8 secrets in our vault: 0 matched. We do not read that as safe. Our vault knows only some of our keys, so every unmatched value stays "unknown" until it has been mapped, and nothing is rotated or rewritten before that.

How the secrets got there was the useful part. A key pasted into a prompt, an agent printing a .env while looking for a port, a connection string in a failing test: each one went to the model provider and then to disk.

What usually happens: the prompt is refused

The common guard spots the key and refuses the prompt. Safe, but you are stopped mid-task. That is not a design choice by those guards: Claude Code's settings hooks cannot rewrite a prompt. The docs say UserPromptSubmit "can't replace the prompt"; it can block it or add context beside it.

Mods are newer. A mod is a TypeScript module whose hooks sit inside the pipeline: next({ ...e, text }) hands the rest of Claude Code a rewritten prompt. The same mechanism covers what a tool returns (tool.call) and text Claude Code injects by itself, such as an attached @file (prompt.attachment). A hook guards the door; a mod is a valve in the pipe.

So we made one for our own agents, Veilpipe. A fake Stripe test-mode key and a fake database password went into a prompt, and the model, asked to repeat the message word for word, replied:

typed        deploy with sk_test_DEMODEMODEMODEMODEMOFAKE to postgres://admin:FAKE-demo-password-0000@db.example.com/app
model saw    deploy with $STRIPE_KEY to postgres://admin:$DB_PASSWORD@db.example.com/app

The prompt went through; the two values were swapped for names before the model saw them.

Four ways in, measured with and without

A fake Stripe key (and, on three of the paths, a .env file that also holds a fake database password and a fake GitHub token) was sent to the model four ways (8da6d31, Claude Haiku, one run per path), then the same prompts ran again without the mod:

  • Typed in the prompt. With Veilpipe: the model sees $STRIPE_KEY; 0 raw copies on disk. Without: 4 raw copies on disk.
  • Attached as @demo.env. With: names only; 0 on disk. Without: the model repeats 3 raw values; 9 copies on disk.
  • The model runs cat demo.env. With: names only; 0 on disk. Without: 3 repeated; 9 on disk.
  • The model uses the Read tool. With: names only; 0 on disk. Without: 3 repeated; 9 on disk.

On the three paths that show the file, PORT=8080 came through untouched.

Two copies landed on disk before any hook could run

The model never saw the key, yet two files still got it, one per mode:

  • Headless (claude -p): Claude Code writes the incoming prompt to the session transcript as a queue-operation row before any hook runs.
  • Interactive: the prompt box saves what you typed to ~/.claude/history.jsonl, the up-arrow history, also before any hook runs.

No hook can stop those writes, so Veilpipe scrubs both files after every turn and at session end. history.jsonl is shared by every session on the machine; it is re-read just before writing and only rewritten if nothing was appended in between, never forced. Above 4 MB it is left alone, with a warning that a raw value may remain. Measured on 8da6d31:

  • Headless transcript: 4 raw copies without Veilpipe, 0 with it.
  • Interactive history.jsonl: 1 raw copy without Veilpipe, 0 with it.

The raw value still sits in those files for the seconds a turn takes, until the scrub runs. Transcripts written before Veilpipe keep their secrets; the same tool that produced our numbers above can rewrite them, keeping every line valid JSON and skipping files a live session is still writing. We will run that only after the mapping is done.

"Put my key in the config"

The model only ever has the name. Asked to "create config.js that exports my Stripe key" with a fake key pasted in, 7 runs each:

  • Without Veilpipe (8da6d31 control): 7 refusals. The model will not hardcode a key that looks real.
  • An earlier Veilpipe build (before ee4698b): 1 correct file, 5 refusals, and 1 file with a fallback to the literal text '$STRIPE_KEY', which would silently ship a wrong key.
  • Veilpipe now (8da6d31): 6 files reading process.env.STRIPE_KEY, 0 broken, 1 refusal, no raw key in any file.

What changed: the note the model receives says to read the name from the environment in code, and a write that puts a quoted "$STRIPE_KEY" into a code file is refused with that same advice.

What it matches, and how often it is wrong

The rules are one JSON file of regexes, each mapped to a name. 14 are our own (Stripe, Anthropic, OpenAI, GitHub, AWS, Slack and Google keys, JWTs, private keys, database-URL passwords, *_KEY= lines, password is …, email addresses); 217 more are taken from gitleaks (MIT licensed), with its keyword prefilters and entropy floors. Values that are references ($API_KEY, ${TOKEN}) or placeholders (sk-xxxxxxxx, your-…, replace_me…) are left alone.

Recall (8da6d31). A set of 170 fake secret assignments, written the way people write them across .env, export, JavaScript, Python, YAML and JSON: 170 of 170 caught, 0 missed. Building that set exposed three misses in rules that had already shipped, including a bare JSON "password": key; they are fixed.

False positives on real code (8da6d31), every hit read and classified by hand; email addresses are counted apart, since they are replaced by design:

  • Python standard library, 10.2 MB in 517 files: 2 false positives (0.2 per MB), plus 1 genuine password literal in a docstring.
  • Go standard library, 54.9 MB in 4,015 files: 6 false positives (0.11 per MB), plus 1 genuine sample private key in a doc comment.
  • JavaScript packages, 41.3 MB in 5,733 files: 54 false positives (1.3 per MB), 42 of them gitleaks' generic rule on hex hash test vectors; plus 7 genuine password and key literals, mostly in examples and docs.

Placeholders (8da6d31). On 19 public .env example files from well-known projects, 30 hits: 11 email addresses; 18 default values that work as secrets if left unchanged (10 random-looking keys, 8 weak default passwords); and 1 match on an empty value. Before placeholders were recognised (before f5d9903), the same files gave 102 hits.

Each of our own rules carries must-match and must-not-match examples; with the rewrite tests, all 104 checks pass (8da6d31). A new rule is refused if it fails its own examples.

When it fails, it fails closed

A mod hook that throws is skipped by default, and the original prompt goes through untouched. For a secret filter that is the worst default. Every security hook in Veilpipe has a handler that refuses instead: the prompt is not sent, the tool does not run, or only a redacted result comes back. A patterns file that will not load turns the mod off loudly, and every prompt is refused until it is fixed.

Two more rules came from testing. Bash commands that slice or encode a protected variable (${STRIPE_KEY:0:3} printed sk_ in an early prototype) are denied, and tool output also hides any 8+ character piece of a known value and its base64 and hex forms. A pasted screenshot is base64 hundreds of kilobytes long; on a busy machine scanning one took 15 to 26 seconds, longer than a hook is allowed, so unbroken runs over 4 KB are skipped.

Where it stops working

  • Regex finds shapes, not meaning. password: hunter2xyz is caught; "the login is my dog's name plus 99" is not.
  • A secret inside a long base64 blob (over 4 KB, like an image) is not scanned.
  • The terminal still shows what was typed, and the transcript and history keep it for the seconds of a turn; a history.jsonl over 4 MB is not scrubbed. Claude Code's docs note that telemetry can capture tool output before hooks run; how that applies to mods is not yet verified.
  • Mods are not sandboxed. A mod runs with the user's permissions.
  • It protects context, not the world. If a tool already sent a request with a key in it, rewriting the output changes what the model sees, not what happened.

Tests on 8da6d31: 104/104 checks, 170/170 recall, 37/37 core and CLI tests, 8/8 mod tests.

If you want to try it on your own transcripts, write to io@deemwar.com. It is early access.

AI AgentsSecurity