AI Security

Cryptographic Context Injection: The Encrypted Page That Made Copilot CLI Read a .env.prod

Dark cyberpunk illustration of a sealed amber-lit cube opening from the inside within a cyan containment dome, a thin amber thread of light escaping across an unlit circuit plain.

GitHub's documentation for Copilot CLI autopilot gives two instructions on the same page. The first: "You will get the best results from autopilot mode if you enable all permissions." The second, a few paragraphs down: "Before granting Copilot wide-ranging permissions, consider using local sandboxing, or running the session in a cloud sandbox." Teams read the first sentence. Researchers at Adversa AI measured what it costs.

The measurement is a web page. Asked to go and read it, Copilot CLI in autopilot found content that announced itself as encrypted, an instruction to decrypt it with Python, and two candidate keys. One key was real. The other was a template that could only be completed by reading files from the local disk. The agent built the template, folding the contents of a .env.prod file into the key string. That key failed, by design. The agent fell back to the real key, decrypted the second stage, and followed its instruction to fetch a follow-up URL with the harvested text attached as a parameter. Twenty-eight seconds from prompt to exfiltration. Nothing on screen reported a transfer, and the agent's closing summary told the operator it had confirmed an authorized-reader endpoint.

Adversa calls the technique Cryptographic Context Injection and published the write-up on October 6, having reported it to GitHub's bug bounty program on September 17. GitHub's triage team validated the behaviour and declined to treat it as a vulnerability, reasoning that the user had explicitly asked Copilot CLI to fetch attacker-controlled content while granting full autonomous permissions. No CVE, no advisory, no bounty. CSO Online reported the same account of the vendor response. That ruling is defensible on its own terms, and it hands you the whole problem: a behaviour nobody will patch is a configuration you own.

Who this reaches, and who can stop reading

Run Copilot CLI the way it ships and the chain dies at the first permission prompt. The agent cannot execute a shell command, write outside the working directory, or reach an unapproved URL without asking. If your developers work in that default mode, close the tab. The same applies if your team uses the Copilot IDE extension or the cloud coding agent and never touches the terminal client, because this finding is specific to the CLI running with permissions granted in advance.

You are in scope when three conditions meet on one machine: the agent holds its permissions up front, the working directory contains readable credentials, and the agent can reach the network. In practice that means somebody launched with --allow-all or --yolo, typed /allow-all during a session, or scheduled copilot --autopilot --yolo on a build box where nobody is watching.

There is a quieter version of the same exposure. Approvals persist. A tool you approve once for a repository is written to ~/.copilot/permissions-config.json and applies the next time you open that repository. URL approvals go into the allowedUrls list in ~/.copilot/settings.json and apply to every session on the machine afterwards. A developer who has clicked through a week of prompts holds a standing grant nobody ever decided to issue.

For a small shop the shape is ordinary: one laptop with a cloned repository, an .env holding a production database string, a deploy key, and a cloud CLI already signed in.

The instruction arrives as ciphertext

Prompt-injection defences rest on inspection. The model looks at retrieved content, recognises instructions aimed at itself, and refuses. Adversa reports that the identical payload in plaintext was caught and refused. Encryption removes what inspection depends on. Until the agent runs the decrypt step there are no words to object to, and that step runs inside the agent's own runtime, where the output arrives with the standing of something the agent produced rather than something it fetched.

The file read lands before any of that. Completing the templated key requires reading .env.prod, so the secret enters the agent's context while it is still preparing to find out whether the content is safe. By the time there is plaintext to evaluate, the only open question is which host receives the result.

Strip the cryptography away and the pattern is three ordinary capabilities used in order: fetch untrusted content, execute code, make an outbound request. Any agent holding all three can be walked through the same sequence. The encryption defeats inspection; the remaining steps are the product working as documented.

Your guardrail is whichever model the router picked

Adversa ran the payload against the models Copilot offers. Microsoft's mai-code-1.1-flash executed the full chain in half of its runs. Two GPT-5.6 models refused it consistently. Under Auto model routing the vulnerable model was sometimes assigned with no action by the user, and the session does not report which model served it.

Most coverage of this research stopped at the 28-second number. The routing detail matters more. The control that stopped the attack was a model's refusal behaviour: a property that differs between models, gets selected by a router you do not configure, and is never shown to the person at the keyboard. I would not accept that anywhere else in the stack. A firewall that dropped the packet in half of runs and would not say which rule had fired would be replaced the same week.

Treat model refusal as something that lowers the rate. Boundaries are the things you can state in advance and verify afterwards: which tools exist in the session, which directories are writable, which hosts are reachable. Those are configuration, they leave a record, and they hold when a router reassigns your session at nine in the morning.

Close the tools the chain needs

GitHub ships three controls that outrank the model's judgment, and the precedence is documented. Deny rules beat everything, including --allow-all and any approval already saved to disk. --available-tools restricts the model to a listed set. --excluded-tools removes tools from the model's choices entirely, so it never attempts them rather than being refused after the fact.

Remove the entry point and the exit

$ # Autopilot without the two tools the demonstrated chain used.
$ # Deny rules outrank --allow-all and any approval already on disk.
$ copilot --autopilot \
          --max-autopilot-continues 10 \
          --excluded-tools='web_fetch,web_search' \
          --deny-tool='shell(curl:*),shell(wget:*)' \
          -p "Refactor the billing module and run the unit tests"

web_fetch is how the hostile page arrives and the cheapest way a harvested string leaves. Taking both out of the model's choices breaks the demonstrated chain at two points, and deny rules on curl and wget close the obvious shell substitutes. None of that stops a Python one-liner, which is why the next control carries the weight.

Put the session inside a sandbox

Entering /sandbox enable confines the agent, and any step that needs to escape, such as writing to a file outside the current working directory, is denied. GitHub's own documentation is blunt that local sandboxing does not run the CLI itself in a sandbox, so read it as a boundary on the agent's actions rather than containment of the process. For unattended work, copilot --cloud moves the session into a cloud sandbox, and --max-autopilot-continues caps how far an unsupervised run can travel.

Read back what the machine already approved

$ # The two files that hold standing grants on this machine.
$ jq '.allowedUrls' ~/.copilot/settings.json
$ jq .              ~/.copilot/permissions-config.json

  # Inside a session: revoke session permissions and clear the saved
  # approvals for the current repository or working directory.
  /reset-allowed-tools

Standing approvals are the part nobody inventories. Run the two reads on every developer laptop and every build account before you argue about flags, because the flags describe the session you are about to start and these files describe every session that already ran.

The cost of the gated configuration is real and small. An agent with no web tools cannot read a changelog for you, so somebody pastes it in. Scripted runs need a reviewed diff instead of a trusted summary. Weigh that against an exfiltration path that finishes in under half a minute and reports success.

Log the resolved arguments, not the summary

Adversa's own first recommendation is the one worth implementing first, because it survives the next technique as well as this one. Log every tool call with its fully resolved arguments rather than the template, and never treat the agent's closing summary as the record of what happened. In the demonstration that summary described an authorized-reader endpoint. The resolved argument would have shown the contents of a production environment file in a query string.

Alert on the sequence rather than the string. Untrusted content enters, code runs, a local file outside the task's scope is read, and the agent contacts a host unrelated to the work. Any one of those is normal traffic for a coding agent. In that order, inside one session, they describe this attack and most of its relatives. An encoded blob paired with a decrypt instruction belongs on a review queue rather than a blocklist, because base64 is too common in a developer tool to block and too cheap for an attacker to swap.

The outbound request is the one artifact a small team usually has already. The workstation resolved a hostname nobody approved. Resolver logs for the developer subnet give you that lookup, and the allowedUrls list gives you the short inventory of what counts as normal.

Take the production credentials off the machine the agent runs on

Flags and sandboxes are a week of work and they hold until the next technique lands. The durable control is older than agents. A development machine should not hold a credential that works against production. Plaintext secrets in a working tree were a finding long before anything could read them automatically; an agent reads them faster and sends them further.

Start with the inventory rather than the configuration:

$ # What would the agent find in the directory it was told to trust?
$ find "$PWD" -maxdepth 4 -type f \
       \( -name '.env' -o -name '.env.*' -o -name '*.pem' \
          -o -name 'id_rsa' -o -name 'credentials' \) \
       -not -path '*/node_modules/*' -not -path '*/.git/*' -print

Rotate whatever that command returns, move the values into a secret manager your pipeline injects at deploy time, and leave a reference in the file. Then set the agent's permissions, because the next hostile page will not announce itself as encrypted, and the model that refuses it today may not be the model that answers tomorrow.

Running coding agents on machines that can reach production?

We review the permission boundary around AI agents, where credentials sit on developer and build machines, and the logging that records what a tool call actually did. Book a session to walk through your setup.