AI Security

CVE-2026-90970: GitLab's AI Gateway Template Sandbox Broke Twice in Eight Months

Dark cyberpunk illustration of two identical amber-lit glass containment chambers standing in an unlit machine hall, each with the same panel shattered at the same point, and a thin cyan beam passing straight through both breaches.

In February, GitLab fixed a CVSS 9.9 bug in its AI Gateway: a crafted Duo Agent Platform flow definition could break out of template expansion and run code on the gateway host. On October 2 it fixed the same thing again. CVE-2026-90970 is CVSS 9.9, in the same component, reached by the same kind of input, and classified under the same weakness. The two CVE records carry the identical CVSS vector string, character for character.

The relevance verdict first, because most readers can stop after this paragraph. This is yours only if you run the GitLab Self-Hosted AI Gateway - the model-gateway container - on your own infrastructure. GitLab states that it has already patched the gateways it operates, so GitLab.com, GitLab Dedicated, and Self-Managed instances that point at a GitLab-hosted gateway have no action to take. The precondition on the rest is an authenticated user with Duo Agent Platform access, so if you self-host the gateway and have never turned Duo on, you still want the patch, but nobody can reach the bug today.

There is a second verdict the advisory does not give you. Upgrading to 19.4.1 removes the sandbox escape. It leaves intact a project Maintainer's ability to write a flow whose toolset includes run_command, which the flow specification documents as intended behavior. The upgrade is this week's task. Deciding who may author a flow is the standing control, and it outlives the patch.

Two advisories, one weakness class

Read the two records next to each other and the repetition is exact, not approximate:

  • CVE-2026-1868, published February 9, 2026. "AI Gateway was vulnerable to insecure template expansion of user supplied data via crafted Duo Agent Platform Flow definitions." CWE-1336. CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H, base 9.9. Fixed in 18.6.2, 18.7.1, and 18.8.1.
  • CVE-2026-90970, published October 2, 2026. An authenticated user with Duo Agent Platform access could "escape the prompt template sandbox via a specially crafted flow configuration, resulting in arbitrary command execution on the AI Gateway." CWE-1336. CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H, base 9.9. Fixed in 19.2.4, 19.3.2, and 19.4.1, credited to HackerOne reporter invisiblemeerkat.

CWE-1336 is "Improper Neutralization of Special Elements Used in a Template Engine", and MITRE lists exactly two mitigations for it: choose a template engine with a sandbox, and use that sandbox. GitLab did both, in February, and the sandbox held for about eight months.

Now do the version arithmetic, because this is the part that bites operators who thought they were already covered. February's fixes landed at 18.6.2, 18.7.1, and 18.8.1. October's affected range runs from 18.1.6 up to 19.2.4. A gateway that was patched in February and has sat at 18.8.1 since then is inside the new vulnerable range. Applying the first fix bought eight months, not immunity.

One honest caveat on urgency: neither CVE is in the CISA Known Exploited Vulnerabilities catalog, and as of October 2 nobody has reported exploitation. Treat this as a scheduled-this-week upgrade on a privileged internal service, not a drop-everything perimeter emergency.

Work out whether you are running one at all

In my experience the hard part of a self-hosted AI advisory is never the patch. It is that the component was stood up by a platform team during an AI pilot, deployed from a Helm chart or a single docker run, and never entered the asset inventory that the security team actually reads. The gateway is a container with its own release tags, patched on its own cadence, independent of the GitLab version you track. Five checks settle it:

# 1. Is a self-hosted gateway running on this host?
docker ps | grep -i model-gateway

# 2. Which image tag is deployed - the tag IS the version you compare
docker ps --no-trunc | grep -o 'ai-assist/model-gateway[^ ]*'

# 3. Kubernetes: same image, any namespace
kubectl get pods --all-namespaces -o yaml \
  | grep -o 'ai-assist/model-gateway[^"]*' | sort -u

# 4. GitLab's own diagnostic, run from the Rails node
sudo gitlab-rake gitlab:duo:verify_self_hosted_setup

# 5. Anything answering on the documented gateway ports
ss -lntp | grep -E ':5052|:50052'

The image you are looking for is registry.gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/model-gateway, with a FIPS variant published under model-gateway/self-hosted-fips. Ports 5052 (HTTP) and 50052 (gRPC) are the documented listeners. If none of those five checks returns anything, close the tab; you are a GitLab-hosted customer and GitLab has already done the work.

What else lives on that host

The scope-change flag in the vector - S:C - is the part worth dwelling on, because it is what separates this from a bug that merely crashes a service. Command execution on the gateway puts an attacker in the same process space as the material the install documentation requires you to put there:

The configuration a gateway host holds

  • AIGW_SELF_SIGNED_JWT__SIGNING_KEY and DUO_WORKFLOW_SELF_SIGNED_JWT__SIGNING_KEY - the keys the gateway uses to sign its own tokens.
  • AIGW_GITLAB_URL and AIGW_GITLAB_API_URL - the inbound path back to your GitLab instance.
  • Model provider credentials for whichever backend you configured, whether that is Anthropic, Gemini, or Azure OpenAI.

Egress makes it worse, and again the documentation says so plainly: the gateway needs outbound access to your GitLab instance, to your model provider endpoints, and to customers.gitlab.com for license validation. A host with a sanctioned, encrypted, high-volume outbound channel to a third-party API is a comfortable place to run a command loop. If you self-host the gateway for data-residency reasons, that outbound allowance deserves an explicit egress policy naming the provider hostnames, not a blanket rule.

The patch closes the escape; the flow file still ships a shell

Here is what the coverage of this advisory missed. The flow registry v1 specification lists run_command as a standard tool, and the documented multi-agent example assigns it to a "tester" agent that "writes and runs automated tests." The specification puts it there on purpose. Custom flows reached general availability in GitLab 19.2, and a flow author composes components, picks a toolset, and writes Jinja2 prompt templates.

The specification also ships the controls, and they are off unless you set them. require_tool_approval forces a human decision before tool calls run. pre_approved_tools carves out the read-only ones so approval fatigue does not train people to click through. Approvals can also arrive workflow-wide through pre_approved_agent_privileges in the startRequest, or per session; GitLab documents the session check as fail-closed, matching on tool name plus canonicalized arguments, so the same tool with different arguments prompts again. Compare the two shapes:

version: "v1"
environment: ambient
components:
  # Shape A - the documented example. A shell, no approval gate.
  - name: "tester"
    type: AgentComponent
    prompt_id: "tester_agent"
    toolset: ["read_file", "create_file_with_contents", "run_command"]

  # Shape B - same job, shell removed, approval forced on what is left.
  - name: "tester_reviewed"
    type: AgentComponent
    prompt_id: "tester_agent"
    require_tool_approval: true
    pre_approved_tools: ["read_file", "list_dir", "find_files"]
    toolset: ["read_file", "list_dir", "find_files", "edit_file"]
routers:
  - from: "tester_reviewed"
    to: "end"
flow:
  entry_point: "tester_reviewed"

That makes the durable control an access review. GitLab's documentation is specific about who can do what: you need the Maintainer or Owner role on a project to create a custom flow and the same roles to enable one, while Developer and above can run it. So the population that can author the YAML is your project Maintainers and Owners. Go count them:

# Flow configs across your repos that hand an agent a shell
grep -rn --include='*.yml' --include='*.yaml' \
  -e 'run_command' -e 'require_tool_approval' -e 'pre_approved' .

# Who can author or enable a flow: access_level 40 (Maintainer) and 50 (Owner)
curl -s --header "PRIVATE-TOKEN: $GITLAB_TOKEN" \
  "https://gitlab.example.com/api/v4/projects/$PROJECT_ID/members/all?per_page=100" \
  | jq -r '.[] | select(.access_level >= 40) | [.username, .access_level] | @tsv'

If that list is longer than the list of people you would trust with a shell on a host holding your model credentials, you have a real finding that no patch addresses.

Self-hosted or GitLab-hosted: the trade you are making

People self-host the gateway for defensible reasons. Prompts and source context never leave the network. You choose the model and the region. An air-gapped or offline-licensed install stays offline. For a defense subcontractor or a firm with contractual data-residency terms, that is not a preference.

The cost is now measurable. You have taken ownership of a component that has produced two CVSS 9.9 remote-code-execution bugs of the same class in eight months, patched on a release train separate from your GitLab upgrades, versioned by container tag, and sitting on a host that holds model credentials and JWT signing keys. GitLab patched its own fleet before publishing and says it ran targeted outreach to self-hosted customers ahead of the release post. The hosted customers got the fix for free. You got an email.

My recommendation has a threshold in it. If you cannot name the person who will see the next AI Gateway patch release within 48 hours and apply it, you should not be running the gateway yourself, and the data-residency argument does not survive a compromised host that can reach your GitLab API. If you can name that person, self-hosting is defensible - treat the gateway as perimeter-grade infrastructure rather than as part of a developer tool.

Pin the gateway tag, then name who may author a flow

Three things this week, in order. Upgrade the gateway to 19.2.4, 19.3.2, or 19.4.1, whichever matches your branch, and write the tag down somewhere a human reads so the next advisory has a baseline to compare against. Pull the Maintainer and Owner list for every project where Duo Agent Platform is enabled, and cut it to the people who should hold a shell on that host. Then set require_tool_approval: true on any flow whose toolset includes run_command, and keep pre_approved_tools limited to the read-only set so the prompts still mean something. If you want someone to review the AI services you have stood up in the last year - what they can reach, what credentials they hold, and who can change their configuration - that review is work we do.

Do you know what your self-hosted AI services can reach?

We inventory and review the AI gateways, agent platforms, and model endpoints that teams stood up during the last year, including who can change their configuration and what credentials sit on those hosts. Book a session and we will map your AI stack and tell you what to change first.