The Fix Was Already Shipped: Lessons From LiteLLM’s 2026
Last Updated on October 6, 2026 by Editorial Team
Author(s): Nick Hystax
Originally published on Towards AI.
The Fix Was Already Shipped: Lessons From LiteLLM’s 2026

In late August, Microsoft’s security researchers published a detailed account of an intrusion into a LiteLLM gateway. The attackers read the gateway’s process environment and pulled out model-provider API keys, the LiteLLM master key, and the database connection string. They used that connection string to dump the tables holding model configuration and proxy-issued virtual keys. Then they installed a cryptominer, added an SSH key to a service account, and locked down their files so cleanup would be harder.
Microsoft assessed that the way in was most likely a known vulnerability chain involving LiteLLM’s MCP test endpoints. LiteLLM had shipped the fix for its side of that chain months earlier.
That detail ran through LiteLLM’s whole year. It is one of the most popular open-source LLM gateways, and in 2026 it took a series of serious hits. In almost every case where attackers actually got in, the code had already been fixed. The deployments had not.
What happened, roughly in order
The year started with a supply-chain attack. On 24 March, two malicious releases, 1.82.7 and 1.82.8, were published to PyPI with stolen publishing credentials. They were built to steal whatever credentials they could find, from SSH keys and Kubernetes tokens to cloud secrets. The second one hid its payload in a .pth file. Python processes those files every time the interpreter starts, so the package never even had to be imported. According to LiteLLM’s own incident report, teams running the official proxy Docker image were not affected, because that image pins its dependencies.
April brought two critical authentication bugs. CVE-2026–35030, rated 9.4, let attackers bypass JWT authentication on deployments that had it enabled. CVE-2026–42208, rated 9.3, was worse: a SQL injection in the code that checks API keys, reachable with no credentials. The project fixed it in a stable release before publishing the advisory. Sysdig still saw targeted exploitation about 36 hours after the advisory was indexed, aimed squarely at the tables that store keys and credentials.
June put MCP in the spotlight. CVE-2026–42271, a command injection disclosed in April in the endpoints used to preview MCP servers, originally required a valid API key. Horizon3.ai showed it could be chained with a host-header bypass in Starlette, the web framework under LiteLLM, to achieve remote code execution without credentials. On 8 June, CISA added the flaw to its Known Exploited Vulnerabilities catalog. A week later, Obsidian Security disclosed a chain of three CVEs, rated 9.9, that let a default low-privilege user promote themselves to proxy admin and run code on the host.
The Microsoft report followed on 26 August. Seven days later, on 2 September, a second LiteLLM vulnerability went onto CISA’s list. CVE-2026–59822 let anyone open an authenticated MCP session with a made-up bearer token, then list and call whatever MCP tools the gateway had configured. The fix, in version 1.84.0, had been public since the end of June.
Every exploited bug already had a fix
Put the dates side by side, and a pattern appears. The SQL injection was patched before anyone outside the project knew about it, and exploitation still began within two days of disclosure. CISA confirmed attacks on the MCP command injection weeks after the fixed release shipped. The MCP authentication bypass had been fixed for two months by the time CISA listed it. Microsoft’s case involved a chain whose LiteLLM half was fixed in the spring.
None of this describes a project that ignores security. The maintainers fixed problems quickly, ran a bug bounty, published advisories, and kept the official Docker image locked down well enough to dodge the March attack. The weak point was somewhere else: the time between a fix existing and that fix running on someone’s server.
That gap is an operations problem. It is also the one part of an open-source gateway that nobody can ship for you.
Why the gateway is such a good target
Most internal services hold a secret or two. A gateway holds nearly all of them. It keeps an API key for every model provider the company uses, a master key that can mint new keys, a database full of virtual keys, and, increasingly, connections to MCP tools that reach code repositories, ticketing systems, and databases. It also sees every prompt and every response in plain text.
The first line of Microsoft’s mitigation guidance is blunt:
“Treat AI gateways as Tier-0 secrets stores.”
Tier-0 is the label security teams normally reserve for domain controllers and identity systems, the components that grant access to everything else.
Obsidian Security made a related point in June. A compromised gateway sits between your agents and the model, so it can alter answers on their way back. When those answers drive tool calls, whoever controls the gateway is effectively steering your agents.
If you run LiteLLM yourself, start here
Most of what follows comes from the advisories themselves and from Microsoft’s mitigation list. None of it is exotic.
1. Check your version today. Anything below 1.84.0 is exposed to a flaw CISA lists as exploited. While you’re at it, find out who would apply the next fix and how long it would take them.
2. Keep management surfaces off the internet. The admin UI and management endpoints should not be reachable from outside, and both the API and the UI should require authentication.
3. Switch off what you don’t use. Two of this year’s worst bugs lived in MCP endpoints. If your gateway doesn’t serve MCP, the CVE-2026–59822 advisory’s own workaround is to disable MCP routes or block /mcp/ at your reverse proxy.
4. Get provider keys out of environment variables. In Microsoft’s case, the attackers read them straight from the process environment. A managed secret store, per-team virtual keys with spend limits, and a master key nobody uses day to day all shrink what a single compromise exposes.
5. Install from pinned, verified artifacts. The official image with pinned dependencies came through March untouched. If you install from PyPI, pin exact versions and hashes so a poisoned release can’t slip in during a routine rebuild.
6. Restrict outbound traffic. A gateway needs to reach a known list of provider endpoints. Deny the rest by default, and stolen keys and data have a much harder time leaving.
7. Watch the gateway process. A model proxy has no business starting a shell, running curl, or reading /proc/1/environ. If it does, treat it as an incident.
8. If any environment ever installed 1.82.7 or 1.82.8, rotate everything it could reach. Removing the package doesn’t undo what was copied off the machine.
Deciding whether to keep running it
That checklist is routine for some teams and unrealistic for others. Three questions usually settle which group you’re in.
Can you ship a gateway upgrade within days, including on a Friday? CISA gave federal agencies two weeks to patch the September listing, and a regulator set that deadline. The April SQL injection showed attackers working on a 36-hour clock.
Is there a named owner? “The platform team” is not an owner. A gateway looked after in whatever time is left over from product work tends to fall a few versions behind, and this year showed exactly what a few versions behind can cost.
Does the gateway have to live inside your perimeter? Data residency or air-gapped requirements can rule out hosted services. They don’t automatically mean your own engineers have to operate the software, since some vendors support on-premises deployments.
If the first two answers are yes, running an open-source gateway yourself is a reasonable choice, and LiteLLM is still capable. If either one is no, the third question points to the alternative: a managed service where hosting is allowed, or a vendor-supported install where it isn’t. Swapping one self-hosted gateway for another won’t help much either way, because the upgrade work comes along with you.
What the year actually showed
Open-source gateways earned their popularity. They are flexible, support a huge range of providers, and let you read every line of code. What they can’t do is upgrade themselves. In 2026, the gap between “fixed upstream” and “fixed in production” is where the damage happened, and closing it is work someone has to own.
• • •
Disclosure: I work on tooling for AI agent governance, so this is a problem I think about professionally — and, yes, with some bias. I’ve kept this piece about the problem, not any product.
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.