LMCache’s distributed architecture contains a critical vulnerability that can let an unauthenticated network client execute code through unsafe Python deserialization. The flaw, CVE-2026-105192, received a CVSS v3.1 score of 9.8 CRITICAL.
JFrog, the CVE Numbering Authority for the issue, published and updated the CVE record on 7 October 2026. The company’s security researchers also disclosed the vulnerability that day and credited Yuval Moravchick with finding it.
The risk is configuration-dependent. LMCache binds the affected transport to localhost by default, but operators may expose it on a routable address for multi-node deployments. At disclosure, the reporting identified no patched LMCache release and provided no verified evidence that attackers had exploited the flaw in deployed systems.
A network message reaches pickle.loads before authentication
LMCache is open-source caching software designed to accelerate large-language-model serving systems such as vLLM. The vulnerable component is its multiprocess mode, also called distributed mode, where LMCache operates as a separate cache server and workers communicate with it through ZeroMQ.
According to the CVE Program record, the server opens a ZeroMQ ROUTER socket that workers use to register and exchange KV-cache blocks. That socket does not authenticate clients.
Messages are encoded with msgpack. During decoding, extension code 1 is sent to DeviceIPCWrapper.Deserialize, which calls Python’s pickle.loads on data controlled by the sender. Crucially, deserialization happens while the server is still processing request arguments, before the relevant message handler runs.
An attacker who can reach the transport can therefore send one crafted ZeroMQ DEALER message and trigger code execution as the account running LMCache. No credentials or user interaction are required.
The assigned vector is:
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
The record classifies the weakness as both CWE-306, missing authentication for a critical function, and CWE-502, deserialization of untrusted data.
Remote exposure depends on how operators bind port 5555
The affected transport uses port 5555 by default. It listens only on localhost unless an operator supplies a routable address through --host, a configuration that allows workers on other nodes to connect.
Consequently, an installation using the default local binding is not remotely reachable from another host through this interface. The exposure changes when the service is bound to a cluster address, all interfaces or another network-accessible endpoint.
The primary report says LMCache’s example Kubernetes deployment starts the server on every network interface. It also reports that an LMCache instance embedded within a single vLLM process does not open the affected port.
JFrog’s description says official LMCache container images run the process as root. Successful exploitation in those images could therefore execute code with root privileges inside the environment where the process operates. That outcome does not automatically apply to installations configured to run LMCache under a less-privileged account, nor does it establish what host resources any particular container can access.
No exploitation in the wild has been verified in the cited records. The technical path supports remote code execution when the socket is reachable, but that finding is separate from evidence of real-world compromise.
Version records differ on the full LMCache range
The official CVE entry marks LMCache 0.3.9 as affected but does not specify an upper version boundary. Its references point to vulnerable code in both v0.3.9 and v0.5.5.
The news report gives a broader range: releases from 0.3.9 through 0.5.5, along with 0.5.6 release candidates and the development branch. It describes 0.5.5 as the latest stable release and says version 0.3.9 was released in October 2025.
These claims have different evidentiary scope. The CNA record confirms an affected entry for 0.3.9, while the wider version and branch coverage comes from the news report rather than a complete affected-range declaration in the supplied CVE record.
No fixed LMCache version was identified at disclosure. The CVE record contains no patched-version entry, and the reporting says a corrected release was not yet available on 7 October 2026.
Network isolation is the immediate defensive measure
JFrog’s interim recommendations, as relayed by the report, focus on preventing untrusted systems from reaching the ZeroMQ transport:
- Do not bind the multiprocess server to a routable address while a fixed release is unavailable.
- Keep the service on localhost where distributed operation is unnecessary.
- If remote workers must connect, limit access to a trusted cluster network.
- Apply firewall controls to port 5555 and allow only required peers.
Firewall filtering reduces the reachable attack surface but does not remove the unsafe deserialization behavior. Any allowed or compromised host that can connect to the socket may still be able to deliver the malicious message.
Operators should also verify which account runs the LMCache process and minimize its privileges. This limits potential consequences but is not a substitute for correcting the vulnerable code.
The report says JFrog’s advisory did not provide a procedure for determining whether exploitation had already occurred. The cited material also supplies no indicators of compromise or forensic query for CVE-2026-105192; that is limited to the available reporting and does not establish that guidance cannot exist elsewhere.
Separate vLLM flaw can terminate EngineCore
A related but technically distinct issue affects vLLM deployments using the built-in LMCache-MP KV connector. CVE-2026-105756 allows a malformed cache_salt value to trigger an uncaught exception and terminate EngineCore.
The vulnerability has a CVSS v3.1 score of 6.5 and a Moderate rating:
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
It is classified as CWE-20, improper input validation, and CWE-248, uncaught exception. Unlike the LMCache RCE, its recorded vector specifies low privileges required and impact only to availability.
Affected OpenAI-compatible models for Completions, Chat Completions and Responses accept any non-empty cache_salt, but the downstream LMCache-MP IPCCacheServerKey imposes additional restrictions. It rejects salts longer than 128 characters or containing @, /, \ or NUL.
Those constraints were not enforced before the value reached the scheduler’s cache lookup. A rejected salt raises a ValueError that is not caught at the connector or scheduler level, according to the GitHub Security Advisory. EngineCore then treats the exception as fatal, disrupting service for concurrent users. The advisory gives cache_salt="/" as an example.
The issue was confirmed on vLLM 0.25.1, commit 752a3a504485, and applies only when the opt-in LMCache-MP connector is enabled. That connector requires lmcache >= 0.4.4.
The NVD description and general advisory range identify versions prior to 0.30.0 as affected and 0.30.0 as the fix. However, another advisory field lists patched versions as >= 30.0.0. That inconsistent value should not be silently treated as equivalent; operators should confirm the package release they deploy. The advisory was published on 23 September 2026, and the report says vLLM 0.30.0 was released on September 22.
Additional LMCache allegations remain unconfirmed
The report also describes six other LMCache security submissions opened by one GitHub account on 6 October 2026. They allege cross-tenant access to cached data and unauthenticated access to network services capable of executing commands.
Those submissions have no CVE assignments, maintainer confirmation or identified fixes in the cited reporting. They rely on proof-of-concept claims and are separate from CVE-2026-105192.
One report alleges that an administrative HTTP server listened on every interface in 0.5.5, while 0.5.6 release candidates restrict it to localhost. Until maintainers validate the underlying claims, these reports should not be presented as confirmed LMCache vulnerabilities.
The unsafe pattern behind CVE-2026-105192 resembles the use of unauthenticated messaging with Python pickle that researchers discussed under the name ShadowMQ in November 2025. The reporting does not establish shared code or a common origin between those earlier findings and LMCache.




