SGLang RCE Flaw Lets Malicious GGUF Models Execute Arbitrary Code
SGLang is an open-source large language model inference framework that helps companies and research institutions deploy AI services. CVE-2026-5760, disclosed in 2026, is a remote code execution vulnerability. The risk extends beyond model output and could directly expose servers hosting inference services and sensitive data to attackers.
The latest investigation found that attackers can embed malicious payloads in Jinja2 templates within GGUF models. Once an administrator loads the model, a request to a specific endpoint may trigger arbitrary code execution. As of July 20, 2026, affected deployments should restrict untrusted models, reduce endpoint exposure, and apply updates or isolation measures in line with SGLang project advisories.
All Coverage
1 original reportsThe Backstory
The history behind this eventThree Critical SGLang Flaws Expose AI Servers to Remote Code Execution
SGLang is an open-source AI inference framework used to deploy and run large language models. If its underlying services are compromised, attackers could take control of computing servers and steal models or data. All three vulnerabilities disclosed by cybersecurity firm Antiproof involve remote code execution, posing a significant threat to internet-facing AI infrastructure.
Antiproof’s latest disclosure covers three high-risk vulnerabilities: CVE-2026-7301, CVE-2026-7302 and CVE-2026-7304. They could allow unauthenticated attackers to execute code directly on SGLang servers. Affected releases include v0.5.5 and versions after v0.4.1.post7. As of July 20, 2026, the project had not released patches, and users were advised to restrict access to the service in the meantime.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.