MLflow's Webhook SSRF Was Exploited Within Hours. The Bigger Bug Is the Open Tracking Server

CVE-2026-64849 lets an unauthenticated caller read cloud metadata through a test endpoint; patch to 3.15.0 and take MLflow off the internet.

2 min read ·

A critical server-side request forgery in MLflow, tracked as CVE-2026-64849 and scored 9.3, was under attack within hours of its CVE being assigned on 17 August. watchTowr's honeypots saw scanning and exploitation of internet-facing MLflow Tracking Servers aimed at cloud credentials, and CISA added the bug to its Known Exploited Vulnerabilities catalogue on 19 August. Every MLflow release before 3.15.0 is affected, and BleepingComputer notes that the default Tracking Server configuration exposes the vulnerable API without authentication.

How the bug works

MLflow's model registry can fire webhooks to notify CI pipelines or chat channels when a model's state changes. Since version 3.10.0 MLflow has checked that a webhook's destination resolves to a public address, specifically to block SSRF, so this CVE is a bypass of an existing defence. The problem is a time-of-check to time-of-use gap. The hostname is validated once, but the delivery code resolves it again when it actually sends the request, and it follows redirects. An attacker who controls DNS for a domain, or simply runs a public endpoint that redirects, can pass validation with a harmless address and have the real request land on 169.254.169.254 or an internal service.

What makes it serious is the POST /api/2.0/mlflow/webhooks/{id}/test endpoint. It is unauthenticated on default deployments and returns the upstream response status and body to the caller. That turns a blind SSRF into a full-read one: the attacker asks MLflow to "test" the webhook, and MLflow sends back whatever the cloud metadata service returned, including temporary IAM credentials for the instance's role.

Why ML infrastructure keeps ending up here

An MLflow server is easy to stand up and is often treated as a lab tool, not production infrastructure. It tends to run on a cloud VM or a Kubernetes pod whose role can read the artifact bucket, often the training-data bucket too, and sometimes much more. Data science teams expose it so collaborators can reach the UI, and authentication is an add-on most deployments never enable. The Cloud Security Alliance's research note describes unauthenticated tracking servers as common in practice, and counts more than thirty published MLflow CVEs since 2022. The webhook bug is the trigger, but the deployment is where the exposure comes from.

This is the pattern behind a lot of AI-infrastructure incidents. The model-serving and experiment-tracking stack grew up in notebooks and was moved into production without the controls the rest of the platform has.

What to do

  • Upgrade to MLflow 3.15.0 or later, including any instances that individual teams run themselves. Search your cloud accounts for them, because they rarely show up in the central inventory.
  • Put authentication in front of the Tracking Server and limit it to trusted networks. A reverse proxy with SSO is a reasonable start, but it does not replace the patch: the CSA note points out that any user who can still call the test endpoint can use it to reach the cloud environment. The default unauthenticated setup should never face the internet.
  • Cut the path to metadata. Block egress from the MLflow host to the instance metadata address. On AWS, require IMDSv2 with a hop limit of one, which makes simple GET-based SSRF against metadata much harder.
  • Audit webhooks for entries nobody recognises, and review the MLflow server's logs for calls to the test endpoint.
  • Assume exposed credentials were taken. If an unpatched server was reachable after 17 August, check your cloud audit logs for the instance role being used from unfamiliar IPs, and scope that role down while you are there.

Federal agencies were given until 2 September. Everyone else should move faster, because the exploitation timeline here was measured in hours, not weeks. The bigger fix is cultural. ML platforms hold some of the most valuable data and credentials in an organisation, and they need the same baseline of authentication, network controls and least-privilege roles as anything else in production.


Sources

Responses (1)

Sign in to leave a response.

  • Every ML team I have worked with has at least one MLflow box someone spun up for a project and forgot about. The inventory step is the hard one.

More from Hana Rahman

Recommended from Horizon