ACT TONIGHT

Ollama / vLLM / LiteLLM: inference API open to the internet with no auth

Risk / Filed 28 Sep 2026 / CVE-2024-37032, CVE-2026-7482

ollamavllmlitellmexposed-apidocker

What: Ollama, vLLM and LiteLLM don't have auth turned on by default. If the port is reachable, anyone can run your models on your GPU, pull or push models, and read what the API exposes. People are actively scanning for these. SentinelLABS and Censys counted about 175,000 exposed Ollama hosts (Oct 2025 to Jan 2026). Nearly half of them advertised tool calling. Pillar Security found a marketplace ("Operation Bizarre Bazaar") that resells access to hijacked endpoints. Between March and June 2026, Sysdig and Zenity caught attackers using exposed Ollama and LiteLLM backends as the brain of automated attack tools.

Who it hits: You're affected if any of these is true:

  • OLLAMA_HOST=0.0.0.0 with 11434 port-forwarded, tunneled or on a VPS.
  • docker run -p 11434:11434 ollama/ollama on a box with a public IP. Docker's published ports get around ufw.

  • vllm serve on its defaults, which bind all interfaces. --api-key only guards /v1-style paths. /invocations, /score, /pooling and a few others stay open.

  • A LiteLLM proxy on 0.0.0.0:4000 with no master_key. Anyone can then spend the upstream OpenAI or Anthropic keys it holds.

Ollama versions below 0.17.1 add CVE-2026-7482 ("Bleeding Llama", CVSS 9.1). It takes three unauthenticated API calls (/api/create then /api/push) to leak heap memory: prompts, env vars and API keys. Versions below 0.1.34 also have Probllama (CVE-2024-37032), an RCE through /api/pull.

Check if you're affected:

# On the box: anything on 0.0.0.0 / [::] / * here is listening on every interface
ss -ltnp | grep -E ':(11434|8000|4000)\b'
# From OUTSIDE your network (phone hotspot, VPS). Any JSON back means you're exposed:
curl -s -m 5 http://YOUR_PUBLIC_IP:11434/api/tags

Do this:

  1. Tonight: stop Ollama listening on every interface. Run sudo systemctl edit ollama, set Environment="OLLAMA_HOST=127.0.0.1:11434", then run sudo systemctl daemon-reload && sudo systemctl restart ollama. For Docker, publish it as -p 127.0.0.1:11434:11434. Start vLLM with --host 127.0.0.1.

  2. Remove the router port-forward. If other LAN machines need the API, only let the LAN in, e.g. sudo ufw allow from 192.168.1.0/24 to any port 11434, and deny the rest. Or reach it over Tailscale or WireGuard instead.

  3. If you need it reachable from outside, put an auth proxy in front: Caddy or nginx with basic auth or forward-auth, or Cloudflare Tunnel with Cloudflare Access. For vLLM, only allow the paths you use through the proxy. Don't rely on --api-key.

  4. LiteLLM: set LITELLM_MASTER_KEY=sk-<long random> (it must start with sk-) and give clients virtual keys. Rotate every upstream provider key the proxy held if it was ever open.

  5. Upgrade Ollama to 0.17.1 or newer: ollama -v.

  6. Confirm the fix: from outside, the curl above should time out. The ss line should show only 127.0.0.1 or your LAN or tailnet IP.

If you were already hit: Look in the Ollama logs for unfamiliar /api/create, /api/push or /api/pull calls, and for models you didn't pull (ollama list). Rotate any keys that were in the process env or passed through LiteLLM. Check GPU usage history for load you didn't cause.

Why this level: It answers yes on all five rubric questions: commonly port-forwarded or tunneled, exploited in live campaigns, a core lab AI tool, big blast radius (keys, prompts, free compute, tool-calling agents), and no auth by default. Mass scanning of 11434 with live LLMjacking campaigns also triggers the act-tonight override.

Sources

How severity is decided. Source file: risks/2026-09-28-exposed-llm-inference-apis.md.