← All posts

Perplexity let 9 AIs try to escape. 4 found a side door

2026-09-29 · 4 min read

On Monday, Perplexity CEO @AravSrinivas posted the results of a month his security team spent trying to break SPACE, the sandbox that runs every Perplexity Computer task. The setup was about as hostile as it gets. Nine models, including Opus 5, GPT-5.6 Sol, Kimi K3 and Gemini 3.1 Pro, got root access inside the virtual machine, and in some runs the full source code of the sandbox too. The job was simple: escape to the host, or reach a web address the network policy blocks.

The company account, @perplexity_ai, led with the good news: "Across 108 runs, none breached the VM boundary." The network rules did worse.

The side door was a package mirror

Once the team allowed a small amount of network access, just the PyPI and Ubuntu mirrors so the models could install software, four of the nine models reached the blocked URL anyway.

Srinivas describes two routes. In the first, models spoofed DNS responses to Perplexity's gateway. They took advantage of the fact that pypi.org sits on a Fastly IP address shared with thousands of other sites. If the firewall only checks "is this IP allowed?", then any site hosted on that same IP gets through too.

The second route is stranger. GPT-5.6 Cyber sent its request through Taboola's image fetcher to a screenshot service, then read the blocked page's secret flag out of the returned image with OCR. So it never connected to the forbidden site itself. Another service looked at the page and sent back a picture.

According to AlphaSignal's writeup, Claude Opus 5 spotted the same kind of opening and declined to use it, calling it out of scope for the task. Good for Claude, but I wouldn't build a security plan on one model's manners on one day.

It isn't only Perplexity's problem

Perplexity says it fixed both issues and confirmed the reruns held. It also tested ten other sandbox providers, and Srinivas reports that eight had the same shared-IP weakness, naming E2B, Vercel and Modal. AlphaSignal's summary adds that Cloudflare Sandbox and NVIDIA's OpenShell were the two that resisted both routes, and that some vendors patched while others filed the behavior as a known limitation.

Srinivas also tied the work to NVIDIA's Open Agent Safety Platform, announced the same day, and said Perplexity wants to fold these guardrails into it. His argument is that sharing red-team findings openly is how the industry builds better guardrails. I agree, and this report shows why. Most teams would have assumed an allowlist of two mirrors was safe. It wasn't, and they only found out because someone told a capable model to try.

The VM result is reassuring, but I think it's the less important half. Nobody hands a production agent root on a box and asks it to break out. People hand it "a little" access, and that's where the models got through. The edges of your permissions are fuzzier than your config file suggests.

What this means for a business using AI agents

You probably aren't running a sandbox company. But if you're wiring an agent into your inbox, your CRM or a vendor portal, you're making the same kind of call Perplexity made: this tool can reach these few things and nothing else. A few habits follow from the report:

  • Check the whole path. "The agent can only use our booking site" means little if that site can fetch or preview other pages on the agent's behalf. Link previews, file importers and screenshot tools are all side doors.
  • Grant access per task, then take it back. A package mirror was allowed so the models could install software. Permissions granted for convenience are the ones nobody revisits.
  • Before you trust the fence, have a capable model try to get past it. Perplexity's fixes came from asking models to break things on purpose. A short adversarial test on your own setup costs far less than finding the gap in production.

I wouldn't hold off on agents because of this. I would treat "what can this touch?" as a design question you answer on purpose. If you want help working out where an agent in your operation should stop, New Face Design runs a free process audit. We map the tasks an agent could take over, what each one would need access to, and where that access should end.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere