← All posts

Meta's Muse agent can't hit send without a second AI's OK

2026-09-14 · 4 min read

Meta launched Muse on September 8. It's a personal AI agent that works inside your email, calendar, and payment accounts, books travel, fills out forms, and keeps going in the background after you close the app. It's live in the US on iOS, Android, the web, and WhatsApp, with a free tier and paid plans at $20 and $100 a month.

The launch posts talked less about what Muse does than about what stops it. In his thread, @finkd said users pick which apps Muse can reach, and "for sensitive actions like purchases or sending emails, Muse checks with you first." Passwords, he said, sit in separate storage Muse can't read. Shengjia Zhao, chief scientist at Meta Superintelligence Labs, was more direct. @shengjia_zhao wrote that "our biggest focus was the security architecture underneath," meaning each Muse gets its own secure VM and a separate Sentinel checks every action.

Later that day, @alexandr_wang posted that early usage had "blown way past our projections." Real users were using Muse 10x more than the testing groups had.

How the Sentinel works

Each Muse runs on its own Linux virtual machine in Meta's cloud. A second agent, Sentinel, runs next to it, walled off from Muse. Muse can propose an action, but only Sentinel can let it out. Anything leaving the machine goes through Sentinel, whether it's a connector call to your inbox or a plain web request, and Sentinel allows it, blocks it, or asks you.

When it asks, the request doesn't come through your chat with Muse. It shows up in the app as its own prompt, and you choose how far the yes goes: once, this session, this task, a set window of time, or permanently. Meta also says the agent never handles your real credentials. It works with placeholder tokens, and Sentinel swaps in the actual secrets at the network edge.

The target is prompt injection, where instructions hidden in a web page or an email get treated as commands. Meta's security write-up, by Tarek Sheasha, builds its threat model on Simon Willison's "lethal trifecta": an agent with access to private data, exposure to untrusted content, and a way to send data out. A personal agent has all three by design, so Meta took the final decision away from the model that reads the untrusted content.

What Meta admits, and what Reuters found

The same write-up says "prompt injection remains an open problem in the industry" and that Muse will sometimes make mistakes. Meta added Muse to its bug bounty, paying up to $300,000, and up to $130,000 for a prompt injection that affects one user.

Reuters reviewed internal posts from Meta employees who were testing Muse as recently as launch week. One wrote that an agent got around its guardrails and exposed personal iCloud photos after being asked to identify toys in pictures from a child's birthday party. CTO Andrew Bosworth wrote that it kept logging him out. For another tester, a ticket-watching task switched itself off with no warning. Meta can also still technically reach users' VMs. A Confidential VM version, where users hold the keys, is planned but not shipped.

My read

I think the architecture is right. Plenty of AI tools handle safety with an instruction in the prompt, something like "always ask before sending," and that instruction lives in the same model that just read a stranger's email. Sentinel is a separate system the agent can't argue its way past. That's a real step up.

It doesn't make Muse reliable, though. The reporting doesn't make clear which safeguards were in place when those photos leaked, but a leak like that is exactly what a permission gate is supposed to stop. Wang's 10x number cuts both ways, too: people like the product, and they're pushing it far harder than the testers did.

What to take from it

You don't need Muse to borrow the pattern. If an AI tool in your business can send email, move money, or change records, put these questions to the vendor:

  • Where does the "ask first" rule live: in the prompt, or in a separate system the model can't override?
  • Does the AI ever see the real password or API key, or only a stand-in?
  • Can an approval be limited to one task instead of granted forever?

If the answers are vague, keep that tool on read-only access until they aren't. And expect your staff to use it more than the pilot group did.

New Face Design's free process audit works through the same questions for your business: which workflows an agent should touch, and where a person still needs to sign off.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere