Muse Rebuilt OpenClaw's Shape. The Prompt Injection Came With It.
Muse wraps a persistent agent around messages, files, calendars, memory, and tools—the same architectural shape OpenClaw made popular. It also recreates the security problem: untrusted content can steer an agent that acts with your authority.
Muse did not invent a new agent architecture.
Meta describes Muse as a personal AI agent that works proactively toward a user's goals. Underneath the launch language is a familiar shape: a long-running agent, persistent memory, connectors into personal data, event-driven triggers, and tools that can act on the user's behalf.
That is also the shape OpenClaw popularized: keep an agent running, connect it to the places where work arrives, let it remember context, and expose skills that turn model output into action. Muse is not necessarily a source-code fork. The important point is that it recreates the same trust topology at consumer scale.
The architecture is compelling because it removes friction. It is dangerous for exactly the same reason. Messages, documents, web pages, and calendar events flow into the same reasoning loop that decides which privileged tool to invoke next.
Copy the shape, inherit the failure modes.
OpenClaw's security problem was never simply that an LLM might say something strange. It was architectural: untrusted text and trusted instructions entered one context, while the resulting model output could reach tools backed by real credentials. Skills and connectors multiplied the blast radius. Persistent memory gave malicious instructions somewhere to survive.
Muse brings those same ingredients together. It can ingest information from systems people treat as private and trustworthy, retain context, and use that context to plan later actions. A polished interface and a first-party brand do not remove the confused-deputy problem. They make the deputy more widely trusted.
The comparison matters because teams keep treating prompt injection as a model-quality bug. It is not. Prompt injection becomes exploitable when architecture turns model interpretation into authority. If content can influence planning and planning can reach a privileged tool, the security boundary is the tool gateway—not the model's promise to ignore bad instructions.
Muse can be prompt-injected because it has to read.
A useful personal agent cannot limit itself to text the user typed directly. It has to read the email someone else sent, the document a coworker shared, the event description an organizer wrote, and the web page a search result returned. Each object is both data for the task and a delivery channel for instructions the user never approved.
The attack does not need to look like ignore previous instructions. It can be styled as document metadata, an accessibility note, a workflow requirement, a quoted policy, or invisible text. The agent has to infer which language describes the world and which language is trying to control it. No model can make that distinction perfectly across every format and language.
Once injected, the attacker does not need to take over the whole agent. A smaller steering effect is enough: add a recipient, choose an attacker-controlled link, summarize the wrong file, disclose one remembered fact, create a recurring event, or store a poisoned instruction for the next session. Persistence turns a one-time read into a durable foothold.
This is the OpenClaw lesson in one line: anything the agent reads can become something the agent does. Muse has not escaped that rule. It has put the rule inside a product meant for everyone.
Consent at connection time is not consent at action time.
The usual OAuth screen asks for broad access once: read your mail, manage your calendar, access your files. An agent may then reuse that authority across thousands of decisions the user never sees. The permission was technically granted, but the intent behind each action remains unknown.
Agentic products need a second layer above OAuth: action-level policy. Reading a flight confirmation to answer a question is not the same as forwarding it. Drafting a calendar event is not the same as inviting twenty people. Finding a subscription is not the same as cancelling it. The underlying service may represent each pair with the same broad scope, but the product cannot treat them as equivalent risk.
The safe default is progressive authority: let the agent observe narrowly, prepare reversibly, and ask before it commits consequential or externally visible actions. High-risk capabilities should be short-lived, purpose-bound, and issued only when the action is ready to execute.
Four controls matter more than another system prompt.
1. Enforce policy outside the model. The tool gateway must validate user identity, resource ownership, allowed operation, risk tier, and current approval state. The model can request an action. It cannot authorize its own request.
2. Separate reading from acting. Content-retrieval tools and state-changing tools should not share unrestricted credentials or an implicit trust boundary. Treat every email, page, attachment, and retrieved document as hostile input, even when it came from a service the user trusts.
3. Make effects inspectable and reversible. Show the exact recipients, fields, amounts, permissions, and resources that will change. Prefer drafts, previews, version history, delayed execution, and rollback over immediate irreversible calls.
4. Log the decision chain. A useful audit record includes the user's request, the untrusted inputs consulted, the policy decisions made, the tool arguments proposed, the approval captured, and the observed result. Logging only the final API call tells you what happened, not why the agent believed it should happen.
Security testing has to follow the whole action path.
Most agent evaluations stop when the model says something unsafe. That misses the failures that matter. The real test starts with hostile content, follows it through retrieval and planning, observes the attempted tool call, and verifies whether the backend blocks the effect.
A finding is not ‘the model appeared willing to send data.’ A verified finding is: the injected document changed the tool arguments, the gateway accepted them without fresh authorization, and data reached a recipient outside the user's intended scope. That end-to-end proof is what separates an alarming transcript from an exploitable vulnerability.
This is why agentic security cannot be reduced to model safety. Model behavior is one control inside a distributed system containing identity, permissions, tool schemas, data provenance, approval UX, APIs, and recovery paths. You have to test the system that ships.
Audit an agentic app as an authorization system
Inventory every agent tool, find authorization that lives only in a prompt, and trace untrusted input to unauthorized effects.
View prompt on GitHub →The move.
Muse points toward the interface people actually want: one agent that understands context and gets things done. That future is plausible precisely because the risky part is no longer hypothetical. Agents are being connected to the systems where people communicate, organize, buy, decide, and store private information.
The winning agent will not be the one with the longest system prompt or the most warnings. It will be the one whose authority is legible, bounded, and continuously verified.
Treat every agent as a new principal in your authorization model. Give it less authority than the user, make it prove intent before every consequential action, and test the complete path from hostile input to real-world effect. Intelligence makes the agent useful. Boundaries make it safe enough to use.
Run Composed on your last 10 PRs.
Apply for design partner access and we’ll show which findings are real, exploitable, and worth fixing — free.
Run it on 10 PRs →