$npx skillfedfor your agent
RESEARCH

nanoMuse defines a personal agent by who runs it, not what model powers it

on: nanoMuse: An Open-Source Personal Agent for Every Device You Own

The central claim here is architectural, not technical: a personal agent is defined by who runs it and where, not by what model powers it. The paper makes this argument by dissecting Meta's Muse — released September 2026 — and then presenting nanoMuse as a GPL-3.0 counterpart that runs on hardware the person already owns.

The Muse analysis is unusually candid about its sources. Three public Meta documents plus a copy of Muse's production system prompt, each statement tagged as documented, prompt, observed, or inferred. What emerges is a picture of a product whose novelty lies not in its tool-calling loop — described here as standard — but in placing every safety mechanism outside the agent's reach. The Sentinel is the sole permission authority; the agent proposes, never grants. Credentials are surrogates swapped at the network boundary. The browser sub-agent sees an accessibility-tree snapshot, never the DOM, which is an explicit capability trade for safety. The phone is a data source and command target, not a screen the agent operates — iOS forbids it entirely, and Android computer use is absent from Muse.

That gap is where nanoMuse begins. The Android app runs a complete agent on the device itself, with a one-action-per-screenshot loop for the phone's screen. The desktop side ports UI-TARS-desktop. Devices share one conversation over an optional relay whose stored contents are enumerated precisely: hashed identifiers, token-count ledgers, conversation text the person chose to sync — never file contents or tool outputs. The relay is a single process on one machine, and the paper says so plainly.

The Sentinel here is a policy boundary in the same process family as the agent, not a privilege boundary outside it. The paper states this limitation without softening: a compromised device is a compromised agent. The taint rule — once a conversation reads private data, any outbound call not on the allow-list becomes an ask — and the decision order mitigate the exposure but do not close it.

Memory is Markdown files. The phone side uses named files (SOUL.md, USER.md, GLOBAL.md, a diary, HEARTBEAT.md). The desktop store tracks changes and allows undo. What it does not yet record is which model wrote a line, when, or with what confidence — the paper calls this out as a real problem: a line a weak model guessed is read as fact by a stronger one later.

The roadmap is honest about what does not exist yet. There is no success rate for the hands, no count of human take-overs. The evaluation suite — AndroidWorld, OSWorld, MemGUI-Bench, OS-Harm — is listed as near-term work, not present capability. The default-on training-data switch on the community relay is flagged as the project's least certain design choice.

Version 0.1.40, October 2026. Minimal is meant literally.

A complete, self-critical blueprint for a personal agent you can actually run — honest about every gap between its Sentinel and Muse's privilege boundary.

Sources & links