Command Palette
Search for a command to run...
nanoMuse: ユーザーが所有するすべてのデバイスのためのオープンソース・パーソナルエージェント
nanoMuse: ユーザーが所有するすべてのデバイスのためのオープンソース・パーソナルエージェント
Guangyi Liu Yong Liu Jiangning Zhang
概要
2011年頃のアシスタントは応答して待つだけであり、2023年頃のエージェントは一つのタスクを実行すると停止した。2026年9月、MetaのMuseは、アカウント、デバイス、記憶、そして持続する会話を備えた一人の人間のためのエージェントを示したが、それはベンダーのクラウド上で、一国において閉鎖的に提供されるものであった。このようなエージェントには、本人のアカウントやデバイスに対して操作を行い、数週間にわたってその人を記憶し、価値があるときには自ら話しかけ、自らが行ったことについて説明することが期待される。それはモデルではなく一種のソフトウェアであり、これまでオープンな対応物は存在しなかった。本報告は、パーソナルエージェントを5つの問いと3つのホライズンによって定義する。Metaの公開記録と本番プロンプトの複製からMuseがどのように構築されているかを読み解き、各記述にはその出典を付す。次に、GPL-3.0の下で提供されるオープンソースの対応物であるnanoMuseを提示する。nanoMuseは、本人が所有するすべてのデバイス上で動作する単一のエージェントであり、携帯電話の画面とコンピュータの画面を操作する手を備える。それらは誰でも運用できるリレーを介して一つの会話を共有し、すべての行動はSentinelを経由し、記憶は本人が読めるファイルであり、モデルは利用者の選択に委ねられる。そのサイズとコストは推定値として示される。何がオープンであるか、来歴付きの記憶、操作の手のための評価スイート、そしてその手のためのオープンモデルが、ロードマップとして提示される。
One-sentence Summary
Researchers from Zhejiang University propose nanoMuse, an open-source GPL-3.0 personal agent that runs on every device a person owns and shares one conversation through a self-hostable relay with screen control, sentinel-gated actions, readable memory files, and user-chosen models, as the open counterpart to Meta's closed Muse.
Key Contributions
- The report defines the personal agent through five questions and three horizons, tracing the lineage from the assistants of 2011 to the agents of September 2026.
- It reconstructs how Meta’s Muse is built from Meta’s public record and a copy of its production prompt, with each statement marked by its source.
- It introduces nanoMuse, an open-source GPL-3.0 counterpart that runs on every device a person owns, with hands on the phone’s screen and the computer’s, a shared conversation relay anyone can run, a Sentinel over every action, memory as files, and model choice. Size and cost are given as estimates, and an open roadmap covers memory with provenance, an evaluation suite for the hands, and an open model for them.
Introduction
The authors trace the shift from task-oriented assistants to personal agents, culminating in the September 2026 wave led by Meta's Muse. A personal agent acts on one person's accounts, devices, and files over time and must be accountable, but commercial offerings like Muse run as closed, vendor-hosted Linux VMs with only a small gadget SDK opened. Prior agents either handled single tasks in vendor clouds or, if open, ran on user infrastructure without full screen control, and cloud-only agents cannot reach everyday mobile apps that lack APIs or web pages. To address this, the authors present nanoMuse, an open personal agent that runs on hardware the user already owns, with hands on phone and computer screens, a Sentinel over actions, memory as files, and reported size and cost.
Method
The authors design Muse as a background personal agent structured across three distinct security zones: the surfaces of the user, a dedicated cloud environment, and isolated host services.
In this framework, each user is allocated a dedicated Linux virtual machine in the cloud, which is replaced approximately hourly while maintaining a persistent home directory. The agent and its tools operate within a container bounded by seccomp filters and reduced capabilities. The system prompt establishes a strict persona that works exclusively for the user, enforcing discretion and hard safety stops. The agent utilizes tools organized in namespaces, such as browser, chat, memory, and device control. A critical component of this architecture is the Sentinel, which acts as the sole permission authority for connector actions and network egress. The agent holds only surrogate tokens, which are swapped for real credentials at the network boundary after approval. The system employs kernel level taint tracking, where a tool process starts clean and becomes tainted upon reading user data, ensuring that only clean requests within a narrow policy pass without a prompt. Approvals are presented directly to the client as scoped capabilities rather than passing through the conversation.
In contrast, the authors develop nanoMuse as a decentralized personal agent where every device operates as a peer carrying a complete agent and its own Sentinel.
The Android implementation runs a complete agent inside the application, featuring a sandboxed Alpine Linux environment, a shell, a browser, and MCP servers. The desktop application utilizes an Electron shell around a harness with a Python runtime. Unlike the cloud hosted approach, nanoMuse places the agent directly on the device, granting it direct access to the phone screen. The Sentinel in this system functions as a policy boundary within the same process family, enforcing a fixed decision order that evaluates denied tools, user rules, allow and ask lists, data taint, and tool risk. Approvals are issued as scoped grants, and sensitive actions trigger mandatory warnings that cannot be bypassed. For screen interaction, the system employs a hybrid approach, climbing a ladder from skills and command line tools to a browser, and finally to direct screen manipulation via a one action per screenshot loop. Memory is managed as plain Markdown files that the user can open and edit, ensuring data portability and transparency. The architecture includes an optional and open relay that handles sign in, model call ledgers, and device synchronization without storing the actual content of the conversations.
The authors outline a structured development pipeline to advance the system capabilities across memory, device integration, and physical world interaction.
The immediate focus involves closing the feature gap by improving memory recall, reducing hand interaction failures, and enabling single command self hosting. The subsequent phase introduces an evaluation suite built from failed runs to score releases objectively, alongside the training of a small open model on user contributed traces to reduce vendor dependency. The long term vision extends the hands of the agent to the physical world, integrating cameras, appliances, and robot arms under the same approval and logging mechanisms, ultimately aiming to build an agent with memory and identity that survive across different models, devices, and vendors.
Experiment
This section outlines nanoMuse’s community roadmap in three horizons: near-term work focuses on memory, proactivity, conversation continuity, hand reliability, and one-command self-hosting; next steps introduce an evaluation suite built from failed runs and benchmarked with AndroidWorld, OSWorld, MemGUI-Bench, and OS-Harm, alongside a small open model for the hands trained on contributed traces to avoid vendor dependence; later work extends the hands to physical devices and long-lived agents with durable memory and identity. The roadmap concludes by inviting contributions such as skills, translations, devices, failing traces, and corrections.
The comparison spans open projects from mid-2025 and 2026 plus a cluster of commercial agents launched in September 2026. Phone screen operation is rare and absent from the September launches, while computer screen operation appears only in two cloud-based systems, Muse and Today. Open licensing is concentrated in the earlier and self-hosted projects, whereas the commercial cloud agents are closed. Only OpenMinis can operate a phone screen, and only on Android; none of the September 2026 agents has phone screen access. Muse and Today are the only listed systems with computer screen operation; OpenMuse and Manus Cue run loops without screen control. Open source systems generally run on the user's own computer, server, or app, while the commercial launches run in per-person cloud computers or virtual machines. Among the September systems there is a tradeoff: Muse and Today offer computer screen use but are closed, while OpenMuse is open but lacks screen operation.
nanoMuse's footprint varies widely by component: the Android client is compact, the desktop client is substantially larger to download and install, and the relay is lightweight enough for a single small server. Estimated self-hosting costs are modest, on the order of a few US dollars per month outside mainland China or tens of yuan inside it, plus model usage. Model call pricing is low enough that ordinary daily chat costs only a few fen. The Android app is the smallest release artifact, while the desktop app is far heavier to download and also uses about half a gigabyte of memory at idle on Linux. The relay is the lightest runtime component, idling at roughly 85 MB of memory and sized for a server with 1 vCPU and 1 GB. Self-hosting a relay is estimated at US4to6 per month outside mainland China or ¥30 to ¥60 inside it before model usage. The hands model costs more per token than the chat model, but a typical day of chat remains a few fen.
The first comparison examines screen-control capabilities and licensing across open and commercial agents, finding that phone screen operation is rare or absent in the September 2026 launches while computer screen operation appears only in the closed cloud agents Muse and Today, so users face a tradeoff between screen access and openness. The second evaluation assesses nanoMuse's component footprint and hosting model, showing that the Android client and relay are lightweight, the desktop client is heavier, and self-hosting remains modest in cost before model usage. Overall, the experiments indicate that nanoMuse is practical to self-host and that current commercial agents remain closed even when they offer screen control.