Command Palette
Search for a command to run...
nanoMuse: وكيل شخصي مفتوح المصدر لكل جهاز تملكه
nanoMuse: وكيل شخصي مفتوح المصدر لكل جهاز تملكه
Guangyi Liu Yong Liu Jiangning Zhang
الملخص
المساعدون الذين ظهروا في 2011 كانوا يجيبون وينتظرون، والوكلاء الذين ظهروا في 2023 كانوا ينجزون مهمة ثم يتوقفون. في سبتمبر 2026، عرضت Muse من Meta وكيلًا لشخص واحد يتضمن حسابات وأجهزة وذاكرة ومحادثة تدوم، لكنه مغلق، في سحابة مزوّد الخدمة، وفي بلد واحد. ويُتوقع من هذا الوكيل أن يعمل على حسابات الشخص وأجهزته، وأن يتذكرها عبر أسابيع، وأن يبادر بالحديث عندما يكون ذلك مجديًا، وأن يقدم حسابًا عما فعله. إنه نوع من البرمجيات لا نموذج، ولم يكن له حتى الآن نظير مفتوح. يعرّف هذا التقرير الوكيل الشخصي في خمسة أسئلة وثلاثة آفاق. ويقرأ كيفية بناء Muse من السجل العام لـ Meta ومن نسخة من موجّه الإنتاج الخاص به، مع تمييز كل عبارة بمصدرها. ثم يقدم nanoMuse، النظير المفتوح المصدر بموجب GPL-3.0، وكيلًا واحدًا على كل جهاز يملكه الشخص، وله يدان على شاشة الهاتف وشاشة الحاسوب. وتتشارك هذه الوكلاء محادثة واحدة عبر مرحّل يمكن لأي شخص تشغيله؛ ويمر كل إجراء عبر Sentinel، والذاكرة عبارة عن ملفات يمكن للشخص قراءتها، والنموذج باختياره. وتُعرض تقديرات حجمه وتكلفته. أما ما هو مفتوح — ذاكرة موثقة المصدر، ومجموعة تقييم للأيدي، ونموذج مفتوح لها — فيُحدد بوصفه خارطة طريق.
One-sentence Summary
Researchers from Zhejiang University propose nanoMuse, an open-source GPL-3.0 personal agent that runs on every device a person owns and shares one conversation through a self-hostable relay with screen control, sentinel-gated actions, readable memory files, and user-chosen models, as the open counterpart to Meta's closed Muse.
Key Contributions
- The report defines the personal agent through five questions and three horizons, tracing the lineage from the assistants of 2011 to the agents of September 2026.
- It reconstructs how Meta’s Muse is built from Meta’s public record and a copy of its production prompt, with each statement marked by its source.
- It introduces nanoMuse, an open-source GPL-3.0 counterpart that runs on every device a person owns, with hands on the phone’s screen and the computer’s, a shared conversation relay anyone can run, a Sentinel over every action, memory as files, and model choice. Size and cost are given as estimates, and an open roadmap covers memory with provenance, an evaluation suite for the hands, and an open model for them.
Introduction
The authors trace the shift from task-oriented assistants to personal agents, culminating in the September 2026 wave led by Meta's Muse. A personal agent acts on one person's accounts, devices, and files over time and must be accountable, but commercial offerings like Muse run as closed, vendor-hosted Linux VMs with only a small gadget SDK opened. Prior agents either handled single tasks in vendor clouds or, if open, ran on user infrastructure without full screen control, and cloud-only agents cannot reach everyday mobile apps that lack APIs or web pages. To address this, the authors present nanoMuse, an open personal agent that runs on hardware the user already owns, with hands on phone and computer screens, a Sentinel over actions, memory as files, and reported size and cost.
Method
The authors design Muse as a background personal agent structured across three distinct security zones: the surfaces of the user, a dedicated cloud environment, and isolated host services.
In this framework, each user is allocated a dedicated Linux virtual machine in the cloud, which is replaced approximately hourly while maintaining a persistent home directory. The agent and its tools operate within a container bounded by seccomp filters and reduced capabilities. The system prompt establishes a strict persona that works exclusively for the user, enforcing discretion and hard safety stops. The agent utilizes tools organized in namespaces, such as browser, chat, memory, and device control. A critical component of this architecture is the Sentinel, which acts as the sole permission authority for connector actions and network egress. The agent holds only surrogate tokens, which are swapped for real credentials at the network boundary after approval. The system employs kernel level taint tracking, where a tool process starts clean and becomes tainted upon reading user data, ensuring that only clean requests within a narrow policy pass without a prompt. Approvals are presented directly to the client as scoped capabilities rather than passing through the conversation.
In contrast, the authors develop nanoMuse as a decentralized personal agent where every device operates as a peer carrying a complete agent and its own Sentinel.
The Android implementation runs a complete agent inside the application, featuring a sandboxed Alpine Linux environment, a shell, a browser, and MCP servers. The desktop application utilizes an Electron shell around a harness with a Python runtime. Unlike the cloud hosted approach, nanoMuse places the agent directly on the device, granting it direct access to the phone screen. The Sentinel in this system functions as a policy boundary within the same process family, enforcing a fixed decision order that evaluates denied tools, user rules, allow and ask lists, data taint, and tool risk. Approvals are issued as scoped grants, and sensitive actions trigger mandatory warnings that cannot be bypassed. For screen interaction, the system employs a hybrid approach, climbing a ladder from skills and command line tools to a browser, and finally to direct screen manipulation via a one action per screenshot loop. Memory is managed as plain Markdown files that the user can open and edit, ensuring data portability and transparency. The architecture includes an optional and open relay that handles sign in, model call ledgers, and device synchronization without storing the actual content of the conversations.
The authors outline a structured development pipeline to advance the system capabilities across memory, device integration, and physical world interaction.
The immediate focus involves closing the feature gap by improving memory recall, reducing hand interaction failures, and enabling single command self hosting. The subsequent phase introduces an evaluation suite built from failed runs to score releases objectively, alongside the training of a small open model on user contributed traces to reduce vendor dependency. The long term vision extends the hands of the agent to the physical world, integrating cameras, appliances, and robot arms under the same approval and logging mechanisms, ultimately aiming to build an agent with memory and identity that survive across different models, devices, and vendors.
Experiment
This section outlines nanoMuse’s community roadmap in three horizons: near-term work focuses on memory, proactivity, conversation continuity, hand reliability, and one-command self-hosting; next steps introduce an evaluation suite built from failed runs and benchmarked with AndroidWorld, OSWorld, MemGUI-Bench, and OS-Harm, alongside a small open model for the hands trained on contributed traces to avoid vendor dependence; later work extends the hands to physical devices and long-lived agents with durable memory and identity. The roadmap concludes by inviting contributions such as skills, translations, devices, failing traces, and corrections.
The comparison spans open projects from mid-2025 and 2026 plus a cluster of commercial agents launched in September 2026. Phone screen operation is rare and absent from the September launches, while computer screen operation appears only in two cloud-based systems, Muse and Today. Open licensing is concentrated in the earlier and self-hosted projects, whereas the commercial cloud agents are closed. Only OpenMinis can operate a phone screen, and only on Android; none of the September 2026 agents has phone screen access. Muse and Today are the only listed systems with computer screen operation; OpenMuse and Manus Cue run loops without screen control. Open source systems generally run on the user's own computer, server, or app, while the commercial launches run in per-person cloud computers or virtual machines. Among the September systems there is a tradeoff: Muse and Today offer computer screen use but are closed, while OpenMuse is open but lacks screen operation.
nanoMuse's footprint varies widely by component: the Android client is compact, the desktop client is substantially larger to download and install, and the relay is lightweight enough for a single small server. Estimated self-hosting costs are modest, on the order of a few US dollars per month outside mainland China or tens of yuan inside it, plus model usage. Model call pricing is low enough that ordinary daily chat costs only a few fen. The Android app is the smallest release artifact, while the desktop app is far heavier to download and also uses about half a gigabyte of memory at idle on Linux. The relay is the lightest runtime component, idling at roughly 85 MB of memory and sized for a server with 1 vCPU and 1 GB. Self-hosting a relay is estimated at US4to6 per month outside mainland China or ¥30 to ¥60 inside it before model usage. The hands model costs more per token than the chat model, but a typical day of chat remains a few fen.
The first comparison examines screen-control capabilities and licensing across open and commercial agents, finding that phone screen operation is rare or absent in the September 2026 launches while computer screen operation appears only in the closed cloud agents Muse and Today, so users face a tradeoff between screen access and openness. The second evaluation assesses nanoMuse's component footprint and hosting model, showing that the Android client and relay are lightweight, the desktop client is heavier, and self-hosting remains modest in cost before model usage. Overall, the experiments indicate that nanoMuse is practical to self-host and that current commercial agents remain closed even when they offer screen control.