HyperAIHyperAI

Command Palette

Search for a command to run...

Agent
LLM

nanoMuse: Ein Open-Source-Personalagent für jedes Gerät, das Sie besitzen

Guangyi Liu Yong Liu Jiangning Zhang

Zusammenfassung

Assistenten von 2011 antworteten und warteten, und Agenten von 2023 erledigten eine Aufgabe und hörten auf. Im September 2026 zeigte Metas Muse einen Agenten für eine Person, mit Konten, Geräten, Gedächtnis und einer dauerhaften Konversation – geschlossen, in der Cloud eines Anbieters, in einem einzigen Land. Von einem solchen Agenten wird erwartet, dass er im Namen einer Person auf deren Konten und Geräte zugreift, sich über Wochen an sie erinnert, von sich aus spricht, wenn es sich lohnt, und Rechenschaft über sein Handeln ablegt. Es handelt sich um eine Art Software, nicht um ein Modell, und bislang gab es kein offenes Gegenstück. Dieser Bericht definiert den persönlichen Agenten anhand von fünf Fragen und drei Horizonten. Er rekonstruiert, wie Muse aufgebaut ist, aus Metas öffentlichen Angaben und einer Kopie seines Produktions-Prompts, wobei jede Aussage mit ihrer Quelle gekennzeichnet ist. Anschließend stellt er nanoMuse vor, das quelloffene Gegenstück unter der GPL-3.0: ein Agent auf jedem Gerät, das eine Person besitzt, mit Händen auf dem Bildschirm des Telefons und dem des Computers. Sie teilen eine gemeinsame Konversation über ein Relay, das jede Person betreiben kann; jede Aktion durchläuft einen Sentinel, das Gedächtnis besteht aus Dateien, die die Person lesen kann, und das Modell ist frei wählbar. Seine Größe und Kosten werden als Schätzungen angegeben. Was offen ist – Gedächtnis mit Provenienz, eine Evaluierungssuite für die Hände und ein offenes Modell dafür – wird als Roadmap dargelegt.

One-sentence Summary

Researchers from Zhejiang University propose nanoMuse, an open-source GPL-3.0 personal agent that runs on every device a person owns and shares one conversation through a self-hostable relay with screen control, sentinel-gated actions, readable memory files, and user-chosen models, as the open counterpart to Meta's closed Muse.

Key Contributions

  • The report defines the personal agent through five questions and three horizons, tracing the lineage from the assistants of 2011 to the agents of September 2026.
  • It reconstructs how Meta’s Muse is built from Meta’s public record and a copy of its production prompt, with each statement marked by its source.
  • It introduces nanoMuse, an open-source GPL-3.0 counterpart that runs on every device a person owns, with hands on the phone’s screen and the computer’s, a shared conversation relay anyone can run, a Sentinel over every action, memory as files, and model choice. Size and cost are given as estimates, and an open roadmap covers memory with provenance, an evaluation suite for the hands, and an open model for them.

Introduction

The authors trace the shift from task-oriented assistants to personal agents, culminating in the September 2026 wave led by Meta's Muse. A personal agent acts on one person's accounts, devices, and files over time and must be accountable, but commercial offerings like Muse run as closed, vendor-hosted Linux VMs with only a small gadget SDK opened. Prior agents either handled single tasks in vendor clouds or, if open, ran on user infrastructure without full screen control, and cloud-only agents cannot reach everyday mobile apps that lack APIs or web pages. To address this, the authors present nanoMuse, an open personal agent that runs on hardware the user already owns, with hands on phone and computer screens, a Sentinel over actions, memory as files, and reported size and cost.

Method

The authors design Muse as a background personal agent structured across three distinct security zones: the surfaces of the user, a dedicated cloud environment, and isolated host services.

In this framework, each user is allocated a dedicated Linux virtual machine in the cloud, which is replaced approximately hourly while maintaining a persistent home directory. The agent and its tools operate within a container bounded by seccomp filters and reduced capabilities. The system prompt establishes a strict persona that works exclusively for the user, enforcing discretion and hard safety stops. The agent utilizes tools organized in namespaces, such as browser, chat, memory, and device control. A critical component of this architecture is the Sentinel, which acts as the sole permission authority for connector actions and network egress. The agent holds only surrogate tokens, which are swapped for real credentials at the network boundary after approval. The system employs kernel level taint tracking, where a tool process starts clean and becomes tainted upon reading user data, ensuring that only clean requests within a narrow policy pass without a prompt. Approvals are presented directly to the client as scoped capabilities rather than passing through the conversation.

In contrast, the authors develop nanoMuse as a decentralized personal agent where every device operates as a peer carrying a complete agent and its own Sentinel.

The Android implementation runs a complete agent inside the application, featuring a sandboxed Alpine Linux environment, a shell, a browser, and MCP servers. The desktop application utilizes an Electron shell around a harness with a Python runtime. Unlike the cloud hosted approach, nanoMuse places the agent directly on the device, granting it direct access to the phone screen. The Sentinel in this system functions as a policy boundary within the same process family, enforcing a fixed decision order that evaluates denied tools, user rules, allow and ask lists, data taint, and tool risk. Approvals are issued as scoped grants, and sensitive actions trigger mandatory warnings that cannot be bypassed. For screen interaction, the system employs a hybrid approach, climbing a ladder from skills and command line tools to a browser, and finally to direct screen manipulation via a one action per screenshot loop. Memory is managed as plain Markdown files that the user can open and edit, ensuring data portability and transparency. The architecture includes an optional and open relay that handles sign in, model call ledgers, and device synchronization without storing the actual content of the conversations.

The authors outline a structured development pipeline to advance the system capabilities across memory, device integration, and physical world interaction.

The immediate focus involves closing the feature gap by improving memory recall, reducing hand interaction failures, and enabling single command self hosting. The subsequent phase introduces an evaluation suite built from failed runs to score releases objectively, alongside the training of a small open model on user contributed traces to reduce vendor dependency. The long term vision extends the hands of the agent to the physical world, integrating cameras, appliances, and robot arms under the same approval and logging mechanisms, ultimately aiming to build an agent with memory and identity that survive across different models, devices, and vendors.

Experiment

This section outlines nanoMuse’s community roadmap in three horizons: near-term work focuses on memory, proactivity, conversation continuity, hand reliability, and one-command self-hosting; next steps introduce an evaluation suite built from failed runs and benchmarked with AndroidWorld, OSWorld, MemGUI-Bench, and OS-Harm, alongside a small open model for the hands trained on contributed traces to avoid vendor dependence; later work extends the hands to physical devices and long-lived agents with durable memory and identity. The roadmap concludes by inviting contributions such as skills, translations, devices, failing traces, and corrections.

The comparison spans open projects from mid-2025 and 2026 plus a cluster of commercial agents launched in September 2026. Phone screen operation is rare and absent from the September launches, while computer screen operation appears only in two cloud-based systems, Muse and Today. Open licensing is concentrated in the earlier and self-hosted projects, whereas the commercial cloud agents are closed. Only OpenMinis can operate a phone screen, and only on Android; none of the September 2026 agents has phone screen access. Muse and Today are the only listed systems with computer screen operation; OpenMuse and Manus Cue run loops without screen control. Open source systems generally run on the user's own computer, server, or app, while the commercial launches run in per-person cloud computers or virtual machines. Among the September systems there is a tradeoff: Muse and Today offer computer screen use but are closed, while OpenMuse is open but lacks screen operation.

nanoMuse's footprint varies widely by component: the Android client is compact, the desktop client is substantially larger to download and install, and the relay is lightweight enough for a single small server. Estimated self-hosting costs are modest, on the order of a few US dollars per month outside mainland China or tens of yuan inside it, plus model usage. Model call pricing is low enough that ordinary daily chat costs only a few fen. The Android app is the smallest release artifact, while the desktop app is far heavier to download and also uses about half a gigabyte of memory at idle on Linux. The relay is the lightest runtime component, idling at roughly 85 MB of memory and sized for a server with 1 vCPU and 1 GB. Self-hosting a relay is estimated at US4to4 to4to6 per month outside mainland China or ¥30 to ¥60 inside it before model usage. The hands model costs more per token than the chat model, but a typical day of chat remains a few fen.

The first comparison examines screen-control capabilities and licensing across open and commercial agents, finding that phone screen operation is rare or absent in the September 2026 launches while computer screen operation appears only in the closed cloud agents Muse and Today, so users face a tradeoff between screen access and openness. The second evaluation assesses nanoMuse's component footprint and hosting model, showing that the Android client and relay are lightweight, the desktop client is heavier, and self-hosting remains modest in cost before model usage. Overall, the experiments indicate that nanoMuse is practical to self-host and that current commercial agents remain closed even when they offer screen control.


KI mit KI entwickeln

Von der Idee bis zum Launch – beschleunigen Sie Ihre KI-Entwicklung mit kostenlosem KI-Co-Coding, sofort einsatzbereiter Umgebung und bestem GPU-Preis.

KI-gestütztes kollaboratives Programmieren
Sofort einsatzbereite GPUs
Die besten Preise

HyperAI Newsletters

Abonnieren Sie unsere neuesten Updates
Wir werden die neuesten Updates der Woche in Ihren Posteingang liefern um neun Uhr jeden Montagmorgen
Unterstützt von MailChimp