HyperAIHyperAI

Command Palette

Search for a command to run...

SuperNav: あらゆるシーンにおけるあらゆるタスクのためのエージェント型ナビゲーションシステム

Jinkai Zhang Jingyi Xu Yuanhong Yu Jiarui Guo Ruizhen Hu Hujun Bao Xiaowei Zhou Sida Peng

概要

汎用サービスロボットには、未知環境で多様な人間の要求を処理できるナビゲーションシステムが必要であり、タスク汎用性とシーン汎用性を兼ね備えることが求められる。既存手法の一部はマルチモーダル大規模言語モデル(MLLM)を微調整してナビゲーション行動を予測させるが、その挙動はナビゲーション訓練データの網羅範囲に依存し、新しい要求や環境への汎化が制限される可能性がある。我々の重要な洞察は、MLLMには要求の解釈・シーン理解・意思決定に集中させ、その汎用能力を維持しつつ、移動実行をナビゲーションツールへ委任することである。この着想を実現するため、我々はSuperNavを提案する。SuperNavは、ナビゲーション固有の微調整を行わずに、事前学習済みMLLMへ専用のエージェントハーネスを組み合わせる。本ハーネスは、ナビゲーションスキル、物理的インタラクションのためのエージェント指向ツール、タスク進捗・文脈管理によって意思決定を支援する。統一されたビジュアルポイントインターフェースが、モデルによる画像内での目的地の直接指定と実行フィードバックに基づく決定修正を可能にし、意思決定と移動を接続する。これらの構成要素は一体として、異なるタスク要件や環境にわたる持続的ナビゲーションを支える。SuperNavは、インスタンスレベルタスク、複数物体タスク、需要駆動型タスクにおいて評価した4つのベースラインを上回る。HM3Dにおけるカテゴリレベル評価と実四足歩行ロボットへの展開は、環境を横断する適用可能性をさらに示している。プロジェクトページ: https://zju3dv.github.io/SuperNav/

One-sentence Summary

Zhejiang University, Shenzhen University, and Causa Robotics propose SuperNav, an agentic navigation system that equips a pretrained MLLM with a specialized harness, Navigation Skills, agent-oriented tools, and a unified visual-point interface, enabling task- and scene-general navigation without navigation-specific fine-tuning, outperforming four baselines on instance-level, multi-object, and demand-driven tasks, and demonstrating applicability through category-level HM3D evaluation and real quadruped deployment.

Key Contributions

  • The paper presents SuperNav, a navigation harness that enables a pretrained multimodal large language model to interpret requests, understand scenes, and make navigation decisions without navigation-specific fine-tuning, while delegating motion execution to external tools.
  • The method introduces agent-oriented tools and optional Navigation Skills, supported by task-progress tracking and context management, and uses a unified visual-point interface to specify destinations in images and revise decisions from execution feedback.
  • Experimental results show that SuperNav outperforms four evaluated baselines on instance-level, multi-object, and demand-driven navigation, with further evidence from category-level evaluation on HM3D and deployment on a real quadruped robot.

Introduction

General-purpose service robots need navigation that can handle diverse requests in unfamiliar environments, from finding a specific object and visiting multiple targets to satisfying abstract needs such as finding a place to rest. These variations change how the robot interprets goals, explores, tracks progress, and decides when to stop, so both task generality and scene generality are important for practical deployment.

Prior work has limitations. Modular zero-shot navigation systems use pretrained vision-language models to measure semantic relevance, but task-specific logic still determines how exploration transitions to confirmation and when execution ends. Changing the request type often requires changing the workflow. End-to-end navigation methods fine-tune vision-language models on navigation data, but their transfer depends on training coverage and often breaks under new visual appearances, spatial layouts, and interaction conditions.

The authors introduce SuperNav, an agent harness that enables a pretrained multimodal large language model to direct navigation without navigation-specific fine-tuning. The MLLM interprets requests, understands visual observations, and reasons about actions, while the harness provides observation, motion, and task management tools, optional navigation skills, progress tracking, and context management. A unified visual-point interface lets the MLLM specify destinations directly in an image, and the navigation tool executes movement and returns feedback. The authors evaluate SuperNav on category-level, instance-level, multi-object, and demand-driven navigation, plus real-robot deployment, and report substantial improvements over action-prediction baselines.

Method

The authors present SuperNav, a system that combines a pretrained Multimodal Large Language Model (MLLM) with a navigation-specific harness to create an embodied navigation agent. This architecture sustains decisions and actions across diverse requests, ranging from specific targets to high-level needs. The MLLM handles interpretation and decision-making, while the harness organizes execution and context through a continuous agent loop. Within this framework, agent-oriented tools provide callable operations for observation, motion, and state management, and navigation skills offer procedural guidance for search, verification, and recovery.

As shown in the figure below:

The agent loop coordinates these capabilities by organizing the request, visual observations, and interaction history into a continuous navigation process. Upon initialization, the system opens a session and returns the first observation, allowing the MLLM to interpret the requested goals and completion conditions. At each subsequent step, the model uses the current context and task progress to select an operation, such as observing, moving, updating the goal state, or reading relevant skill guidance. The returns from these tools and retrieved instructions update the context for the next decision, enabling execution feedback to dynamically alter the plan.

To manage long-horizon execution, the harness maintains goal records for pending and completed items alongside a textual interaction history that retains inspected areas, candidate clues, and failed attempts. To control the growing image context, the system employs a media-only pruning strategy. It removes the media payload of an image from subsequent requests once it has been processed and superseded by newer observations, while preserving the textual history, task state, and image paths. The MLLM can still request earlier images from the workspace when necessary. Finally, the MLLM judges whether the request is satisfied or if execution is blocked, selecting a termination tool to record the outcome and close the session.

Agent-oriented tools connect the MLLM navigation decisions to executable environment operations. The suite includes initialization, observation, turning, visual-point navigation, goal-progress management, and termination. The central observation and motion interface pairs visual destinations with updated observations and execution feedback. An observation returns four RGB images labeled front, right, back, and left relative to the robot heading. These labels allow the MLLM to associate a candidate or passage with a specific view and select a destination using the PointNav function:

PointNav(d,[u,v]),u,v∈[0,1]\mathrm{PointNav} (d, [ u, v ]), \qquad u, v \in [0, 1]PointNav(d,[u,v]),u,v∈[0,1]

where ddd identifies the view and u,vu, vu,v are normalized horizontal and vertical coordinates. The tool executes the movement internally and returns new four-view images, execution status, and concise diagnostics.

This visual-point interface separates destination selection from the execution mechanism, supporting interchangeable motion backends. The geometric backend, or Geo-based Executor, uses depth, camera calibration, and pose to associate the selected point with an environment location, then plans and follows a path. The learned backend, or Learned Executor, replaces destination-image goals with marked points in reference observations. It predicts local displacement increments and a stopping signal from the selected image-point pair and recent RGB observations, repeatedly executing short predicted motion segments within a single tool call. Both backends accept visual destinations and return observations with execution feedback, ensuring the agent retains a consistent decision workflow.

Navigation skills organize reusable navigation experience into procedural guidance that the MLLM can read on demand. Each skill is structured as a Markdown package describing applicable situations, evidence to inspect, recommended steps, and linked references. The bundle includes a main navigation skill, backend-specific tool-use instructions, and references for exploration, recovery, and entrance search. The harness exposes skill names and descriptions so the MLLM can discover the relevant package and load its instructions into the decision context.

This guidance addresses three recurring decisions: where to search, how to recover, and when to finish. Search guidance prompts the MLLM to record inspected areas and visible openings, preserving alternatives for later exploration. Recovery guidance recommends changing the point or route after failed motion, stepping back when a candidate is difficult to inspect, or seeking an entrance when glass blocks access. Completion guidance calls for checking the candidate appearance, requested relations, and surrounding context before recording verified goals. These skills influence decisions by entering the context but do not programmatically enforce action sequences, leaving the final operational choices to the MLLM based on the current request and feedback.

Experiment

SuperNav is evaluated with GPT-5.6 Terra and either a Geo-based or Learned Executor across Habitat-GS instance navigation, AI2-THOR demand-driven navigation, HM3D category navigation, harness ablations, and real-world deployment, with comparisons to NaVid, UniNaVid, StreamVLN, and OmniNav. The framework improves over baselines on multi-object and inferred-goal tasks, and the Geo-based executor generally preserves success while producing more efficient paths than the Learned executor. Ablations indicate that Navigation Skills are particularly important, external grounding does not beat direct point selection, and replacing the decision model can improve success and path efficiency, while real-world tests show route revision and goal transitions.

SuperNav with the Geo-based Executor leads all evaluated methods across single-object, multi-object, and demand-driven navigation, with especially large gains in single-object tasks. Ordered multi-object navigation remains difficult for every method, with success rates falling sharply relative to single-object results. Demand-driven navigation shows clear improvements over baselines in both success and path efficiency, though long sequences remain challenging. SuperNav with Geo-based Executor achieves the best success and path efficiency in every reported benchmark setting. Single-object success more than doubles the strongest baseline, while multi-object performance drops substantially across all methods. In demand-driven navigation, SuperNav with Geo-based Executor clears the best baseline by a wide margin in both success rate and SPL.

On HM3D-OVON val-unseen, SoftNav leads the listed external methods in both success and SPL, while MTU3D is substantially lower. In HM3Dv2 comparisons, AstraNav-Memory edges out OmniNav in both success and path efficiency under the same threshold. The accompanying text reports that the Learned Executor preserves much of the Geo-based Executor's success but sacrifices path efficiency, and the Geo-based Executor remains competitive with these published baselines. SoftNav attains the highest HM3D-OVON val-unseen success and SPL among listed external results, with MTU3D notably lower. AstraNav-Memory slightly outperforms OmniNav in success rate and SPL on the shared HM3Dv2 threshold. Learned Executor keeps much of Geo-based Executor's success but is less path-efficient, while Geo-based Executor is competitive with published methods.

Ablations with the Geo-based Executor show that removing navigation skills reduces success rate more than restricting the interface to a front view, especially at the stricter distance threshold. The full SuperNav harness performs best overall, and substituting the Astra medium decision model for Terra high improves both success rate and path efficiency. Removing navigation skills lowers success rate more than front-only interaction, with the largest gap at the stricter 0.25 m threshold. The full SuperNav configuration achieves the highest success rate and SPL under both distance thresholds. Substituting the Astra medium decision model for Terra high improves both success rate and SPL while keeping the harness fixed. External grounding does not outperform the agent's direct point selection.

The experiments evaluate SuperNav with a Geo-based Executor across single-object, multi-object, and demand-driven navigation, and compare it with external methods on HM3D-OVON and HM3Dv2 benchmarks. SuperNav with the Geo-based Executor consistently leads in success and path efficiency, with single-object success more than doubling the strongest baseline, while ordered multi-object navigation remains difficult for every method. In external comparisons, SoftNav leads HM3D-OVON results and AstraNav-Memory slightly outperforms OmniNav on HM3Dv2, whereas the Learned Executor retains much of the Geo-based Executor's success but is less path-efficient. Ablations show that removing navigation skills hurts success more than front-only interaction, the full configuration performs best, and substituting the Astra medium decision model for Terra high improves both success and path efficiency.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています