HyperAIHyperAI

Command Palette

Search for a command to run...

PanoVLN: Auf dem Weg zu effektiver panoramischer Vision-und-Sprache-Navigation

Zhen Wang Changpeng Wang Zhe Liu Zhangyang Qi Yuxiang Lu Zimo Zeng Donglian Qi Xi Chen

Zusammenfassung

Jüngste Vision-Sprache-Modelle (VLMs) haben die Vision-und-Sprache-Navigation (VLN) vorangebracht und ermöglichen es Modellen, Navigationsaktionen aus visuellen Beobachtungen und Sprachanweisungen vorherzusagen. In dieser Arbeit untersuchen wir VLN mit panoramischen Beobachtungen und stellen PanoVLN vor. Die Motivation ist naheliegend: Ein vollständigerer visueller Kontext sollte fundiertere Navigationsentscheidungen ermöglichen. Beispielsweise kann ein Panorama einen Durchgang außerhalb des Sichtfelds einer perspektivischen Kamera zeigen, sodass das Modell die beabsichtigte Route ohne zusätzliche Exploration erkennen kann. Wir stellen jedoch fest, dass das bloße Ersetzen perspektivischer Bilder durch Panoramen nur begrenzte Verbesserungen bringt. Unsere Diagnose legt nahe, dass die vollständige Nutzung der größeren Sichtbarkeit Änderungen an der Aktionsvorhersage, der Trainingssupervision und der visuellen Repräsentation erfordert. Erstens unterstützt eine größere Sichtbarkeit eine längerfristige Aktionsplanung. Wir lassen das Modell längere Aktionssequenzen vorhersagen, wodurch größere Drehungen und anschließende Bewegungen aus einem einzigen Panorama möglich werden. Konkret führen wir eine konfidenzgesteuerte Ausführungsstrategie (CGE) ein, die dynamisch bestimmt, wie viele vorhergesagte Aktionen vor einer Neuplanung ausgeführt werden. Zweitens bringt eine größere Sichtbarkeit auch komplexere Routenwahlmöglichkeiten mit sich. Daher konstruieren wir Trainingsrouten mit häufigen Verzweigungspunkten und klaren Anweisungen, um eine gezielte Supervision für die Routenauswahl bereitzustellen. Drittens erfordert panoramische Navigation ein Verständnis räumlicher Beziehungen über Blickrichtungen hinweg, das über das Erkennen einzelner Landmarken hinausgeht. Wir kombinieren semantische und geometrische Merkmale aus RGB-Panoramen, um sowohl den Szeneninhalt als auch das räumliche Layout zu erfassen, ohne zusätzliche visuelle Token hinzuzufügen. Mit einem 4B-Backbone und reinem RGB-Input übertrifft PanoVLN den bisherigen SOTA um 11,9 % bzw. 8,7 % Erfolgsrate auf R2R-CE und RxR-CE Val-Unseen. Realweltexperimente auf einem Vierbeiner zeigen darüber hinaus eine schnellere Navigation mit weniger Pausen als frühere VLN-Methoden. Unsere Projektseite ist verfügbar unter https://wangzhen-w.github.io/PanoVLN/.

One-sentence Summary

Researchers from Zhejiang University and The University of Hong Kong propose PanoVLN, a panoramic vision-and-language navigation model that uses confidence-guided execution for longer-horizon actions, branching-point training supervision, and combined semantic-geometric RGB panorama features, surpassing the previous state of the art by 11.9%11.9\%11.9% and 8.7%8.7\%8.7% in success rate on R2R-CE and RxR-CE Val-Unseen and enabling faster real-world quadruped navigation with fewer pauses.

Key Contributions

  • PanoVLN is a panoramic vision-and-language navigation method that predicts longer action sequences from a single panorama and uses confidence-guided execution to determine how many predicted actions to execute before replanning.
  • A decision-centric training dataset of 98K trajectories across 800 HM3D scenes provides frequent branching points and visually grounded instructions, with denser sampling around turns and stopping points for route selection and completion supervision.
  • The method combines semantic VLM features with geometric PanoVGGT features from the same RGB panorama; with a 4B RGB-only backbone, it achieves 77.3% success on R2R-CE Val-Unseen and 78.0% on RxR-CE Val-Unseen, improving over the previous state of the art by 11.9% and 8.7% and enabling faster real-world quadruped navigation with fewer pauses.

Introduction

Vision-and-language navigation (VLN) requires an agent to follow natural-language instructions through an environment, and recent vision-language models have improved this task by predicting navigation actions from visual observations. Most prior work uses perspective images, which limit the visual context available at each decision point. The authors investigate equirectangular panoramas, which provide a 360-degree view and can reveal passages, landmarks, and route alternatives. They find that simply replacing perspective images with panoramas does not improve performance under the same setup, so they propose PanoVLN, which adapts action prediction with longer horizons and confidence-guided execution, creates a 98K-trajectory decision-centric training dataset, and fuses semantic VLM features with panoramic geometric features. This approach achieves state-of-the-art success rates on R2R-CE and RxR-CE and enables faster real-world robot navigation with fewer pauses.

Dataset

Dataset composition and sources

  • The authors construct 98K navigation trajectories across 800 HM3D scenes.
  • Trajectories are designed to contain frequent branching points, where the agent must choose among multiple visible traversable paths.
  • Each trajectory is paired with an instruction that identifies the chosen path and stopping location.

Route construction and filtering

  • Walkable space is divided into connected areas using the navigation mesh.
  • A branching point is defined as having at least two visible, traversable paths to different areas, excluding the incoming path.
  • Endpoints are sampled in different areas, and routes passing through branching points are retained.
  • Rendering-quality checks remove candidates with mesh holes or incomplete geometry, followed by near-duplicate removal.
  • An expert converts remaining routes into primitive action sequences.
  • Replay verifies goal reachability and confirms that the chosen path and its alternatives are visible at each branching point.

Instruction construction and verification

  • Trajectories are divided into travel, branching, and arrival segments.
  • Travel segments use first-person video with the expert path marked on the ground.
  • Branching and arrival segments additionally use eight-view compass images.
  • Qwen3.8-27B describes movement, identifies the chosen path from visible cues, and specifies the stopping location.
  • Descriptions are combined in route order, with repetition removed and wording refined.
  • Verification uses clean videos and compass images without instruction or route overlays.
  • Three checks are applied: motion consistency, choice grounding, and stop grounding.
  • Mismatched segments are revised locally and reverified; only samples passing all three checks are retained.

Training sample construction and usage

  • The data is used to provide supervision for learning path selection from panoramic observations.
  • For H = 18, the authors use a stride-six grid to reduce overlap between adjacent grid targets from 17 to 12 actions.
  • Additional states are added at sustained-turn onsets and near termination to supervise turning and stopping.
  • Each state is paired with its H-step expert action sequence.
  • The grid preserves route coverage, while added states emphasize action transitions.
  • The provided section does not specify train/eval split or mixture ratios.

Method

The authors propose a method to fully exploit the complete visual context provided by panoramas in the Vision-and-Language Navigation task. Starting from a baseline model that takes panoramas as input, they introduce three key adaptations: longer action-sequence supervision and execution, decision-centric data construction, and geometry-aware visual representations.

To leverage the wider visibility of panoramic observations, the authors extend the action prediction horizon. Instead of predicting a single step, the policy is trained to predict a sequence of the next HHH expert actions. The training objective minimizes the negative log-likelihood of the expert action sequence using teacher forcing:

Lact=−1H∑i=1Hlog⁡pθ(at,i∗∣Ot,At,<i∗)\mathcal{L}_{\mathrm{act}} = - \frac{1}{H} \sum_{i=1}^{H} \log p_{\theta} \left(a_{t,i}^{*} \mid \mathcal{O}_{t}, \mathbf{A}_{t,<i}^{*}\right)Lact​=−H1​i=1∑H​logpθ​(at,i∗​∣Ot​,At,<i∗​)

where At,<i∗\mathbf{A}_{t,<i}^{*}At,<i∗​ contains the preceding expert actions and pθp_{\theta}pθ​ is the VLM next-token distribution.

During inference, the execution length is adapted based on prediction uncertainty through a mechanism called Confidence-Guided Execution. The uncertainty for a generated action is defined as the negative log probability of the predicted action. As shown in the figure below:

Mean uncertainty rises after an initial dip and exhibits substantial variation across policy calls. To handle this, the execution mechanism extends the executed prefix as long as the cumulative uncertainty Ut(k)U_t(k)Ut​(k) remains within a predefined budget BBB, ensuring at least Emin⁡E_{\min}Emin​ actions are executed:

Et=max⁡{k∈{1,…,H}:k≤Emin⁡ or Ut(k)≤B}E_{t} = \max \left\{k \in \{1, \dots, H\}: k \leq E_{\min} \text{ or } U_{t}(k) \leq B \right\}Et​=max{k∈{1,…,H}:k≤Emin​ or Ut​(k)≤B}

The agent executes this prefix and then reobserves the environment unless it predicts a stop action.

To provide better supervision for selecting the correct path among multiple visible options, the authors construct a decision-centric training dataset. They generate navigation trajectories across various scenes, specifically targeting branching points where at least two traversable paths are visible. After filtering out routes with rendering issues or near-duplicates, an expert converts the valid routes into primitive action sequences. Instructions are constructed by dividing trajectories into travel, branching, and arrival segments. A large language model describes the movement and identifies the chosen path using first-person video and compass images. The authors rigorously verify motion consistency and choice grounding, revising and retaining only the segments that pass all checks. To reduce overlap between adjacent training states, they employ a stride-six grid for sampling and add specific states at sustained-turn onsets and near termination to emphasize action transitions.

Finally, the authors develop a geometry-aware visual representation to better understand the spatial relationships within a panoramic observation. They allocate a larger number of tokens to the current equirectangular panorama and fewer tokens to each history frame. To fuse geometric information without adding extra visual tokens, a pretrained PanoVGGT encoder extracts geometric features from the current RGB panorama. These features are resampled in ERP coordinates and grouped to align with the VLM merged current tokens. A trainable MLP projects the aligned geometric groups into the visual-token embedding space for residual fusion:

Vˉt=Vt+αfψ(Gt)\bar{V}_{t} = V_{t} + \alpha f_{\psi}(G_{t})Vˉt​=Vt​+αfψ​(Gt​)

where α\alphaα is a fixed residual scale. This fusion combines semantics and geometry from corresponding ERP regions while preserving the token count and order, ultimately conditioning the action prediction alongside the instruction and history tokens. The VLM and projection layers are trained jointly, while the geometry encoder remains frozen.

Experiment

Experiments evaluate RGB-only PanoVLN on R2R-CE and RxR-CE Val-Unseen splits in Matterport3D using Habitat, with metrics including navigation error, success rate, SPL, and nDTW. Simulation results show PanoVLN achieves state-of-the-art success rates on both benchmarks and benefits from panoramic context, while real-world tests on a Unitree Go2 across hallway, office, and campus settings demonstrate reliable indoor route following, outdoor transfer, and efficient execution through longer predicted segments with confidence-guided execution. Ablations confirm that longer ERP prediction horizons, turn- and termination-aware sampling, confidence-guided execution, PanoVGGT panoramic features, and decision-centric training trajectories all improve navigation and stopping behavior.

PanoVLN achieves the highest success rate and SPL on both R2R-CE and RxR-CE, setting a new state of the art by large margins over prior panoramic navigation methods. A restricted-data version of PanoVLN also leads its training-data group in success rate and SPL. The results connect panoramic context and decision-centric trajectories to better generalization in unseen scenes. PanoVLN surpasses previous best success rates by 11.9 percentage points on R2R-CE and 8.7 percentage points on RxR-CE. PanoVLN trained without navigation data beyond R2R-CE and RxR-CE still leads the restricted-data group in SR and SPL on both benchmarks. Adding decision-centric trajectories further improves performance in unseen scenes.

Across the real-world routes, PanoVLN achieves the shortest navigation duration and the highest travel speed among all compared methods. It also spends the least time waiting for policy responses and records far fewer pauses and policy calls. Although its per-request inference latency is not the lowest, its overall execution flow is the most efficient. PanoVLN combines the shortest navigation duration with the highest travel speed, while JanusVLN is the slowest by a wide margin. PanoVLN has the lowest waiting fraction, fewest pauses, and fewest policy calls, despite not having the lowest per-request latency.

Compared with random-start sampling, the proposed turn- and termination-aware sampling lowers navigation error while keeping oracle success nearly unchanged. It also increases success rate and SPL by a clear margin, indicating more reliable termination and route completion when turn and stop states receive stronger supervision. The proposed sampling reduces navigation error relative to random-start sampling while maintaining comparable oracle success. It improves success rate and SPL, suggesting more reliable termination in the goal region and better supervision of turn and stop states.

Confidence-guided execution achieves the best navigation performance on both benchmarks, outperforming fixed execution lengths and random execution. Fixed execution at one action is the strongest fixed setting, while longer fixed horizons generally reduce success and path quality. The results indicate that adapting execution to model confidence is more reliable than committing to a preset or random number of actions. Confidence-guided execution records the lowest navigation error and the highest success and path-quality metrics on both benchmarks. Among fixed strategies, one-action execution is strongest, and longer fixed horizons degrade success and SPL, especially on RxR-CE. Random execution over one to eighteen actions underperforms confidence-guided execution and trails the one-action fixed strategy on RxR-CE.

Geometry encoder comparison under matched fusion and execution settings shows mixed effects for existing encoders. PanoVGGT achieves the highest success rate on both R2R-CE and RxR-CE, and also leads on R2R-CE oracle success and SPL and on RxR-CE nDTW. The results suggest its panoramic geometric features add spatial cues that support route selection. PanoVGGT leads all compared encoders in success rate on both benchmarks, with the best R2R-CE oracle success and SPL and the best RxR-CE nDTW. Alternative encoders have mixed effects: UniK3D improves R2R-CE navigation error and SPL but not RxR-CE success, while DA^2 and DAP tend to reduce success on both benchmarks.

The experiments benchmark PanoVLN on R2R-CE and RxR-CE, real-world navigation routes, and ablations of sampling, execution, and geometry encoders. PanoVLN sets new state-of-the-art success and SPL on both benchmarks by large margins, and its real-world runs achieve the shortest duration and highest speed with fewer pauses and policy calls despite not having the lowest per-request latency. Turn- and termination-aware sampling and confidence-guided execution improve success and path quality over random or fixed baselines, while among geometry encoders PanoVGGT gives the strongest overall results and other encoders yield mixed effects.


KI mit KI entwickeln

Von der Idee bis zum Launch – beschleunigen Sie Ihre KI-Entwicklung mit kostenlosem KI-Co-Coding, sofort einsatzbereiter Umgebung und bestem GPU-Preis.

KI-gestütztes kollaboratives Programmieren
Sofort einsatzbereite GPUs
Die besten Preise

HyperAI Newsletters

Abonnieren Sie unsere neuesten Updates
Wir werden die neuesten Updates der Woche in Ihren Posteingang liefern um neun Uhr jeden Montagmorgen
Unterstützt von MailChimp