← all papers · overview

CMMR-VLN: Vision-and-language Navigation Via Continual Multimodal Memory Retrieval

Abstract

Although large language models (LLMs) are introduced into vision-and-language navigation (VLN) to improve instruction comprehension and generalization, existing LLM- based VLN lacks the ability to selectively recall and use relevant priori experiences to help navigation tasks, limiting their performance in long-horizon and unfamiliar scenarios. In this work, we propose CMMR-VLN (Continual Multimod

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).