← all papers · overview

FLAME: Learning To Navigate With Multimodal LLM In Urban Environments

Abstract

Large Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation tasks, yielding suboptimal performance compared to specialized VLN models. We introduce FLAME (FLAMingo-Architected Embodied Agent), a novel Multimodal LLM-base

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).