← all papers · overview

ERA: Transforming Vlms Into Embodied Agents Via Embodied Prior Learning And Online Reinforcement Learning

Abstract

Recent advances in embodied AI highlight the potential of vision language models (VLMs) as agents capable of perception, reasoning, and interaction in complex environments. However, top-performing systems rely on large-scale models that are costly to deploy, while smaller VLMs lack the necessary knowledge and skills to succeed. To bridge this gap, we present \textit\{Embodied Reasoning Agent (ERA)\}, a two-stage framework that integrates prior knowledge learning and online reinforcement learning (RL). The first stage, \textit\{Embodied Prior Learning\}, distills foundational knowledge from three types of data: (1) Trajectory-Augmented Priors, which enrich existing trajectory data with structured reasoning generated by stronger models; (2) Environment-Anchored Priors, which provide in-environment knowledge and grounding supervision; and (3) External Knowledge Priors, which transfer general knowledge from out-of-environment datasets. In the second stage, we develop an online RL pipeline

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).