← all papers · overview

Any4d: Open-prompt 4D Generation From Natural Language And Images

Abstract

While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of embodied data fundamentally limit the alignment granularity between language and actions and exacerbate the challenge of long-horizon video generation--hindering generative models from achieving a \textit\{"GPT moment"\} in the embodied domain. There is a naive observation: \textit\{the diversity of embodied data far exceeds the relatively small space of possible primitive motions\}. Based on this insight, we propose \textbf\{Primitive Embodied World Models\} (PEWM), which restricts video generation to fixed shorter horizons, our approach \textit\{1) enables\} fine-grained alignment between linguistic concepts and visual representations of robotic actions, \textit\{2) reduces\} learning complexity, \textit\{3) improves\} data efficiency in embodied data co

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).