RoboTwin 2.0
Emerging10papers using it
2025first seen
RoboTwin~2.0 is a benchmark integrated into the StarVLA codebase that is used to evaluate Vision-Language-Action (VLA) approaches in the context of developing generalist embodied agents.
Papers using RoboTwin 2.0 (10)
- Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought SupervisionRotVLA: Rotational Latent Action for Vision-Language-Action ModelFrom Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action ModelStarVLA: A Lego-like Codebase for Vision-Language-Action Model DevelopingVLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image ReasoningJEPA-VLA: Video Predictive Embedding is Needed for VLA ModelsUniversal Pose Pretraining for Generalizable Vision-Language-Action PoliciesPEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual ManipulationMM-ACT: Learn from Multimodal Parallel Generation to ActSimpleVLA-RL: Scaling VLA Training via Reinforcement Learning