VLABench
Emerging7papers using it
2025first seen
VLABench is a benchmark dataset used to evaluate the performance of Vision-Language-Action models in detecting execution failures during robotic task execution.
Papers using VLABench (7)
- Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime MonitoringBridge-WA: Predicting Where and How the World Changes for Robotic ActionRevisiting Embodied Chain-of-Thought for Generalizable Robot ManipulationDIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action ModelE0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete DiffusionFrom Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation