← all papers · overview

Training-time Action Conditioning For Efficient Real-time Chunking

Abstract

Real-time chunking (RTC) enables vision-language-action models (VLAs) to generate smooth, reactive robot trajectories by asynchronously predicting action chunks and conditioning on previously committed actions via inference-time inpainting. However, this inpainting method introduces computational overhead that increases inference latency. In this work, we propose a simple alternative: simulating inference delay at training time and conditioning on action prefixes directly, eliminating any inference-time overhead. Our method requires no modifications to the model architecture or robot runtime, and can be implemented with only a few additional lines of code. In simulated experiments, we find that training-time RTC outperforms inference-time RTC at higher inference delays. In real-world experiments on box building and espresso making tasks with the VLA, we demonstrate that training-time RTC maintains both task performance and speed parity with inference-time RTC while bein

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).