Abstract
A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG). Despite recent progress, existing STVG studies remain largely confined to object-centric and descriptive instructions, neglecting the task-oriented reasoning that is crucial for embodied agents to accomplish goal-directed interactions. To bridge this gap, we introduce \textbf\{ToG-Bench\}, the first task-oriented spatio-temporal video grounding benchmark for egocentric videos. ToG-Bench is characterized by three key features: (1) \textbf\{Task-oriented Grounding\}, which requires identifying and localizing objects based on intended tasks rather than straightforward descriptions; (2) \textbf\{Explicit-Implicit Dual Grounding\}, where target objects can be either explicitly mentioned or implicitly inferred by contextual reasoning; (3) \textbf\{One-to-Many Grounding\}, where a single instruction may correspon