← all papers · overview

INSIGHT: Inference-time Sequence Introspection For Generating Help Triggers In Vision-language-action Models

Abstract

Recent Vision-Language-Action (VLA) models show strong generalization capabilities, yet they lack introspective mechanisms for anticipating failures and requesting help from a human supervisor. We present \textbf\{INSIGHT\}, a learning framework for leveraging token-level uncertainty signals to predict when a VLA should request help. Using -FAST as the underlying model, we extract per-token *entropy*, *log-probability*, and Dirichlet-based estimates of *aleatoric and epistemic uncertainty*, and train compact transformer classifiers to map these sequences to help triggers. We explore supervision regimes for strong or weak supervision, and extensively compare them across in-distribution and out-of-distribution tasks. Our results show a trade-off: strong labels enable models to capture fine-grained uncertainty dynamics for reliable help detection, while weak labels, though noisier, still support competitive introspection when training and evaluation are aligned, offering a scalab

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).