Abstract
Detecting general obstacles is critical for autonomous driving, especially in long-tail scenarios with rare or unseen objects. Existing methods rely on supervision or predefined categories, limiting generalization. We propose a training-free approach that combines multimodal foundation models with geometric reasoning for 3D obstacle detection. Our key idea is to detect obstacles as deviations from the road surface, segmented in 2D and localized in 3D via temporal LiDAR aggregation. The pipeline operates in a zero-shot manner without task-specific training. Experiments show accurate localization up to 100 meters and 10-25% recall gains from foundation model priors, while enabling scalable autolabeling.