Abstract
Multimodal recommender systems often suffer from “semantic distortion” when constructing item graphs via modality similarity, misaligning feature proximity with user intent. Existing behavior-based denoising methods remain limited by shallow high-order modeling or noise introduction. We propose MINERec, which introduces Interest-Tree-Guided Representation Purification. MINERec first constructs a Behavior-Aware Interest Tree (BAIT) based on item co-occurrence in user interaction histories to anchor authentic user intents. By cross-validating modality similarity against this co-occurrence structure, we identify items with “high modality similarity but zero co-occurrence” as hard negatives that represent semantic traps. Coupled with Structural-Semantic Adaptive Contrastive Learning (SACL), our method dynamically repels these negatives to filter noise at the representation level prior to graph reconstruction. Experiments on three datasets show MINERec outperforms state-of-the-art baselines, achieving an average improvement of 3.16%.