← all papers · overview

Pixrec: Leveraging Visual Context For Next-item Prediction In Sequential Recommendation

Abstract

Large Language Models (LLMs) have recently shown strong potential for usage in sequential recommendation tasks through text-only models, which combine advanced prompt design, contrastive alignment, and fine-tuning on downstream domain-specific data. While effective, these approaches overlook the rich visual information present in many real-world recommendation scenarios, particularly in e-commerce. This paper proposes PixRec - a vision-language framework that incorporates both textual attributes and product images into the recommendation pipeline. Our architecture leverages a vision-language model backbone capable of jointly processing image-text sequences, maintaining a dual-tower structure and mixed training objective while aligning multi-modal feature projections for both item-item and user-item interactions. Using the Amazon Reviews dataset augmented with product images, our experiments demonstrate and 40% improvements in top-rank and top-10 rank accuracy over text-only

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).