Vision-Language Models
loadingβ¦
loadingβ¦
Vision-Language Models is one of the most active areas in Awesome Multimodal β 9,603 papers in this collection, evaluated on datasets like COCO, VQA, LIBERO. A strong starting point is "Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation".