← all papers · overview

Grandes Modelos De Linguagem Multimodais (mllms): Da Teoria à Prática

Abstract

Multimodal Large Language Models (MLLMs) combine the natural language understanding and generation capabilities of LLMs with perception skills in modalities such as image and audio, representing a key advancement in contemporary AI. This chapter presents the main fundamentals of MLLMs and emblematic models. Practical techniques for preprocessing, prompt engineering, and building multimodal pipelin

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).