← all papers · overview

A Review Of Multi-modal Large Language And Vision Models

Abstract

Large Language Models (LLMs) have recently emerged as a focal point of research and application, driven by their unprecedented ability to understand and generate text with human-like quality. Even more recently, LLMs have been extended into multi-modal large language models (MM-LLMs) which extends their capabilities to deal with image, video and audio information, in addition to text. This opens u

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).