← all papers · overview

Scaling-up Perceptual Video Quality Assessment

Abstract

The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of scaling law remains unprecedented due to the scarcity of labeled resources and the insufficient scale of datasets. To address this, we propose \textbf\{OmniVQA\}, an efficient framework designed to efficiently build high-quality, human-in-the-loop VQA multi-modal instruction databases (MIDBs). We then scale up to create \textbf\{OmniVQA-Chat-400K\}, the largest MIDB in the VQA field concurrently. Our focus is on the technical and aesthetic quality dimensions, with abundant in-context instruction data to provide fine-grained VQA knowledge. Additionally, we have built the \textbf\{OmniVQA-MOS-20K\} dataset to enhance the model's quantitative quality rating capabilities. We then introduce a \textbf\{complementary\} training strategy that effectively leverages th

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).