← all papers · overview

Bi-level Prompt Optimization For Multimodal Llm-as-a-judge

Abstract

Large language models (LLMs) have become widely adopted as automated judges for evaluating AI-generated content. Despite their success, aligning LLM-based evaluations with human judgments remains challenging. While supervised fine-tuning on human-labeled data can improve alignment, it is costly and inflexible, requiring new training for each task or dataset. Recent progress in auto prompt optimiza

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).