← all papers · overview

RPRA: Predicting An Llm-judge For Efficient But Performant Inference

Abstract

Large language models (LLMs) face a fundamental trade-off between computational efficiency (e.g., number of parameters) and output quality, especially when deployed on computationally limited devices such as phones or laptops. One way to address this challenge is by following the example of humans and have models ask for help when they believe they

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).