← all papers · overview

Query-opt: Optimizing Inference Of Large Language Models Via Multi-query Instructions In Meeting Summarization

Abstract

This work focuses on the task of query-based meeting summarization in which the summary of a context (meeting transcript) is generated in response to a specific query. When using Large Language Models (LLMs) for this task, usually a new call to the LLM inference endpoint/API is triggered for each new query, even if the context stays the same. However, repeated calls to the LLM inference endpoints

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).