← all papers · overview

Selectllm: Can Llms Select Important Instructions To Annotate?

Abstract

Instruction tuning benefits from large and diverse datasets; however, creating such datasets involves a high cost of human labeling. While synthetic datasets generated by large language models (LLMs) have partly solved this issue, they often contain low-quality data. One effective solution is selectively annotating unlabelled instructions, especially given the relative ease of acquiring unlabeled

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).