← all papers · overview

Learning From "silly" Questions Improves Large Language Models, But Only Slightly

Abstract

Constructing high-quality Supervised Fine-Tuning (SFT) datasets is critical for the training of large language models (LLMs). Recent studies have shown that using data from a specific source, Ruozhiba, a Chinese website where users ask "silly" questions to better understand certain topics, can lead to better fine-tuning performance. This paper aims to explore some hidden factors: the potential int

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).