← all papers · overview

The Self-improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities Without External Scaffolding?

Abstract

Self-improving large language models (LLMs) -- i.e., to improve the performance of an LLM by fine-tuning it with synthetic data generated by itself -- is a promising way to advance the capabilities of LLMs while avoiding extensive supervision. Existing approaches to self-improvement often rely on external supervision signals in the form of seed data and/or assistance from third-party models. This

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).