← all papers · overview

Revisiting The Test-time Scaling Of O1-like Models: Do They Truly Possess Test-time Scaling Capabilities?

Abstract

The advent of test-time scaling in large language models (LLMs), exemplified by OpenAI's o1 series, has advanced reasoning capabilities by scaling computational resource allocation during inference. While successors like QwQ, Deepseek-R1 (R1) and LIMO replicate these advancements, whether these models truly possess test-time scaling capabilities remains underexplored. This study found that longer

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).