← all papers · overview

Resonance Rope: Improving Context Length Generalization Of Large Language Models

Abstract

This paper addresses the challenge of train-short-test-long (TSTL) scenarios in Large Language Models (LLMs) equipped with Rotary Position Embedding (RoPE), where models pre-trained on shorter sequences face difficulty with out-of-distribution (OOD) token positions in longer sequences. We introduce Resonance RoPE, a novel approach designed to narrow the generalization gap in TSTL scenarios by refi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).