← all papers · overview

Sequential-niah: A Needle-in-a-haystack Benchmark For Extracting Sequential Needles From Long Contexts

Abstract

Evaluating the ability of large language models (LLMs) to process lengthy contexts is critical, especially for retrieving query-relevant information embedded within them. We introduce Sequential-NIAH, a benchmark specifically designed to evaluate the capability of LLMs to extract sequential information items (known as *needles*) from long contexts. The benchmark includes three needle generation pi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).