← all papers · overview

Menatqa: A New Dataset For Testing The Temporal Comprehension And Reasoning Abilities Of Large Language Models

Abstract

Large language models (LLMs) have shown nearly saturated performance on many natural language processing (NLP) tasks. As a result, it is natural for people to believe that LLMs have also mastered abilities such as time understanding and reasoning. However, research on the temporal sensitivity of LLMs has been insufficiently emphasized. To fill this gap, this paper constructs Multiple Sensitive Fac

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).