← all papers · overview

Video-em: Event-centric Episodic Memory For Long-form Video Understanding

Abstract

Video Large Language Models (Video-LLMs) have shown strong video understanding, yet their application to long-form videos remains constrained by limited context windows. A common workaround is to compress long videos into a handful of representative frames via retrieval or summarization. However, most existing pipelines score frames in isolation, implicitly assuming that frame-level saliency is sufficient for downstream reasoning. This often yields redundant selections, fragmented temporal evidence, and weakened narrative grounding for long-form video question answering. We present \textbf\{Video-EM\}, a training-free, event-centric episodic memory framework that reframes long-form VideoQA as *episodic event construction* followed by *memory refinement*. Instead of treating retrieved keyframes as independent visuals, Video-EM employs an LLM as an active memory agent to orchestrate off-the-shelf tools: it first localizes query-relevant moments via multi-grained semantic matching, then g

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).