LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
LVSum is a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment. It comprises 72 diverse videos (avg. 16 min) across 13 domains, each with up to 10 human-generated summaries containing temporal references. Evaluation reveals that transcripts contribute more than visual frames, a significant gap remains between model and human summaries, and MLLMs exhibit weaknesses in temporal grounding, instruction adherence, and cross-modal coherence.
content type paperpublished July 2026
LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
AuthorsAlkesh Patel*, Melis Ozyildirim*, Ying-Chang Cheng, Ganesh Nagarajan
View publication
View source code (GitHub)
Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. We introduce LVSum, a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment. LVSum comprises 72 diverse videos spanning 13 domains with an average duration of 16 minutes, each annotated with up to 10 human-generated summaries containing temporal references. We conduct a comprehensive evaluation of leading proprietary and open-source MLLMs using newly introduced LLM-based metrics for content relevance and modality coherence, alongside standard automatic metrics. Our experiments reveal three key findings: (1) transcripts contribute substantially more to summarization quality than visual frames alone, (2) a significant performance gap persists between model-generated and human-written summaries, and (3) current MLLMs exhibit systematic weaknesses in temporal grounding, instruction adherence, and cross-modal coherence.
- Equal contribution
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
July 9, 2026research area Computer Vision, research area Methods and Algorithms
Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning depends on the correct ordering and evolution of events. This deficiency stems in part from training objectives that fail to explicitly reward temporal reasoning and instead rely on frame-level spatial shortcuts. To address this limitation, we propose…
Read more
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
January 6, 2026research area Computer Vision, research area Data Science and Annotationconference ECCV
Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored. True narrative understanding requires grounding who is doing what, when, and where, maintaining coherent entity representations across dynamic visual and temporal contexts. We introduce NarrativeTrack, the first benchmark to evaluate narrative…
Read more