ChronusOmni: Improving Time Awareness of Omni Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yijing, Wu, Yihan, Guan, Kaisi, Ren, Yuchen, Wang, Yuyue, Song, Ruihua, Ru, Liyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917137316052992
author Chen, Yijing
Wu, Yihan
Guan, Kaisi
Ren, Yuchen
Wang, Yuyue
Song, Ruihua
Ru, Liyun
author_facet Chen, Yijing
Wu, Yihan
Guan, Kaisi
Ren, Yuchen
Wang, Yuyue
Song, Ruihua
Ru, Liyun
contents Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target vision-language scenarios and focus on the explicit temporal grounding questions, such as identifying when a visual event occurs or determining what event happens at aspecific time. However, they often make insufficient use of the audio modality, and overlook implicit temporal grounding across modalities--for example, identifying what is visually present when a character speaks, or determining what is said when a visual event occurs--despite such cross-modal temporal relations being prevalent in real-world scenarios. In this paper, we propose ChronusOmni, an omni large language model designed to enhance temporal awareness for both explicit and implicit audiovisual temporal grounding. First, we interleave text-based timestamp tokens with visual and audio representations at each time unit, enabling unified temporal modeling across modalities. Second, to enforce correct temporal ordering and strengthen fine-grained temporal reasoning, we incorporate reinforcement learning with specially designed reward functions. Moreover, we construct ChronusAV, a temporally-accurate, modality-complete, and cross-modal-aligned dataset to support the training and evaluation on audiovisual temporal grounding task. Experimental results demonstrate that ChronusOmni achieves state-of-the-art performance on ChronusAV with more than 30% improvement and top results on most metrics upon other temporal grounding benchmarks. This highlights the strong temporal awareness of our model across modalities, while preserving general video and audio understanding capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09841
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChronusOmni: Improving Time Awareness of Omni Large Language Models
Chen, Yijing
Wu, Yihan
Guan, Kaisi
Ren, Yuchen
Wang, Yuyue
Song, Ruihua
Ru, Liyun
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target vision-language scenarios and focus on the explicit temporal grounding questions, such as identifying when a visual event occurs or determining what event happens at aspecific time. However, they often make insufficient use of the audio modality, and overlook implicit temporal grounding across modalities--for example, identifying what is visually present when a character speaks, or determining what is said when a visual event occurs--despite such cross-modal temporal relations being prevalent in real-world scenarios. In this paper, we propose ChronusOmni, an omni large language model designed to enhance temporal awareness for both explicit and implicit audiovisual temporal grounding. First, we interleave text-based timestamp tokens with visual and audio representations at each time unit, enabling unified temporal modeling across modalities. Second, to enforce correct temporal ordering and strengthen fine-grained temporal reasoning, we incorporate reinforcement learning with specially designed reward functions. Moreover, we construct ChronusAV, a temporally-accurate, modality-complete, and cross-modal-aligned dataset to support the training and evaluation on audiovisual temporal grounding task. Experimental results demonstrate that ChronusOmni achieves state-of-the-art performance on ChronusAV with more than 30% improvement and top results on most metrics upon other temporal grounding benchmarks. This highlights the strong temporal awareness of our model across modalities, while preserving general video and audio understanding capabilities.
title ChronusOmni: Improving Time Awareness of Omni Large Language Models
topic Computation and Language
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2512.09841