OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Zhangquan, Tao, Jiale, Li, Ruihuang, Hu, Yihao, Chen, Ruitao, Yang, Zhantao, Yu, Xinlei, Jing, Haodong, Zhang, Manyuan, Shao, Shuai, Wang, Biao, Lu, Qinglin, Huang, Ruqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
di: Chen, Zhangquan, et al.
Pubblicazione: (2026)
di: Chen, Zhangquan, et al.
Pubblicazione: (2026)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
di: Wang, Yukun, et al.
Pubblicazione: (2026)
di: Wang, Yukun, et al.
Pubblicazione: (2026)
Physical oceanography during DISCOVERY cruise D210
di: Miller, Bill
Pubblicazione: (2013)
di: Miller, Bill
Pubblicazione: (2013)
UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation
di: Liao, Xinyao, et al.
Pubblicazione: (2025)
di: Liao, Xinyao, et al.
Pubblicazione: (2025)
Pigment concentrations in sea surface water from cruise DI210
di: Barlow, Raymond G
Pubblicazione: (2004)
di: Barlow, Raymond G
Pubblicazione: (2004)
Raw physical oceanography CTD data and positions from RV METEOR cruise M210
di: Walter, Maren, et al.
Pubblicazione: (2026)
di: Walter, Maren, et al.
Pubblicazione: (2026)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
di: Feng, Yigui, et al.
Pubblicazione: (2026)
di: Feng, Yigui, et al.
Pubblicazione: (2026)
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
di: Dresvyanskiy, Denis, et al.
Pubblicazione: (2024)
di: Dresvyanskiy, Denis, et al.
Pubblicazione: (2024)
Raw Miniature Autonomous Plume recorder (MAPR) data from RV METEOR cruise M210
di: Walter, Maren, et al.
Pubblicazione: (2026)
di: Walter, Maren, et al.
Pubblicazione: (2026)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
di: Yang, Baoyao, et al.
Pubblicazione: (2025)
di: Yang, Baoyao, et al.
Pubblicazione: (2025)
FocusedAD: Character-centric Movie Audio Description
di: Ye, Xiaojun, et al.
Pubblicazione: (2025)
di: Ye, Xiaojun, et al.
Pubblicazione: (2025)
A Comparative Analysis of Recurrent and Attention Architectures for Isolated Sign Language Recognition
di: Alishzade, Nigar, et al.
Pubblicazione: (2025)
di: Alishzade, Nigar, et al.
Pubblicazione: (2025)
EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation
di: Yang, Shiyuan, et al.
Pubblicazione: (2026)
di: Yang, Shiyuan, et al.
Pubblicazione: (2026)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
di: Karki, Siddhant, et al.
Pubblicazione: (2025)
di: Karki, Siddhant, et al.
Pubblicazione: (2025)
Paleomagnetic and rock magnetic characterization of ODP Leg 210 cores
di: Zhao, Xixi, et al.
Pubblicazione: (2007)
di: Zhao, Xixi, et al.
Pubblicazione: (2007)
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
di: Ma, Jie, et al.
Pubblicazione: (2024)
di: Ma, Jie, et al.
Pubblicazione: (2024)
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
di: Su, Qile, et al.
Pubblicazione: (2025)
di: Su, Qile, et al.
Pubblicazione: (2025)
Traffic-Aware Pedestrian Intention Prediction
di: Nia, Fahimeh Orvati, et al.
Pubblicazione: (2025)
di: Nia, Fahimeh Orvati, et al.
Pubblicazione: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
Polonium-210 and Lead-210 activities measured on 9 water bottle profiles during POLARSTERN cruise ANT-XXIV/3
di: Friedrich, Jana, et al.
Pubblicazione: (2011)
di: Friedrich, Jana, et al.
Pubblicazione: (2011)
HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition
di: Chen, Honghui, et al.
Pubblicazione: (2024)
di: Chen, Honghui, et al.
Pubblicazione: (2024)
Physical oceanography during ALKOR cruise AL210
di: Hinrichsen, Hans-Harald, et al.
Pubblicazione: (2014)
di: Hinrichsen, Hans-Harald, et al.
Pubblicazione: (2014)
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
di: Su, Qile, et al.
Pubblicazione: (2026)
di: Su, Qile, et al.
Pubblicazione: (2026)
Distribution of 210Pb and 210 Po in seawater at site M32_032 of the North Atlantic (Appendix)
di: Bacon, Michael P
Pubblicazione: (1977)
di: Bacon, Michael P
Pubblicazione: (1977)
Radionuclides (210Po and 210Pb) measured at station DI183_11869#19
di: Shimmield, Graham
Pubblicazione: (2004)
di: Shimmield, Graham
Pubblicazione: (2004)
Radionuclides (210Po and 210Pb) measured at station DI183_11872#40
di: Shimmield, Graham
Pubblicazione: (2004)
di: Shimmield, Graham
Pubblicazione: (2004)
Radionuclides (210Po and 210Pb) measured at station DI183_11872#51
di: Shimmield, Graham
Pubblicazione: (2004)
di: Shimmield, Graham
Pubblicazione: (2004)
Radionuclides (210Po and 210Pb) measured at station DI183_11872#91
di: Shimmield, Graham
Pubblicazione: (2004)
di: Shimmield, Graham
Pubblicazione: (2004)
Radionuclides (210Po and 210Pb) measured at station DI183_11869#35
di: Shimmield, Graham
Pubblicazione: (2004)
di: Shimmield, Graham
Pubblicazione: (2004)
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
di: Schneider, David, et al.
Pubblicazione: (2025)
di: Schneider, David, et al.
Pubblicazione: (2025)
M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark
di: Zhu, Morui, et al.
Pubblicazione: (2025)
di: Zhu, Morui, et al.
Pubblicazione: (2025)
Hydrochemistry measured on water bottle samples during ALKOR cruise AL210
di: Hinrichsen, Hans-Harald, et al.
Pubblicazione: (2014)
di: Hinrichsen, Hans-Harald, et al.
Pubblicazione: (2014)
AVControl: Efficient Framework for Training Audio-Visual Controls
di: Ben-Yosef, Matan, et al.
Pubblicazione: (2026)
di: Ben-Yosef, Matan, et al.
Pubblicazione: (2026)
Raw data of POLAR 5 campaign PAMARCMIP 2018
di: Herber, Andreas, et al.
Pubblicazione: (2019)
di: Herber, Andreas, et al.
Pubblicazione: (2019)
Distribution of 210Pb and 210 Po in seawater at site M32_023 of the North Atlantic (Appendix)
di: Bacon, Michael P
Pubblicazione: (1977)
di: Bacon, Michael P
Pubblicazione: (1977)
Distribution of 210Pb and 210 Po in seawater at site M32_015 of the North Atlantic (Appendix)
di: Bacon, Michael P
Pubblicazione: (1977)
di: Bacon, Michael P
Pubblicazione: (1977)
Distribution of 210Pb and 210 Po in seawater at site M32_021 of the North Atlantic (Appendix)
di: Bacon, Michael P
Pubblicazione: (1977)
di: Bacon, Michael P
Pubblicazione: (1977)
Documenti analoghi
-
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025) -
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
di: Chen, Zhangquan, et al.
Pubblicazione: (2026) -
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
di: Chen, Zhangquan, et al.
Pubblicazione: (2025) -
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025) -
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
di: Wang, Yukun, et al.
Pubblicazione: (2026)