Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
Fuente:
arXiv
Saved in:
| Main Authors: | Meng, Jiahao, Li, Xiangtai, Wang, Haochen, Tan, Yue, Zhang, Tao, Kong, Lingdong, Tong, Yunhai, Wang, Anran, Teng, Zhiyang, Wang, Yujing, Wang, Zhuochen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
by: Meng, Jiahao, et al.
Published: (2026)
by: Meng, Jiahao, et al.
Published: (2026)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
Music Grounding by Short Video
by: Xin, Zijie, et al.
Published: (2024)
by: Xin, Zijie, et al.
Published: (2024)
Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction
by: Wang, Dali, et al.
Published: (2026)
by: Wang, Dali, et al.
Published: (2026)
GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing
by: Lin, Zihao, et al.
Published: (2026)
by: Lin, Zihao, et al.
Published: (2026)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
by: Meng, Yiran, et al.
Published: (2025)
by: Meng, Yiran, et al.
Published: (2025)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
by: Tong, Xinyi, et al.
Published: (2025)
by: Tong, Xinyi, et al.
Published: (2025)
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
by: Fang, Pengcheng, et al.
Published: (2026)
by: Fang, Pengcheng, et al.
Published: (2026)
Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web
by: Heo, Ryang, et al.
Published: (2026)
by: Heo, Ryang, et al.
Published: (2026)
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
by: Wang, Feng, et al.
Published: (2024)
by: Wang, Feng, et al.
Published: (2024)
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
by: Yang, Dingyi, et al.
Published: (2024)
by: Yang, Dingyi, et al.
Published: (2024)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
by: Pramanick, Shraman, et al.
Published: (2025)
by: Pramanick, Shraman, et al.
Published: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
by: Qu, Mengxue, et al.
Published: (2024)
by: Qu, Mengxue, et al.
Published: (2024)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
by: Zeng, Xiangyu, et al.
Published: (2024)
by: Zeng, Xiangyu, et al.
Published: (2024)
Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding
by: Wang, Jieyi, et al.
Published: (2026)
by: Wang, Jieyi, et al.
Published: (2026)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
by: Luu, Thanh-Danh, et al.
Published: (2025)
by: Luu, Thanh-Danh, et al.
Published: (2025)
StreamOptix: A Cross-layer Adaptive Video Delivery Scheme
by: Liu, Mufan, et al.
Published: (2024)
by: Liu, Mufan, et al.
Published: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
by: Qin, You, et al.
Published: (2024)
by: Qin, You, et al.
Published: (2024)
Compression Metadata-assisted RoI Extraction and Adaptive Inference for Efficient Video Analytics
by: Wang, Chengzhi, et al.
Published: (2025)
by: Wang, Chengzhi, et al.
Published: (2025)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering
by: Xu, Zhaoyang, et al.
Published: (2026)
by: Xu, Zhaoyang, et al.
Published: (2026)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Consistency-aware Fake Videos Detection on Short Video Platforms
by: Wang, Junxi, et al.
Published: (2025)
by: Wang, Junxi, et al.
Published: (2025)
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
by: Wang, Junxi, et al.
Published: (2025)
by: Wang, Junxi, et al.
Published: (2025)
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
by: Yu, Jiashuo, et al.
Published: (2025)
by: Yu, Jiashuo, et al.
Published: (2025)
PRVR: Partially Relevant Video Retrieval
by: Chen, Xianke, et al.
Published: (2022)
by: Chen, Xianke, et al.
Published: (2022)
Exposure Completing for Temporally Consistent Neural High Dynamic Range Video Rendering
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
by: Pu, Junfu, et al.
Published: (2026)
by: Pu, Junfu, et al.
Published: (2026)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
CLIPRerank: An Extremely Simple Method for Improving Ad-hoc Video Search
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR Headsets
by: Jin, Yili, et al.
Published: (2024)
by: Jin, Yili, et al.
Published: (2024)
Interactive $360^{\circ}$ Video Streaming Using FoV-Adaptive Coding with Temporal Prediction
by: Mao, Yixiang, et al.
Published: (2024)
by: Mao, Yixiang, et al.
Published: (2024)
Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation
by: Yan, Mingxuan, et al.
Published: (2024)
by: Yan, Mingxuan, et al.
Published: (2024)
Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction
by: Yuan, Xiang, et al.
Published: (2025)
by: Yuan, Xiang, et al.
Published: (2025)
Similar Items
-
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
by: Meng, Jiahao, et al.
Published: (2026) -
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025) -
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025) -
Music Grounding by Short Video
by: Xin, Zijie, et al.
Published: (2024) -
Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction
by: Wang, Dali, et al.
Published: (2026)