WAT: Online Video Understanding Needs Watching Before Thinking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Zifan, Sun, Hongbo, Xu, Jinglin, Tang, Canhui, Lei, Yulong, Zhang, Xuchong, Sun, Hongbin, He, Zhongjiang, Sun, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
Object-fabrication Targeted Attack for Object Detection
von: Zhang, Xuchong, et al.
Veröffentlicht: (2022)
von: Zhang, Xuchong, et al.
Veröffentlicht: (2022)
Latent Feature and Attention Dual Erasure Attack against Multi-View Diffusion Models for 3D Assets Protection
von: Sun, Jingwei, et al.
Veröffentlicht: (2024)
von: Sun, Jingwei, et al.
Veröffentlicht: (2024)
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
von: Gao, Tianyi, et al.
Veröffentlicht: (2025)
von: Gao, Tianyi, et al.
Veröffentlicht: (2025)
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
von: Wang, Lu, et al.
Veröffentlicht: (2026)
von: Wang, Lu, et al.
Veröffentlicht: (2026)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
ProTA: Probabilistic Token Aggregation for Text-Video Retrieval
von: Fang, Han, et al.
Veröffentlicht: (2024)
von: Fang, Han, et al.
Veröffentlicht: (2024)
From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
von: Cheng, Cheng, et al.
Veröffentlicht: (2025)
von: Cheng, Cheng, et al.
Veröffentlicht: (2025)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
von: Zheng, Haojie, et al.
Veröffentlicht: (2024)
von: Zheng, Haojie, et al.
Veröffentlicht: (2024)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Expand Your SCOPE: Semantic Cognition over Potential-Based Exploration for Embodied Visual Navigation
von: Wang, Ningnan, et al.
Veröffentlicht: (2025)
von: Wang, Ningnan, et al.
Veröffentlicht: (2025)
Boosting Robust AIGI Detection with LoRA-based Pairwise Training
von: Xia, Ruiyang, et al.
Veröffentlicht: (2026)
von: Xia, Ruiyang, et al.
Veröffentlicht: (2026)
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
von: Zhang, Ke, et al.
Veröffentlicht: (2025)
von: Zhang, Ke, et al.
Veröffentlicht: (2025)
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
von: Qian, Yijie, et al.
Veröffentlicht: (2025)
von: Qian, Yijie, et al.
Veröffentlicht: (2025)
Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
von: Yin, Zijin, et al.
Veröffentlicht: (2026)
von: Yin, Zijin, et al.
Veröffentlicht: (2026)
Structure-Aware Prototype Guided Trusted Multi-View Classification
von: Huang, Haojian, et al.
Veröffentlicht: (2025)
von: Huang, Haojian, et al.
Veröffentlicht: (2025)
PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
von: Yuan, Kun, et al.
Veröffentlicht: (2024)
Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification
von: Zhang, Quan, et al.
Veröffentlicht: (2026)
von: Zhang, Quan, et al.
Veröffentlicht: (2026)
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026)
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026)
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding
von: Gastager, David, et al.
Veröffentlicht: (2025)
von: Gastager, David, et al.
Veröffentlicht: (2025)
Watch and Learn: Learning to Use Computers from Online Videos
von: Song, Chan Hee, et al.
Veröffentlicht: (2025)
von: Song, Chan Hee, et al.
Veröffentlicht: (2025)
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
von: Wei, Runpu, et al.
Veröffentlicht: (2025)
von: Wei, Runpu, et al.
Veröffentlicht: (2025)
How Can Objects Help Video-Language Understanding?
von: Tang, Zitian, et al.
Veröffentlicht: (2025)
von: Tang, Zitian, et al.
Veröffentlicht: (2025)
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
TempCompass: Do Video LLMs Really Understand Videos?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2024)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2024)
Detailed Object Description with Controllable Dimensions
von: Wang, Xinran, et al.
Veröffentlicht: (2024)
von: Wang, Xinran, et al.
Veröffentlicht: (2024)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
von: Sun, Shitong, et al.
Veröffentlicht: (2026)
von: Sun, Shitong, et al.
Veröffentlicht: (2026)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion Prior
von: Bai, Weimin, et al.
Veröffentlicht: (2025)
von: Bai, Weimin, et al.
Veröffentlicht: (2025)
Disentangle and denoise: Tackling context misalignment for video moment retrieval
von: Ma, Kaijing, et al.
Veröffentlicht: (2024)
von: Ma, Kaijing, et al.
Veröffentlicht: (2024)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
von: Bai, Purui, et al.
Veröffentlicht: (2026)
von: Bai, Purui, et al.
Veröffentlicht: (2026)
Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification
von: Huang, Haojian, et al.
Veröffentlicht: (2024)
von: Huang, Haojian, et al.
Veröffentlicht: (2024)
Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains
von: Tang, Zitian, et al.
Veröffentlicht: (2023)
von: Tang, Zitian, et al.
Veröffentlicht: (2023)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Cambrian-P: Pose-Grounded Video Understanding
von: Yang, Jihan, et al.
Veröffentlicht: (2026)
von: Yang, Jihan, et al.
Veröffentlicht: (2026)
Watch Before You Answer: Learning from Visually Grounded Post-Training
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
von: Tang, Canhui, et al.
Veröffentlicht: (2025) -
Object-fabrication Targeted Attack for Object Detection
von: Zhang, Xuchong, et al.
Veröffentlicht: (2022) -
Latent Feature and Attention Dual Erasure Attack against Multi-View Diffusion Models for 3D Assets Protection
von: Sun, Jingwei, et al.
Veröffentlicht: (2024) -
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
von: Gao, Tianyi, et al.
Veröffentlicht: (2025) -
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
von: Wang, Lu, et al.
Veröffentlicht: (2026)