ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Fangxu, Lu, Ziyao, Niu, Liqiang, Meng, Fandong, Zhou, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
von: Chen, Wenchao, et al.
Veröffentlicht: (2024)
von: Chen, Wenchao, et al.
Veröffentlicht: (2024)
Self-supervised Representation Learning for Cell Event Recognition through Time Arrow Prediction
von: Chen, Cangxiong, et al.
Veröffentlicht: (2024)
von: Chen, Cangxiong, et al.
Veröffentlicht: (2024)
Seeing the Arrow of Time in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
von: Omasa, Takamitsu, et al.
Veröffentlicht: (2025)
von: Omasa, Takamitsu, et al.
Veröffentlicht: (2025)
Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs
von: Han, Peitao, et al.
Veröffentlicht: (2026)
von: Han, Peitao, et al.
Veröffentlicht: (2026)
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
D2C: Unlocking the Potential of Continuous Autoregressive Image Generation with Discrete Tokens
von: Wang, Panpan, et al.
Veröffentlicht: (2025)
von: Wang, Panpan, et al.
Veröffentlicht: (2025)
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
von: Lan, Zhibin, et al.
Veröffentlicht: (2025)
von: Lan, Zhibin, et al.
Veröffentlicht: (2025)
Continuous Visual Autoregressive Generation via Score Maximization
von: Shao, Chenze, et al.
Veröffentlicht: (2025)
von: Shao, Chenze, et al.
Veröffentlicht: (2025)
Can Visual Encoder Learn to See Arrows?
von: Terashita, Naoyuki, et al.
Veröffentlicht: (2025)
von: Terashita, Naoyuki, et al.
Veröffentlicht: (2025)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?
von: Gesny, Enoal, et al.
Veröffentlicht: (2026)
von: Gesny, Enoal, et al.
Veröffentlicht: (2026)
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
von: Zhang, Hui, et al.
Veröffentlicht: (2026)
von: Zhang, Hui, et al.
Veröffentlicht: (2026)
ArrowPose: Segmentation, Detection, and 5 DoF Pose Estimation Network for Colorless Point Clouds
von: Hagelskjaer, Frederik
Veröffentlicht: (2025)
von: Hagelskjaer, Frederik
Veröffentlicht: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
von: Li, Yifei, et al.
Veröffentlicht: (2025)
von: Li, Yifei, et al.
Veröffentlicht: (2025)
Object-Shot Enhanced Grounding Network for Egocentric Video
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
Frozen Vision Transformers for Dense Prediction on Small Datasets: A Case Study in Arrow Localization
von: Shepherd, Maxwell
Veröffentlicht: (2026)
von: Shepherd, Maxwell
Veröffentlicht: (2026)
Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding
von: Tu, Xuezhen, et al.
Veröffentlicht: (2026)
von: Tu, Xuezhen, et al.
Veröffentlicht: (2026)
EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
von: Fu, Zhiheng, et al.
Veröffentlicht: (2026)
von: Fu, Zhiheng, et al.
Veröffentlicht: (2026)
EVDI++: Event-based Video Deblurring and Interpolation via Self-Supervised Learning
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Revisit Event Generation Model: Self-Supervised Learning of Event-to-Video Reconstruction with Implicit Neural Representations
von: Wang, Zipeng, et al.
Veröffentlicht: (2024)
von: Wang, Zipeng, et al.
Veröffentlicht: (2024)
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding
von: Deng, Andong, et al.
Veröffentlicht: (2024)
von: Deng, Andong, et al.
Veröffentlicht: (2024)
Learning Event Completeness for Weakly Supervised Video Anomaly Detection
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
von: Yan, Pengyu, et al.
Veröffentlicht: (2026)
von: Yan, Pengyu, et al.
Veröffentlicht: (2026)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining
von: Peng, Bo, et al.
Veröffentlicht: (2026)
von: Peng, Bo, et al.
Veröffentlicht: (2026)
TimeRewind: Rewinding Time with Image-and-Events Video Diffusion
von: Chen, Jingxi, et al.
Veröffentlicht: (2024)
von: Chen, Jingxi, et al.
Veröffentlicht: (2024)
VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
von: Dong, Lu, et al.
Veröffentlicht: (2025)
von: Dong, Lu, et al.
Veröffentlicht: (2025)
Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
von: Zhao, Jinghan, et al.
Veröffentlicht: (2025)
von: Zhao, Jinghan, et al.
Veröffentlicht: (2025)
Exploiting Auxiliary Caption for Video Grounding
von: Li, Hongxiang, et al.
Veröffentlicht: (2023)
von: Li, Hongxiang, et al.
Veröffentlicht: (2023)
Harnessing Object Grounding for Time-Sensitive Video Understanding
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
von: Chen, Wenchao, et al.
Veröffentlicht: (2024) -
Self-supervised Representation Learning for Cell Event Recognition through Time Arrow Prediction
von: Chen, Cangxiong, et al.
Veröffentlicht: (2024) -
Seeing the Arrow of Time in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2025) -
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
von: Omasa, Takamitsu, et al.
Veröffentlicht: (2025) -
Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs
von: Han, Peitao, et al.
Veröffentlicht: (2026)