E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nie, Jiahao, An, Wenbin, Zhang, Gongjie, Xu, Yicheng, Tan, Yap-Peng, Kot, Alex C., Lu, Shijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation
von: Nie, Jiahao, et al.
Veröffentlicht: (2026)
von: Nie, Jiahao, et al.
Veröffentlicht: (2026)
Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence Mining
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
Boosting SAM for Cross-Domain Few-Shot Segmentation via Conditional Point Sparsification
von: Nie, Jiahao, et al.
Veröffentlicht: (2026)
von: Nie, Jiahao, et al.
Veröffentlicht: (2026)
SimBase: A Simple Baseline for Temporal Video Grounding
von: Bao, Peijun, et al.
Veröffentlicht: (2024)
von: Bao, Peijun, et al.
Veröffentlicht: (2024)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton Matching
von: Yang, Siyuan, et al.
Veröffentlicht: (2023)
von: Yang, Siyuan, et al.
Veröffentlicht: (2023)
EventFly: Event Camera Perception from Ground to the Sky
von: Kong, Lingdong, et al.
Veröffentlicht: (2025)
von: Kong, Lingdong, et al.
Veröffentlicht: (2025)
Backdoor Attacks against No-Reference Image Quality Assessment Models via a Scalable Trigger
von: Yu, Yi, et al.
Veröffentlicht: (2024)
von: Yu, Yi, et al.
Veröffentlicht: (2024)
Purify Unlearnable Examples via Rate-Constrained Variational Autoencoders
von: Yu, Yi, et al.
Veröffentlicht: (2024)
von: Yu, Yi, et al.
Veröffentlicht: (2024)
Color Space Learning for Cross-Color Person Re-Identification
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching
von: Liu, Lanmiao, et al.
Veröffentlicht: (2026)
von: Liu, Lanmiao, et al.
Veröffentlicht: (2026)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
von: An, Wenbin, et al.
Veröffentlicht: (2025)
von: An, Wenbin, et al.
Veröffentlicht: (2025)
Towards Model Resistant to Transferable Adversarial Examples via Trigger Activation
von: Yu, Yi, et al.
Veröffentlicht: (2025)
von: Yu, Yi, et al.
Veröffentlicht: (2025)
MTL-UE: Learning to Learn Nothing for Multi-Task Learning
von: Yu, Yi, et al.
Veröffentlicht: (2025)
von: Yu, Yi, et al.
Veröffentlicht: (2025)
Robust and Transferable Backdoor Attacks Against Deep Image Compression With Selective Frequency Prior
von: Yu, Yi, et al.
Veröffentlicht: (2024)
von: Yu, Yi, et al.
Veröffentlicht: (2024)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
Unlearnable Examples Detection via Iterative Filtering
von: Yu, Yi, et al.
Veröffentlicht: (2024)
von: Yu, Yi, et al.
Veröffentlicht: (2024)
SparseCoop: Cooperative Perception with Kinematic-Grounded Queries
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering
von: Xing, Yun, et al.
Veröffentlicht: (2026)
von: Xing, Yun, et al.
Veröffentlicht: (2026)
Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding
von: Woo, Jongbhin, et al.
Veröffentlicht: (2024)
von: Woo, Jongbhin, et al.
Veröffentlicht: (2024)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection
von: Lu, Hui, et al.
Veröffentlicht: (2025)
von: Lu, Hui, et al.
Veröffentlicht: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild
von: Bao, Peijun, et al.
Veröffentlicht: (2024)
von: Bao, Peijun, et al.
Veröffentlicht: (2024)
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
von: Jiang, Xueying, et al.
Veröffentlicht: (2026)
von: Jiang, Xueying, et al.
Veröffentlicht: (2026)
Event-Based Visual Odometry on Non-Holonomic Ground Vehicles
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
von: Lu, Lidong, et al.
Veröffentlicht: (2025)
von: Lu, Lidong, et al.
Veröffentlicht: (2025)
Visual Grounding from Event Cameras
von: Kong, Lingdong, et al.
Veröffentlicht: (2025)
von: Kong, Lingdong, et al.
Veröffentlicht: (2025)
Modeling Continuous Motion for 3D Point Cloud Object Tracking
von: Luo, Zhipeng, et al.
Veröffentlicht: (2023)
von: Luo, Zhipeng, et al.
Veröffentlicht: (2023)
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
On the Generalization Capacities of MLLMs for Spatial Intelligence
von: Zhang, Gongjie, et al.
Veröffentlicht: (2026)
von: Zhang, Gongjie, et al.
Veröffentlicht: (2026)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
von: Nie, Jiahao, et al.
Veröffentlicht: (2024) -
Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation
von: Nie, Jiahao, et al.
Veröffentlicht: (2026) -
Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence Mining
von: Nie, Jiahao, et al.
Veröffentlicht: (2024) -
Boosting SAM for Cross-Domain Few-Shot Segmentation via Conditional Point Sparsification
von: Nie, Jiahao, et al.
Veröffentlicht: (2026) -
SimBase: A Simple Baseline for Temporal Video Grounding
von: Bao, Peijun, et al.
Veröffentlicht: (2024)