EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Ahn, Geo, Han, Jiwook, Kim, Youngrae, Lee, Joonseok, Choi, Jinwoo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026)
by: Han, Jiwook, et al.
Published: (2026)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
by: Bae, Kyungho, et al.
Published: (2023)
by: Bae, Kyungho, et al.
Published: (2023)
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024)
by: Lee, Jongseo, et al.
Published: (2024)
Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding
by: Ren, Junlong, et al.
Published: (2025)
by: Ren, Junlong, et al.
Published: (2025)
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
by: Lee, Jin-Seop, et al.
Published: (2025)
by: Lee, Jin-Seop, et al.
Published: (2025)
Infusing Environmental Captions for Long-Form Video Language Grounding
by: Lee, Hyogun, et al.
Published: (2024)
by: Lee, Hyogun, et al.
Published: (2024)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
by: Lee, Inseo, et al.
Published: (2025)
by: Lee, Inseo, et al.
Published: (2025)
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
by: Yue, Feng, et al.
Published: (2025)
by: Yue, Feng, et al.
Published: (2025)
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
by: Ran, Ran, et al.
Published: (2026)
by: Ran, Ran, et al.
Published: (2026)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
by: Zheng, Zelin, et al.
Published: (2026)
by: Zheng, Zelin, et al.
Published: (2026)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025)
by: Bae, Kyungho, et al.
Published: (2025)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023)
by: Lee, Dongho, et al.
Published: (2023)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
by: Lee, Seunghun, et al.
Published: (2025)
by: Lee, Seunghun, et al.
Published: (2025)
Class-Continuous Conditional Generative Neural Radiance Field
by: Kim, Jiwook, et al.
Published: (2023)
by: Kim, Jiwook, et al.
Published: (2023)
Feature Augmentation based Test-Time Adaptation
by: Cho, Younggeol, et al.
Published: (2024)
by: Cho, Younggeol, et al.
Published: (2024)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
by: Fan, Rong, et al.
Published: (2026)
by: Fan, Rong, et al.
Published: (2026)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
by: Jung, Minjoon, et al.
Published: (2026)
by: Jung, Minjoon, et al.
Published: (2026)
VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
by: Lee, Eunsoo, et al.
Published: (2026)
by: Lee, Eunsoo, et al.
Published: (2026)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding
by: Kang, Minseok, et al.
Published: (2025)
by: Kang, Minseok, et al.
Published: (2025)
Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
Moment Quantization for Video Temporal Grounding
by: Sun, Xiaolong, et al.
Published: (2025)
by: Sun, Xiaolong, et al.
Published: (2025)
Universal Domain Adaptation for Semantic Segmentation
by: Choe, Seun-An, et al.
Published: (2025)
by: Choe, Seun-An, et al.
Published: (2025)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
by: Lv, Guannan, et al.
Published: (2026)
by: Lv, Guannan, et al.
Published: (2026)
Towards Long-Form Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2026)
by: Gu, Xin, et al.
Published: (2026)
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
by: T V, Sethuraman, et al.
Published: (2026)
by: T V, Sethuraman, et al.
Published: (2026)
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
by: Moon, WonJun, et al.
Published: (2023)
by: Moon, WonJun, et al.
Published: (2023)
Few-Shot Adaptation of Grounding DINO for Agricultural Domain
by: Singh, Rajhans, et al.
Published: (2025)
by: Singh, Rajhans, et al.
Published: (2025)
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
by: Lee, Seonho, et al.
Published: (2024)
by: Lee, Seonho, et al.
Published: (2024)
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
by: Li, Jiaqi, et al.
Published: (2026)
by: Li, Jiaqi, et al.
Published: (2026)
Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations
by: Kim, Minung, et al.
Published: (2025)
by: Kim, Minung, et al.
Published: (2025)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
Similar Items
-
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026) -
DEVIAS: Learning Disentangled Video Representations of Action and Scene
by: Bae, Kyungho, et al.
Published: (2023) -
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024) -
Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding
by: Ren, Junlong, et al.
Published: (2025) -
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
by: Lee, Jin-Seop, et al.
Published: (2025)