MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Zhiyi, Wu, Xiaoyu, Liu, Zihao, Yang, Linlin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2026)
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2026)
Rethinking Metrics and Benchmarks of Video Anomaly Detection
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions
von: Liu, Zihao, et al.
Veröffentlicht: (2026)
von: Liu, Zihao, et al.
Veröffentlicht: (2026)
Language-guided Open-world Video Anomaly Detection under Weak Supervision
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
von: Dong, Xinxin, et al.
Veröffentlicht: (2025)
von: Dong, Xinxin, et al.
Veröffentlicht: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
von: Zhu, Zhiyi, et al.
Veröffentlicht: (2025)
von: Zhu, Zhiyi, et al.
Veröffentlicht: (2025)
ME-Mamba: Multi-Expert Mamba with Efficient Knowledge Capture and Fusion for Multimodal Survival Analysis
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2025)
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2025)
CVA: Context-aware Video-text Alignment for Video Temporal Grounding
von: Moon, Sungho, et al.
Veröffentlicht: (2026)
von: Moon, Sungho, et al.
Veröffentlicht: (2026)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
Optical Flow Representation Alignment Mamba Diffusion Model for Medical Video Generation
von: Wang, Zhenbin, et al.
Veröffentlicht: (2024)
von: Wang, Zhenbin, et al.
Veröffentlicht: (2024)
M4V: Multi-Modal Mamba for Text-to-Video Generation
von: Huang, Jiancheng, et al.
Veröffentlicht: (2025)
von: Huang, Jiancheng, et al.
Veröffentlicht: (2025)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
von: An, Joungbin, et al.
Veröffentlicht: (2025)
von: An, Joungbin, et al.
Veröffentlicht: (2025)
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
von: He, Allen, et al.
Veröffentlicht: (2026)
von: He, Allen, et al.
Veröffentlicht: (2026)
Multi-Modal Video Feature Extraction for Popularity Prediction
von: Liu, Haixu, et al.
Veröffentlicht: (2025)
von: Liu, Haixu, et al.
Veröffentlicht: (2025)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery
von: Hao, Pengfei, et al.
Veröffentlicht: (2025)
von: Hao, Pengfei, et al.
Veröffentlicht: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
von: Liu, Wenqi, et al.
Veröffentlicht: (2026)
von: Liu, Wenqi, et al.
Veröffentlicht: (2026)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
Context-Gated Cross-Modal Perception with Visual Mamba for PET-CT Lung Tumor Segmentation
von: Ayllón, Elena Mulero, et al.
Veröffentlicht: (2025)
von: Ayllón, Elena Mulero, et al.
Veröffentlicht: (2025)
GFE-Mamba: Mamba-based AD Multi-modal Progression Assessment via Generative Feature Extraction from MCI
von: Fang, Zhaojie, et al.
Veröffentlicht: (2024)
von: Fang, Zhaojie, et al.
Veröffentlicht: (2024)
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
von: Xing, Yifei, et al.
Veröffentlicht: (2024)
von: Xing, Yifei, et al.
Veröffentlicht: (2024)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
von: Hua, Hang, et al.
Veröffentlicht: (2024)
von: Hua, Hang, et al.
Veröffentlicht: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
von: Qin, You, et al.
Veröffentlicht: (2024)
von: Qin, You, et al.
Veröffentlicht: (2024)
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
von: Zhang, Chunhui, et al.
Veröffentlicht: (2023)
von: Zhang, Chunhui, et al.
Veröffentlicht: (2023)
Unsupervised Audio-Visual Segmentation with Modality Alignment
von: Bhosale, Swapnil, et al.
Veröffentlicht: (2024)
von: Bhosale, Swapnil, et al.
Veröffentlicht: (2024)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment
von: Bi, Xiaowei, et al.
Veröffentlicht: (2025)
von: Bi, Xiaowei, et al.
Veröffentlicht: (2025)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
End-to-End Multi-Modal Diffusion Mamba
von: Lu, Chunhao, et al.
Veröffentlicht: (2025)
von: Lu, Chunhao, et al.
Veröffentlicht: (2025)
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding
von: Tu, Xuezhen, et al.
Veröffentlicht: (2026)
von: Tu, Xuezhen, et al.
Veröffentlicht: (2026)
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
von: Chen, Ruizhe, et al.
Veröffentlicht: (2025)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2025)
BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
von: Zhu, Zihao, et al.
Veröffentlicht: (2026)
von: Zhu, Zihao, et al.
Veröffentlicht: (2026)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
von: Wei, Lai, et al.
Veröffentlicht: (2023)
von: Wei, Lai, et al.
Veröffentlicht: (2023)
Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment
von: Zheng, Yuze, et al.
Veröffentlicht: (2024)
von: Zheng, Yuze, et al.
Veröffentlicht: (2024)
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
von: Wang, Jiazhen, et al.
Veröffentlicht: (2023)
von: Wang, Jiazhen, et al.
Veröffentlicht: (2023)
FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
von: Cao, Zhuo, et al.
Veröffentlicht: (2024)
von: Cao, Zhuo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2026) -
Rethinking Metrics and Benchmarks of Video Anomaly Detection
von: Liu, Zihao, et al.
Veröffentlicht: (2025) -
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
von: Li, Wenrui, et al.
Veröffentlicht: (2024) -
ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions
von: Liu, Zihao, et al.
Veröffentlicht: (2026) -
Language-guided Open-world Video Anomaly Detection under Weak Supervision
von: Liu, Zihao, et al.
Veröffentlicht: (2025)