Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Yunzhuo, Liu, Xinyue, Li, Yanyang, Wu, Nanding, Xu, Yifang, Zong, Linlin, Zhang, Xianchao, Liang, Wenxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
von: Xu, Yifang, et al.
Veröffentlicht: (2024)
von: Xu, Yifang, et al.
Veröffentlicht: (2024)
Hawkes based Representation Learning for Reasoning over Scale-free Community-structured Temporal Knowledge Graphs
von: Du, Yuwei, et al.
Veröffentlicht: (2024)
von: Du, Yuwei, et al.
Veröffentlicht: (2024)
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
von: Zhu, Zhiyi, et al.
Veröffentlicht: (2025)
von: Zhu, Zhiyi, et al.
Veröffentlicht: (2025)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
von: An, Joungbin, et al.
Veröffentlicht: (2025)
von: An, Joungbin, et al.
Veröffentlicht: (2025)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2024)
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2024)
HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
von: Zhao, Henghao, et al.
Veröffentlicht: (2023)
von: Zhao, Henghao, et al.
Veröffentlicht: (2023)
Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
Improve Meta-learning for Few-Shot Text Classification with All You Can Acquire from the Tasks
von: Liu, Xinyue, et al.
Veröffentlicht: (2024)
von: Liu, Xinyue, et al.
Veröffentlicht: (2024)
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
von: Moon, WonJun, et al.
Veröffentlicht: (2023)
von: Moon, WonJun, et al.
Veröffentlicht: (2023)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Mitigating Memorization in Text-to-Image Diffusion via Region-Aware Prompt Augmentation and Multimodal Copy Detection
von: Chen, Yunzhuo, et al.
Veröffentlicht: (2026)
von: Chen, Yunzhuo, et al.
Veröffentlicht: (2026)
QD-VMR: Query Debiasing with Contextual Understanding Enhancement for Video Moment Retrieval
von: Gao, Chenghua, et al.
Veröffentlicht: (2024)
von: Gao, Chenghua, et al.
Veröffentlicht: (2024)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba
von: Wu, Zizhao, et al.
Veröffentlicht: (2025)
von: Wu, Zizhao, et al.
Veröffentlicht: (2025)
Number it: Temporal Grounding Videos like Flipping Manga
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
Triple Path Enhanced Neural Architecture Search for Multimodal Fake News Detection
von: Xu, Bo, et al.
Veröffentlicht: (2025)
von: Xu, Bo, et al.
Veröffentlicht: (2025)
HGP-Mamba: Integrating Histology and Generated Protein Features for Mamba-based Multimodal Survival Risk Prediction
von: Dai, Jing, et al.
Veröffentlicht: (2026)
von: Dai, Jing, et al.
Veröffentlicht: (2026)
Diversifying Query: Region-Guided Transformer for Temporal Sentence Grounding
von: Sun, Xiaolong, et al.
Veröffentlicht: (2024)
von: Sun, Xiaolong, et al.
Veröffentlicht: (2024)
DIFFUMA: High-Fidelity Spatio-Temporal Video Prediction via Dual-Path Mamba and Diffusion Enhancement
von: Xie, Xinyu, et al.
Veröffentlicht: (2025)
von: Xie, Xinyu, et al.
Veröffentlicht: (2025)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
von: Vu, Huu-An, et al.
Veröffentlicht: (2025)
von: Vu, Huu-An, et al.
Veröffentlicht: (2025)
Matten: Video Generation with Mamba-Attention
von: Gao, Yu, et al.
Veröffentlicht: (2024)
von: Gao, Yu, et al.
Veröffentlicht: (2024)
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Moment Quantization for Video Temporal Grounding
von: Sun, Xiaolong, et al.
Veröffentlicht: (2025)
von: Sun, Xiaolong, et al.
Veröffentlicht: (2025)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
von: Fan, Rong, et al.
Veröffentlicht: (2026)
von: Fan, Rong, et al.
Veröffentlicht: (2026)
Beyond Alignment: Blind Video Face Restoration via Parsing-Guided Temporal-Coherent Transformer
von: Xu, Kepeng, et al.
Veröffentlicht: (2024)
von: Xu, Kepeng, et al.
Veröffentlicht: (2024)
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory
von: Deng, Tianchen, et al.
Veröffentlicht: (2026)
von: Deng, Tianchen, et al.
Veröffentlicht: (2026)
Unmixing-Guided Spatial-Spectral Mamba with Clustering Tokens for Hyperspectral Image Classification
von: Zhu, Yimin, et al.
Veröffentlicht: (2026)
von: Zhu, Yimin, et al.
Veröffentlicht: (2026)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
Category Prompt Mamba Network for Nuclei Segmentation and Classification
von: Zhang, Ye, et al.
Veröffentlicht: (2025)
von: Zhang, Ye, et al.
Veröffentlicht: (2025)
Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2025)
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2025)
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
von: Yang, Zaiquan, et al.
Veröffentlicht: (2025)
von: Yang, Zaiquan, et al.
Veröffentlicht: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
von: Alansari, Mohamad, et al.
Veröffentlicht: (2026)
von: Alansari, Mohamad, et al.
Veröffentlicht: (2026)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
von: Xu, Yifang, et al.
Veröffentlicht: (2025) -
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
von: Xu, Yifang, et al.
Veröffentlicht: (2024) -
Hawkes based Representation Learning for Reasoning over Scale-free Community-structured Temporal Knowledge Graphs
von: Du, Yuwei, et al.
Veröffentlicht: (2024) -
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
von: Xu, Yifang, et al.
Veröffentlicht: (2025) -
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
von: Zhu, Zhiyi, et al.
Veröffentlicht: (2025)