MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tang, Haoran, Cao, Meng, Huang, Jinfa, Liu, Ruyang, Jin, Peng, Li, Ge, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
par: Cao, Meng, et autres
Publié: (2024)
par: Cao, Meng, et autres
Publié: (2024)
Video Spatial Reasoning with Object-Centric 3D Rollout
par: Tang, Haoran, et autres
Publié: (2025)
par: Tang, Haoran, et autres
Publié: (2025)
ST-LLM: Large Language Models Are Effective Temporal Learners
par: Liu, Ruyang, et autres
Publié: (2024)
par: Liu, Ruyang, et autres
Publié: (2024)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
par: Cao, Meng, et autres
Publié: (2024)
par: Cao, Meng, et autres
Publié: (2024)
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
par: Cao, Meng, et autres
Publié: (2026)
par: Cao, Meng, et autres
Publié: (2026)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
par: Liu, Ruyang, et autres
Publié: (2025)
par: Liu, Ruyang, et autres
Publié: (2025)
TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba
par: Chen, Xiuwei, et autres
Publié: (2025)
par: Chen, Xiuwei, et autres
Publié: (2025)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
par: Sun, Shangkun, et autres
Publié: (2024)
par: Sun, Shangkun, et autres
Publié: (2024)
MLP Can Be A Good Transformer Learner
par: Lin, Sihao, et autres
Publié: (2024)
par: Lin, Sihao, et autres
Publié: (2024)
RI-Mamba: Rotation-Invariant Mamba for Robust Text-to-Shape Retrieval
par: Nguyen, Khanh, et autres
Publié: (2026)
par: Nguyen, Khanh, et autres
Publié: (2026)
GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections
par: Zhang, Shiyue, et autres
Publié: (2024)
par: Zhang, Shiyue, et autres
Publié: (2024)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
par: Cao, Meng, et autres
Publié: (2025)
par: Cao, Meng, et autres
Publié: (2025)
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
par: Wang, Xiaodong, et autres
Publié: (2025)
par: Wang, Xiaodong, et autres
Publié: (2025)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
par: Liu, Ruyang, et autres
Publié: (2023)
par: Liu, Ruyang, et autres
Publié: (2023)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
par: Yuan, Shenghai, et autres
Publié: (2024)
par: Yuan, Shenghai, et autres
Publié: (2024)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
par: Liu, Weijia, et autres
Publié: (2025)
par: Liu, Weijia, et autres
Publié: (2025)
Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark
par: Deng, Yifei, et autres
Publié: (2026)
par: Deng, Yifei, et autres
Publié: (2026)
VMRNN: Integrating Vision Mamba and LSTM for Efficient and Accurate Spatiotemporal Forecasting
par: Tang, Yujin, et autres
Publié: (2024)
par: Tang, Yujin, et autres
Publié: (2024)
HS-Mamba: Full-Field Interaction Multi-Groups Mamba for Hyperspectral Image Classification
par: Peng, Hongxing, et autres
Publié: (2025)
par: Peng, Hongxing, et autres
Publié: (2025)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
par: Ying, Xinru, et autres
Publié: (2025)
par: Ying, Xinru, et autres
Publié: (2025)
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
par: Cao, Meng, et autres
Publié: (2025)
par: Cao, Meng, et autres
Publié: (2025)
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
par: Huang, Minbin, et autres
Publié: (2024)
par: Huang, Minbin, et autres
Publié: (2024)
LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
par: Wan, Zhongwei, et autres
Publié: (2024)
par: Wan, Zhongwei, et autres
Publié: (2024)
Trusted Mamba Contrastive Network for Multi-View Clustering
par: Zhu, Jian, et autres
Publié: (2024)
par: Zhu, Jian, et autres
Publié: (2024)
MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic Expansion
par: Peng, Fei, et autres
Publié: (2025)
par: Peng, Fei, et autres
Publié: (2025)
Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI Generation
par: Huang, Guohong, et autres
Publié: (2025)
par: Huang, Guohong, et autres
Publié: (2025)
DB-MSMUNet:Dual Branch Multi-scale Mamba UNet for Pancreatic CT Scans Segmentation
par: Guan, Qiu, et autres
Publié: (2026)
par: Guan, Qiu, et autres
Publié: (2026)
Qihoo-T2X: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Any-Task
par: Wang, Jing, et autres
Publié: (2024)
par: Wang, Jing, et autres
Publié: (2024)
MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification
par: Yang, Zijiang, et autres
Publié: (2025)
par: Yang, Zijiang, et autres
Publié: (2025)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
par: Yin, Liang, et autres
Publié: (2025)
par: Yin, Liang, et autres
Publié: (2025)
BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving
par: Tang, Tao, et autres
Publié: (2024)
par: Tang, Tao, et autres
Publié: (2024)
TextMamba: Scene Text Detector with Mamba
par: Zhao, Qiyan, et autres
Publié: (2025)
par: Zhao, Qiyan, et autres
Publié: (2025)
Motion Mamba: Efficient and Long Sequence Motion Generation
par: Zhang, Zeyu, et autres
Publié: (2024)
par: Zhang, Zeyu, et autres
Publié: (2024)
Trajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSM
par: Huang, Yizhou, et autres
Publié: (2025)
par: Huang, Yizhou, et autres
Publié: (2025)
UAVPairs: A Challenging Benchmark for Match Pair Retrieval of Large-scale UAV Images
par: Liu, Junhuan, et autres
Publié: (2025)
par: Liu, Junhuan, et autres
Publié: (2025)
Versatile and Efficient Medical Image Super-Resolution Via Frequency-Gated Mamba
par: Huang, Wenfeng, et autres
Publié: (2025)
par: Huang, Wenfeng, et autres
Publié: (2025)
MVQA: Mamba with Unified Sampling for Efficient Video Quality Assessment
par: Mi, Yachun, et autres
Publié: (2025)
par: Mi, Yachun, et autres
Publié: (2025)
Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
par: Zhao, Zecheng, et autres
Publié: (2026)
par: Zhao, Zecheng, et autres
Publié: (2026)
ML-Mamba: Efficient Multi-Modal Large Language Model Utilizing Mamba-2
par: Huang, Wenjun, et autres
Publié: (2024)
par: Huang, Wenjun, et autres
Publié: (2024)
GeoMamba: A Geometry-driven MambaVision Framework and Dataset for Fine-grained Optical-SAR Object Retrieval
par: Fang, Tiantong, et autres
Publié: (2026)
par: Fang, Tiantong, et autres
Publié: (2026)
Documents similaires
-
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
par: Cao, Meng, et autres
Publié: (2024) -
Video Spatial Reasoning with Object-Centric 3D Rollout
par: Tang, Haoran, et autres
Publié: (2025) -
ST-LLM: Large Language Models Are Effective Temporal Learners
par: Liu, Ruyang, et autres
Publié: (2024) -
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
par: Cao, Meng, et autres
Publié: (2024) -
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
par: Cao, Meng, et autres
Publié: (2026)