Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jiaxin, Wei, Xiao-Yong, Li, Qing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's Education
von: Jia, Yanhao, et al.
Veröffentlicht: (2025)
von: Jia, Yanhao, et al.
Veröffentlicht: (2025)
ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning
von: Luo, Pengfei, et al.
Veröffentlicht: (2025)
von: Luo, Pengfei, et al.
Veröffentlicht: (2025)
On the Brittleness of CLIP Text Encoders
von: Tran, Allie, et al.
Veröffentlicht: (2025)
von: Tran, Allie, et al.
Veröffentlicht: (2025)
Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
von: Wu, Qiyu, et al.
Veröffentlicht: (2025)
von: Wu, Qiyu, et al.
Veröffentlicht: (2025)
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
von: Jing, Xiaolun, et al.
Veröffentlicht: (2024)
von: Jing, Xiaolun, et al.
Veröffentlicht: (2024)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Semantic Item Graph Enhancement for Multimodal Recommendation
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2025)
REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts
von: Lin, Xinkui, et al.
Veröffentlicht: (2025)
von: Lin, Xinkui, et al.
Veröffentlicht: (2025)
Personalized Image Generation with Large Multimodal Models
von: Xu, Yiyan, et al.
Veröffentlicht: (2024)
von: Xu, Yiyan, et al.
Veröffentlicht: (2024)
Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
von: Nareti, Utsav Kumar, et al.
Veröffentlicht: (2024)
von: Nareti, Utsav Kumar, et al.
Veröffentlicht: (2024)
CrossPT-EEG: A Benchmark for Cross-Participant and Cross-Time Generalization of EEG-based Visual Decoding
von: Zhu, Shuqi, et al.
Veröffentlicht: (2024)
von: Zhu, Shuqi, et al.
Veröffentlicht: (2024)
Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond
von: Wei, Tianxin, et al.
Veröffentlicht: (2024)
von: Wei, Tianxin, et al.
Veröffentlicht: (2024)
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval
von: Zou, Qiang, et al.
Veröffentlicht: (2025)
von: Zou, Qiang, et al.
Veröffentlicht: (2025)
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
von: Wang, Yuting, et al.
Veröffentlicht: (2023)
von: Wang, Yuting, et al.
Veröffentlicht: (2023)
VCR: Video representation for Contextual Retrieval
von: Nir, Oron, et al.
Veröffentlicht: (2024)
von: Nir, Oron, et al.
Veröffentlicht: (2024)
VDCook:DIY video data cook your MLLMs
von: Wu, Chengwei
Veröffentlicht: (2026)
von: Wu, Chengwei
Veröffentlicht: (2026)
A Survey of Multimodal Composite Editing and Retrieval
von: Li, Suyan, et al.
Veröffentlicht: (2024)
von: Li, Suyan, et al.
Veröffentlicht: (2024)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
Enhancing Image-Text Matching with Adaptive Feature Aggregation
von: Wang, Zuhui, et al.
Veröffentlicht: (2024)
von: Wang, Zuhui, et al.
Veröffentlicht: (2024)
Semantic Codebook Learning for Dynamic Recommendation Models
von: Lv, Zheqi, et al.
Veröffentlicht: (2024)
von: Lv, Zheqi, et al.
Veröffentlicht: (2024)
Agentic Mixed-Source Multi-Modal Misinformation Detection with Adaptive Test-Time Scaling
von: Jiang, Wei, et al.
Veröffentlicht: (2026)
von: Jiang, Wei, et al.
Veröffentlicht: (2026)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
A Comprehensive Survey on Composed Image Retrieval
von: Song, Xuemeng, et al.
Veröffentlicht: (2025)
von: Song, Xuemeng, et al.
Veröffentlicht: (2025)
PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
von: Xu, Tianyi, et al.
Veröffentlicht: (2026)
von: Xu, Tianyi, et al.
Veröffentlicht: (2026)
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
von: Williams-Lekuona, Mikel, et al.
Veröffentlicht: (2025)
von: Williams-Lekuona, Mikel, et al.
Veröffentlicht: (2025)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
von: He, Xiao, et al.
Veröffentlicht: (2025)
von: He, Xiao, et al.
Veröffentlicht: (2025)
Anchor-aware Deep Metric Learning for Audio-visual Retrieval
von: Zeng, Donghuo, et al.
Veröffentlicht: (2024)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2024)
LightThinker++: From Reasoning Compression to Memory Management
von: Zhu, Yuqi, et al.
Veröffentlicht: (2026)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2026)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
EventCast: Hybrid Demand Forecasting in E-Commerce with LLM-Based Event Knowledge
von: Hu, Congcong, et al.
Veröffentlicht: (2026)
von: Hu, Congcong, et al.
Veröffentlicht: (2026)
ChatDiet: Empowering Personalized Nutrition-Oriented Food Recommender Chatbots through an LLM-Augmented Framework
von: Yang, Zhongqi, et al.
Veröffentlicht: (2024)
von: Yang, Zhongqi, et al.
Veröffentlicht: (2024)
A Multimodal Single-Branch Embedding Network for Recommendation in Cold-Start and Missing Modality Scenarios
von: Ganhör, Christian, et al.
Veröffentlicht: (2024)
von: Ganhör, Christian, et al.
Veröffentlicht: (2024)
Multimodal Music Recommendation System using LLMs
von: Kandagatla, Srikar Prabhas, et al.
Veröffentlicht: (2026)
von: Kandagatla, Srikar Prabhas, et al.
Veröffentlicht: (2026)
On the Origin of Synthetic Information by Means of Steganographic Inheritance
von: Chang, Ching-Chun, et al.
Veröffentlicht: (2026)
von: Chang, Ching-Chun, et al.
Veröffentlicht: (2026)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
von: Gigant, Théo, et al.
Veröffentlicht: (2024)
von: Gigant, Théo, et al.
Veröffentlicht: (2024)
LazyVLM: Neuro-Symbolic Approach to Video Analytics
von: Jian, Xiangru, et al.
Veröffentlicht: (2025)
von: Jian, Xiangru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's Education
von: Jia, Yanhao, et al.
Veröffentlicht: (2025) -
ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning
von: Luo, Pengfei, et al.
Veröffentlicht: (2025) -
On the Brittleness of CLIP Text Encoders
von: Tran, Allie, et al.
Veröffentlicht: (2025) -
Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
von: Nie, Zhijie, et al.
Veröffentlicht: (2024) -
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
von: Wu, Qiyu, et al.
Veröffentlicht: (2025)