MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
Fuente:
arXiv
Salvato in:
| Autori principali: | Ju, Yeong-Joon, Kim, Ho-Joong, Lee, Seong-Whan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
di: Kong, Fanheng, et al.
Pubblicazione: (2025)
di: Kong, Fanheng, et al.
Pubblicazione: (2025)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
di: Liu, Han, et al.
Pubblicazione: (2025)
di: Liu, Han, et al.
Pubblicazione: (2025)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
di: Li, Po-han, et al.
Pubblicazione: (2024)
di: Li, Po-han, et al.
Pubblicazione: (2024)
Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation
di: Ong, Rongqing Kenneth, et al.
Pubblicazione: (2024)
di: Ong, Rongqing Kenneth, et al.
Pubblicazione: (2024)
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
di: Fu, Junchen, et al.
Pubblicazione: (2026)
di: Fu, Junchen, et al.
Pubblicazione: (2026)
From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
di: Wang, Zheng, et al.
Pubblicazione: (2025)
di: Wang, Zheng, et al.
Pubblicazione: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
A Survey of Multimodal Composite Editing and Retrieval
di: Li, Suyan, et al.
Pubblicazione: (2024)
di: Li, Suyan, et al.
Pubblicazione: (2024)
Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
di: Wu, Yiming, et al.
Pubblicazione: (2024)
di: Wu, Yiming, et al.
Pubblicazione: (2024)
Interactive Multi-Turn Retrieval for Health Videos
di: Wu, Chengzheng, et al.
Pubblicazione: (2026)
di: Wu, Chengzheng, et al.
Pubblicazione: (2026)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
di: Messina, Nicola, et al.
Pubblicazione: (2024)
di: Messina, Nicola, et al.
Pubblicazione: (2024)
A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
di: Han, Haochen, et al.
Pubblicazione: (2024)
di: Han, Haochen, et al.
Pubblicazione: (2024)
From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
di: Ju, Yeong-Joon, et al.
Pubblicazione: (2025)
di: Ju, Yeong-Joon, et al.
Pubblicazione: (2025)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
di: Ning, Hailong, et al.
Pubblicazione: (2025)
di: Ning, Hailong, et al.
Pubblicazione: (2025)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
di: Li, Jun, et al.
Pubblicazione: (2025)
di: Li, Jun, et al.
Pubblicazione: (2025)
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
di: Ma, Hongjian, et al.
Pubblicazione: (2026)
di: Ma, Hongjian, et al.
Pubblicazione: (2026)
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
di: Li, Jun, et al.
Pubblicazione: (2026)
di: Li, Jun, et al.
Pubblicazione: (2026)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
di: Zhang, Jiahao, et al.
Pubblicazione: (2025)
di: Zhang, Jiahao, et al.
Pubblicazione: (2025)
The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval
di: Fu, Junchen, et al.
Pubblicazione: (2026)
di: Fu, Junchen, et al.
Pubblicazione: (2026)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
di: Gatti, Prajwal, et al.
Pubblicazione: (2025)
di: Gatti, Prajwal, et al.
Pubblicazione: (2025)
Multimodal Learned Sparse Retrieval for Image Suggestion
di: Nguyen, Thong, et al.
Pubblicazione: (2024)
di: Nguyen, Thong, et al.
Pubblicazione: (2024)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
di: Fang, Xiang, et al.
Pubblicazione: (2022)
di: Fang, Xiang, et al.
Pubblicazione: (2022)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
di: Xiao, Jian, et al.
Pubblicazione: (2024)
di: Xiao, Jian, et al.
Pubblicazione: (2024)
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval
di: Zou, Qiang, et al.
Pubblicazione: (2025)
di: Zou, Qiang, et al.
Pubblicazione: (2025)
Attribute-driven Disentangled Representation Learning for Multimodal Recommendation
di: Li, Zhenyang, et al.
Pubblicazione: (2023)
di: Li, Zhenyang, et al.
Pubblicazione: (2023)
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
di: Rossetto, Luca, et al.
Pubblicazione: (2025)
di: Rossetto, Luca, et al.
Pubblicazione: (2025)
Very Efficient Listwise Multimodal Reranking for Long Documents
di: Sun, Yiqun, et al.
Pubblicazione: (2026)
di: Sun, Yiqun, et al.
Pubblicazione: (2026)
Multimodal Neural Databases
di: Trappolini, Giovanni, et al.
Pubblicazione: (2023)
di: Trappolini, Giovanni, et al.
Pubblicazione: (2023)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
di: Xiao, Jian, et al.
Pubblicazione: (2025)
di: Xiao, Jian, et al.
Pubblicazione: (2025)
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval
di: Messina, Nicola, et al.
Pubblicazione: (2024)
di: Messina, Nicola, et al.
Pubblicazione: (2024)
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
di: Deng, Jiaqi, et al.
Pubblicazione: (2025)
di: Deng, Jiaqi, et al.
Pubblicazione: (2025)
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
di: Zhan, Hao, et al.
Pubblicazione: (2026)
di: Zhan, Hao, et al.
Pubblicazione: (2026)
Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation
di: Kim, Wongyu, et al.
Pubblicazione: (2025)
di: Kim, Wongyu, et al.
Pubblicazione: (2025)
DREAM: A Dual Representation Learning Model for Multimodal Recommendation
di: Zhang, Kangning, et al.
Pubblicazione: (2024)
di: Zhang, Kangning, et al.
Pubblicazione: (2024)
OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
di: Li, Minghan, et al.
Pubblicazione: (2026)
di: Li, Minghan, et al.
Pubblicazione: (2026)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
di: Shih, Yu-Fei, et al.
Pubblicazione: (2025)
di: Shih, Yu-Fei, et al.
Pubblicazione: (2025)
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
di: Barrios, Wayner, et al.
Pubblicazione: (2026)
di: Barrios, Wayner, et al.
Pubblicazione: (2026)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
di: Rothermel, Mark, et al.
Pubblicazione: (2026)
di: Rothermel, Mark, et al.
Pubblicazione: (2026)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
di: Luo, Tianci, et al.
Pubblicazione: (2026)
di: Luo, Tianci, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
di: Kong, Fanheng, et al.
Pubblicazione: (2025) -
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
di: Liu, Han, et al.
Pubblicazione: (2025) -
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
di: Li, Po-han, et al.
Pubblicazione: (2024) -
Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation
di: Ong, Rongqing Kenneth, et al.
Pubblicazione: (2024) -
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
di: Fu, Junchen, et al.
Pubblicazione: (2026)