CART: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Minghui, Ji, Shengpeng, Zuo, Jialong, Huang, Hai, Xia, Yan, Zhu, Jieming, Cheng, Xize, Yang, Xiaoda, Liu, Wenrui, Wang, Gang, Dong, Zhenhua, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
by: Hu, Ruofan, et al.
Published: (2025)
by: Hu, Ruofan, et al.
Published: (2025)
EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration
by: Hong, Minjie, et al.
Published: (2025)
by: Hong, Minjie, et al.
Published: (2025)
UNGER: Generative Recommendation with A Unified Code via Semantic and Collaborative Integration
by: Xiao, Longtao, et al.
Published: (2025)
by: Xiao, Longtao, et al.
Published: (2025)
MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
by: Li, Tianyuan, et al.
Published: (2025)
by: Li, Tianyuan, et al.
Published: (2025)
CoST: Contrastive Quantization based Semantic Tokenization for Generative Recommendation
by: Zhu, Jieming, et al.
Published: (2024)
by: Zhu, Jieming, et al.
Published: (2024)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Multimodal Pretraining and Generation for Recommendation: A Tutorial
by: Zhu, Jieming, et al.
Published: (2024)
by: Zhu, Jieming, et al.
Published: (2024)
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Recall-Augmented Ranking: Enhancing Click-Through Rate Prediction Accuracy with Cross-Stage Data
by: Huang, Junjie, et al.
Published: (2024)
by: Huang, Junjie, et al.
Published: (2024)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Counteracting Duration Bias in Video Recommendation via Counterfactual Watch Time
by: Zhao, Haiyuan, et al.
Published: (2024)
by: Zhao, Haiyuan, et al.
Published: (2024)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
EAGER: Two-Stream Generative Recommender with Behavior-Semantic Collaboration
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
A Unified Optimal Transport Framework for Cross-Modal Retrieval with Noisy Labels
by: Han, Haochen, et al.
Published: (2024)
by: Han, Haochen, et al.
Published: (2024)
FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
by: Zhao, Xinping, et al.
Published: (2024)
by: Zhao, Xinping, et al.
Published: (2024)
Towards Cross-Modal Text-Molecule Retrieval with Better Modality Alignment
by: Song, Jia, et al.
Published: (2024)
by: Song, Jia, et al.
Published: (2024)
RREH: Reconstruction Relations Embedded Hashing for Semi-Paired Cross-Modal Retrieval
by: Wang, Jianzong, et al.
Published: (2024)
by: Wang, Jianzong, et al.
Published: (2024)
CORONA: A Coarse-to-Fine Framework for Graph-based Recommendation with Large Language Models
by: Chen, Junze, et al.
Published: (2025)
by: Chen, Junze, et al.
Published: (2025)
XR: Cross-Modal Agents for Composed Image Retrieval
by: Yang, Zhongyu, et al.
Published: (2026)
by: Yang, Zhongyu, et al.
Published: (2026)
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
by: Zhou, Sashuai, et al.
Published: (2025)
by: Zhou, Sashuai, et al.
Published: (2025)
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
by: Hu, Ruofan, et al.
Published: (2026)
by: Hu, Ruofan, et al.
Published: (2026)
Learning Multi-Aspect Item Palette: A Semantic Tokenization Framework for Generative Recommendation
by: Liu, Qijiong, et al.
Published: (2024)
by: Liu, Qijiong, et al.
Published: (2024)
Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search
by: Zhou, Zhigong, et al.
Published: (2025)
by: Zhou, Zhigong, et al.
Published: (2025)
TayFCS: Towards Light Feature Combination Selection for Deep Recommender Systems
by: Wang, Xianquan, et al.
Published: (2025)
by: Wang, Xianquan, et al.
Published: (2025)
FairFS: Addressing Deep Feature Selection Biases for Recommender System
by: Wang, Xianquan, et al.
Published: (2026)
by: Wang, Xianquan, et al.
Published: (2026)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
by: Huang, Jinghao, et al.
Published: (2025)
by: Huang, Jinghao, et al.
Published: (2025)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
by: Liu, Qijiong, et al.
Published: (2025)
by: Liu, Qijiong, et al.
Published: (2025)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
by: Xiao, Jian, et al.
Published: (2025)
by: Xiao, Jian, et al.
Published: (2025)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Retrieval Augmented Cross-Modal Tag Recommendation in Software Q&A Sites
by: Lu, Sijin, et al.
Published: (2024)
by: Lu, Sijin, et al.
Published: (2024)
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
by: Wang, Tianshi, et al.
Published: (2023)
by: Wang, Tianshi, et al.
Published: (2023)
Improving Semantic Proximity in Information Retrieval through Cross-Lingual Alignment
by: Hong, Seongtae, et al.
Published: (2026)
by: Hong, Seongtae, et al.
Published: (2026)
Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation
by: Dai, Ji, et al.
Published: (2026)
by: Dai, Ji, et al.
Published: (2026)
HASH-RAG: Bridging Deep Hashing with Retriever for Efficient, Fine Retrieval and Augmented Generation
by: Guo, Jinyu, et al.
Published: (2025)
by: Guo, Jinyu, et al.
Published: (2025)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Evaluating Large Language Models for Cross-Lingual Retrieval
by: Zuo, Longfei, et al.
Published: (2025)
by: Zuo, Longfei, et al.
Published: (2025)
FAIR: Focused Attention Is All You Need for Generative Recommendation
by: Xiao, Longtao, et al.
Published: (2025)
by: Xiao, Longtao, et al.
Published: (2025)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
Similar Items
-
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
by: Hu, Ruofan, et al.
Published: (2025) -
EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration
by: Hong, Minjie, et al.
Published: (2025) -
UNGER: Generative Recommendation with A Unified Code via Semantic and Collaborative Integration
by: Xiao, Longtao, et al.
Published: (2025) -
MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
by: Li, Tianyuan, et al.
Published: (2025) -
CoST: Contrastive Quantization based Semantic Tokenization for Generative Recommendation
by: Zhu, Jieming, et al.
Published: (2024)