Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Huatuan, Ma, Yunshan, Wu, Changguang, Zhang, Yanxin, Wang, Pengfei, Du, Xiaoyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual-Diffusional Generative Fashion Recommendation
by: Yu, Mingzhe, et al.
Published: (2026)
by: Yu, Mingzhe, et al.
Published: (2026)
FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
by: Yu, Mingzhe, et al.
Published: (2025)
by: Yu, Mingzhe, et al.
Published: (2025)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
Balancing Semantic Relevance and Engagement in Related Video Recommendations
by: Jaspal, Amit, et al.
Published: (2025)
by: Jaspal, Amit, et al.
Published: (2025)
Learning Item Representations Directly from Multimodal Features for Effective Recommendation
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation
by: Ong, Rongqing Kenneth, et al.
Published: (2024)
by: Ong, Rongqing Kenneth, et al.
Published: (2024)
Multimodal Pretraining, Adaptation, and Generation for Recommendation: A Survey
by: Liu, Qijiong, et al.
Published: (2024)
by: Liu, Qijiong, et al.
Published: (2024)
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
by: Ma, Hongjian, et al.
Published: (2026)
by: Ma, Hongjian, et al.
Published: (2026)
BeFA: A General Behavior-driven Feature Adapter for Multimedia Recommendation
by: Fan, Qile, et al.
Published: (2024)
by: Fan, Qile, et al.
Published: (2024)
OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions
by: Zhou, Hongyu, et al.
Published: (2025)
by: Zhou, Hongyu, et al.
Published: (2025)
Multimodal Graph Neural Network for Recommendation with Dynamic De-redundancy and Modality-Guided Feature De-noisy
by: Mo, Feng, et al.
Published: (2024)
by: Mo, Feng, et al.
Published: (2024)
CM$^3$: Calibrating Multimodal Recommendation
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
by: Zhan, Hao, et al.
Published: (2026)
by: Zhan, Hao, et al.
Published: (2026)
Robust Relevance Feedback for Interactive Known-Item Video Search
by: Ma, Zhixin, et al.
Published: (2025)
by: Ma, Zhixin, et al.
Published: (2025)
DREAM: A Dual Representation Learning Model for Multimodal Recommendation
by: Zhang, Kangning, et al.
Published: (2024)
by: Zhang, Kangning, et al.
Published: (2024)
MMSRARec: Summarization and Retrieval Augumented Sequential Recommendation Based on Multimodal Large Language Model
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Multimodal Pre-training Framework for Sequential Recommendation via Contrastive Learning
by: Zhang, Lingzi, et al.
Published: (2023)
by: Zhang, Lingzi, et al.
Published: (2023)
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
by: Wang, Tianshi, et al.
Published: (2023)
by: Wang, Tianshi, et al.
Published: (2023)
From ID-based to ID-free: Rethinking ID Effectiveness in Multimodal Collaborative Filtering Recommendation
by: Li, Guohao, et al.
Published: (2025)
by: Li, Guohao, et al.
Published: (2025)
A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
Beyond Static Collision Handling: Adaptive Semantic ID Learning for Multimodal Recommendation at Industrial Scale
by: Pan, Yongsen, et al.
Published: (2026)
by: Pan, Yongsen, et al.
Published: (2026)
Enhancing Image-Text Matching with Adaptive Feature Aggregation
by: Wang, Zuhui, et al.
Published: (2024)
by: Wang, Zuhui, et al.
Published: (2024)
CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential Recommendation
by: Xu, Jinfeng, et al.
Published: (2026)
by: Xu, Jinfeng, et al.
Published: (2026)
Knowledge-aware Diffusion-Enhanced Multimedia Recommendation
by: Mo, Xian, et al.
Published: (2025)
by: Mo, Xian, et al.
Published: (2025)
HistLLM: A Unified Framework for LLM-Based Multimodal Recommendation with User History Encoding and Compression
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Breaking the Curse of Knowledge: Towards Effective Multimodal Recommendation using Knowledge Soft Integration
by: Ouyang, Kai, et al.
Published: (2023)
by: Ouyang, Kai, et al.
Published: (2023)
Attribute-driven Disentangled Representation Learning for Multimodal Recommendation
by: Li, Zhenyang, et al.
Published: (2023)
by: Li, Zhenyang, et al.
Published: (2023)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
Disentangled Graph Variational Auto-Encoder for Multimodal Recommendation with Interpretability
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
MDF: A Dynamic Fusion Model for Multi-modal Fake News Detection
by: Lv, Hongzhen, et al.
Published: (2024)
by: Lv, Hongzhen, et al.
Published: (2024)
CIRP: Cross-Item Relational Pre-training for Multimodal Product Bundling
by: Ma, Yunshan, et al.
Published: (2024)
by: Ma, Yunshan, et al.
Published: (2024)
VCR: Video representation for Contextual Retrieval
by: Nir, Oron, et al.
Published: (2024)
by: Nir, Oron, et al.
Published: (2024)
Results of the 2025 Video Browser Showdown
by: Rossetto, Luca, et al.
Published: (2025)
by: Rossetto, Luca, et al.
Published: (2025)
Results of the 2024 Video Browser Showdown
by: Rossetto, Luca, et al.
Published: (2024)
by: Rossetto, Luca, et al.
Published: (2024)
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
by: Tourani, Ali, et al.
Published: (2025)
by: Tourani, Ali, et al.
Published: (2025)
Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network
by: Chen, Zheyu, et al.
Published: (2024)
by: Chen, Zheyu, et al.
Published: (2024)
Small Stickers, Big Meanings: A Multilingual Sticker Semantic Understanding Dataset with a Gamified Approach
by: Chee, Heng Er Metilda, et al.
Published: (2025)
by: Chee, Heng Er Metilda, et al.
Published: (2025)
Leveraging User-Generated Metadata of Online Videos for Cover Song Identification
by: Hachmeier, Simon, et al.
Published: (2024)
by: Hachmeier, Simon, et al.
Published: (2024)
U-Sticker: A Large-Scale Multi-Domain User Sticker Dataset for Retrieval and Personalization
by: Chee, Heng Er Metilda, et al.
Published: (2025)
by: Chee, Heng Er Metilda, et al.
Published: (2025)
Similar Items
-
Dual-Diffusional Generative Fashion Recommendation
by: Yu, Mingzhe, et al.
Published: (2026) -
FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
by: Yu, Mingzhe, et al.
Published: (2025) -
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025) -
Balancing Semantic Relevance and Engagement in Related Video Recommendations
by: Jaspal, Amit, et al.
Published: (2025) -
Learning Item Representations Directly from Multimodal Features for Effective Recommendation
by: Zhou, Xin, et al.
Published: (2025)