Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
Fuente:
arXiv
Guardado en:
| Autores principales: | Nareti, Utsav Kumar, Adak, Chandranath, Chattopadhyay, Soumi, Wang, Pichao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
por: Kumar, Suraj, et al.
Publicado: (2025)
por: Kumar, Suraj, et al.
Publicado: (2025)
Demystifying Visual Features of Movie Posters for Multi-Label Genre Identification
por: Nareti, Utsav Kumar, et al.
Publicado: (2023)
por: Nareti, Utsav Kumar, et al.
Publicado: (2023)
Adaptive Data-Resilient Multi-Modal Hierarchical Multi-Label Book Genre Identification
por: Nareti, Utsav Kumar, et al.
Publicado: (2025)
por: Nareti, Utsav Kumar, et al.
Publicado: (2025)
Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
por: Nie, Zhijie, et al.
Publicado: (2024)
por: Nie, Zhijie, et al.
Publicado: (2024)
CrossPT-EEG: A Benchmark for Cross-Participant and Cross-Time Generalization of EEG-based Visual Decoding
por: Zhu, Shuqi, et al.
Publicado: (2024)
por: Zhu, Shuqi, et al.
Publicado: (2024)
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
por: Wu, Qiyu, et al.
Publicado: (2025)
por: Wu, Qiyu, et al.
Publicado: (2025)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
por: Zhang, Jiahao, et al.
Publicado: (2025)
por: Zhang, Jiahao, et al.
Publicado: (2025)
Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond
por: Wei, Tianxin, et al.
Publicado: (2024)
por: Wei, Tianxin, et al.
Publicado: (2024)
Personalized Image Generation with Large Multimodal Models
por: Xu, Yiyan, et al.
Publicado: (2024)
por: Xu, Yiyan, et al.
Publicado: (2024)
REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts
por: Lin, Xinkui, et al.
Publicado: (2025)
por: Lin, Xinkui, et al.
Publicado: (2025)
Movie Recommendation with Poster Attention via Multi-modal Transformer Feature Fusion
por: Xia, Linhan, et al.
Publicado: (2024)
por: Xia, Linhan, et al.
Publicado: (2024)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
por: Ju, Yeong-Joon, et al.
Publicado: (2024)
por: Ju, Yeong-Joon, et al.
Publicado: (2024)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
por: Fang, Xiang, et al.
Publicado: (2022)
por: Fang, Xiang, et al.
Publicado: (2022)
A Multimodal Single-Branch Embedding Network for Recommendation in Cold-Start and Missing Modality Scenarios
por: Ganhör, Christian, et al.
Publicado: (2024)
por: Ganhör, Christian, et al.
Publicado: (2024)
Semantic Item Graph Enhancement for Multimodal Recommendation
por: Zhang, Xiaoxiong, et al.
Publicado: (2025)
por: Zhang, Xiaoxiong, et al.
Publicado: (2025)
On the Brittleness of CLIP Text Encoders
por: Tran, Allie, et al.
Publicado: (2025)
por: Tran, Allie, et al.
Publicado: (2025)
ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning
por: Luo, Pengfei, et al.
Publicado: (2025)
por: Luo, Pengfei, et al.
Publicado: (2025)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
por: Wu, Jiaxin, et al.
Publicado: (2025)
por: Wu, Jiaxin, et al.
Publicado: (2025)
Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's Education
por: Jia, Yanhao, et al.
Publicado: (2025)
por: Jia, Yanhao, et al.
Publicado: (2025)
ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
por: Nareti, Utsav Kumar, et al.
Publicado: (2025)
por: Nareti, Utsav Kumar, et al.
Publicado: (2025)
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval
por: Zou, Qiang, et al.
Publicado: (2025)
por: Zou, Qiang, et al.
Publicado: (2025)
Modeling Musical Genre Trajectories through Pathlet Learning
por: Marey, Lilian, et al.
Publicado: (2025)
por: Marey, Lilian, et al.
Publicado: (2025)
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
por: Wang, Tianshi, et al.
Publicado: (2023)
por: Wang, Tianshi, et al.
Publicado: (2023)
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
por: Zhan, Hao, et al.
Publicado: (2026)
por: Zhan, Hao, et al.
Publicado: (2026)
Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation
por: Ong, Rongqing Kenneth, et al.
Publicado: (2024)
por: Ong, Rongqing Kenneth, et al.
Publicado: (2024)
Agentic Multi-Persona Framework for Evidence-Aware Fake News Detection
por: Bukke, Roopa, et al.
Publicado: (2025)
por: Bukke, Roopa, et al.
Publicado: (2025)
ChatDiet: Empowering Personalized Nutrition-Oriented Food Recommender Chatbots through an LLM-Augmented Framework
por: Yang, Zhongqi, et al.
Publicado: (2024)
por: Yang, Zhongqi, et al.
Publicado: (2024)
Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning
por: Bin, Yi, et al.
Publicado: (2024)
por: Bin, Yi, et al.
Publicado: (2024)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
por: Li, Yongqi, et al.
Publicado: (2024)
por: Li, Yongqi, et al.
Publicado: (2024)
EventCast: Hybrid Demand Forecasting in E-Commerce with LLM-Based Event Knowledge
por: Hu, Congcong, et al.
Publicado: (2026)
por: Hu, Congcong, et al.
Publicado: (2026)
Semantic Codebook Learning for Dynamic Recommendation Models
por: Lv, Zheqi, et al.
Publicado: (2024)
por: Lv, Zheqi, et al.
Publicado: (2024)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
por: Duan, Yue, et al.
Publicado: (2024)
por: Duan, Yue, et al.
Publicado: (2024)
VDCook:DIY video data cook your MLLMs
por: Wu, Chengwei
Publicado: (2026)
por: Wu, Chengwei
Publicado: (2026)
Multimodal Music Recommendation System using LLMs
por: Kandagatla, Srikar Prabhas, et al.
Publicado: (2026)
por: Kandagatla, Srikar Prabhas, et al.
Publicado: (2026)
On the Origin of Synthetic Information by Means of Steganographic Inheritance
por: Chang, Ching-Chun, et al.
Publicado: (2026)
por: Chang, Ching-Chun, et al.
Publicado: (2026)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
por: Gigant, Théo, et al.
Publicado: (2024)
por: Gigant, Théo, et al.
Publicado: (2024)
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
por: Ma, Hongjian, et al.
Publicado: (2026)
por: Ma, Hongjian, et al.
Publicado: (2026)
Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs
por: Zeng, Donghuo, et al.
Publicado: (2026)
por: Zeng, Donghuo, et al.
Publicado: (2026)
Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
por: Sun, Huatuan, et al.
Publicado: (2025)
por: Sun, Huatuan, et al.
Publicado: (2025)
Emotion-Driven Personalized Recommendation for AI-Generated Content Using Multi-Modal Sentiment and Intent Analysis
por: Hu, Zheqi, et al.
Publicado: (2025)
por: Hu, Zheqi, et al.
Publicado: (2025)
Ejemplares similares
-
Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
por: Kumar, Suraj, et al.
Publicado: (2025) -
Demystifying Visual Features of Movie Posters for Multi-Label Genre Identification
por: Nareti, Utsav Kumar, et al.
Publicado: (2023) -
Adaptive Data-Resilient Multi-Modal Hierarchical Multi-Label Book Genre Identification
por: Nareti, Utsav Kumar, et al.
Publicado: (2025) -
Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning
por: Nie, Zhijie, et al.
Publicado: (2024) -
CrossPT-EEG: A Benchmark for Cross-Participant and Cross-Time Generalization of EEG-based Visual Decoding
por: Zhu, Shuqi, et al.
Publicado: (2024)