Saved in:
| Main Authors: | Chen, Siran, Chen, Boyu, Yu, Chenyun, Luo, Yuxiao, Yi, Ouyang, Cheng, Lei, Zhuo, Chengxiang, Li, Zang, Wang, Yali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.02626 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior Simulation
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network
by: Chen, Zheyu, et al.
Published: (2024)
by: Chen, Zheyu, et al.
Published: (2024)
Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design
by: Yang, Yuxuan, et al.
Published: (2026)
by: Yang, Yuxuan, et al.
Published: (2026)
MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
by: Yuan, Xiang, et al.
Published: (2026)
by: Yuan, Xiang, et al.
Published: (2026)
Intent-Driven Semantic ID Generation for Grounded Conversational News Recommendation
by: Su, Hongyang, et al.
Published: (2026)
by: Su, Hongyang, et al.
Published: (2026)
Breaking the Curse of Knowledge: Towards Effective Multimodal Recommendation using Knowledge Soft Integration
by: Ouyang, Kai, et al.
Published: (2023)
by: Ouyang, Kai, et al.
Published: (2023)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
PRINTER:Deformation-Aware Adversarial Learning for Virtual IHC Staining with In Situ Fidelity
by: Yuan, Yizhe, et al.
Published: (2025)
by: Yuan, Yizhe, et al.
Published: (2025)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights
by: Yang, Jasmine, et al.
Published: (2026)
by: Yang, Jasmine, et al.
Published: (2026)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
by: Li, Qingcao, et al.
Published: (2026)
by: Li, Qingcao, et al.
Published: (2026)
Joint Optimization of Buffer Delay and HARQ for Video Communications
by: Cheng, Baoping, et al.
Published: (2024)
by: Cheng, Baoping, et al.
Published: (2024)
Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings
by: Liu, Shimiao, et al.
Published: (2025)
by: Liu, Shimiao, et al.
Published: (2025)
FairStream: Fair Multimedia Streaming Benchmark for Reinforcement Learning Agents
by: Weil, Jannis, et al.
Published: (2024)
by: Weil, Jannis, et al.
Published: (2024)
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
by: Rong, Yan, et al.
Published: (2025)
by: Rong, Yan, et al.
Published: (2025)
Startup Delay Aware Short Video Ordering: Problem, Model, and A Reinforcement Learning based Algorithm
by: Gao, Zhipeng, et al.
Published: (2024)
by: Gao, Zhipeng, et al.
Published: (2024)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
by: Duan, Chengqi, et al.
Published: (2025)
by: Duan, Chengqi, et al.
Published: (2025)
Anchoring Trends: Mitigating Social Media Popularity Prediction Drift via Feature Clustering and Expansion
by: Lee, Chia-Ming, et al.
Published: (2025)
by: Lee, Chia-Ming, et al.
Published: (2025)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection
by: Gu, Yimeng, et al.
Published: (2025)
by: Gu, Yimeng, et al.
Published: (2025)
Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation
by: Yan, Mingxuan, et al.
Published: (2024)
by: Yan, Mingxuan, et al.
Published: (2024)
QoE Optimization for Semantic Self-Correcting Video Transmission in Multi-UAV Networks
by: Chen, Xuyang, et al.
Published: (2025)
by: Chen, Xuyang, et al.
Published: (2025)
PREMISE: Matching-based Prediction for Accurate Review Recommendation
by: Han, Wei, et al.
Published: (2025)
by: Han, Wei, et al.
Published: (2025)
Generative Video Compression: Towards 0.01% Compression Rate for Video Transmission
by: Chen, Xiangyu, et al.
Published: (2025)
by: Chen, Xiangyu, et al.
Published: (2025)
Balancing Semantic Relevance and Engagement in Related Video Recommendations
by: Jaspal, Amit, et al.
Published: (2025)
by: Jaspal, Amit, et al.
Published: (2025)
Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
by: Chen, Yi-Chun
Published: (2025)
by: Chen, Yi-Chun
Published: (2025)
Symmetric Entropy-Constrained Video Coding for Machines
by: Sun, Yuxiao, et al.
Published: (2025)
by: Sun, Yuxiao, et al.
Published: (2025)
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
by: Guan, Yiwen, et al.
Published: (2025)
by: Guan, Yiwen, et al.
Published: (2025)
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
by: Wang, Junxi, et al.
Published: (2025)
by: Wang, Junxi, et al.
Published: (2025)
SAGER: Self-Evolving User Policy Skills for Recommendation Agent
by: Tao, Zhen, et al.
Published: (2026)
by: Tao, Zhen, et al.
Published: (2026)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
by: Zhu, Xiaofei, et al.
Published: (2024)
by: Zhu, Xiaofei, et al.
Published: (2024)
Dual-Diffusional Generative Fashion Recommendation
by: Yu, Mingzhe, et al.
Published: (2026)
by: Yu, Mingzhe, et al.
Published: (2026)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
Artic: AI-oriented Real-time Communication for MLLM Video Assistant
by: Wu, Jiangkai, et al.
Published: (2026)
by: Wu, Jiangkai, et al.
Published: (2026)
Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction
by: Wang, Dali, et al.
Published: (2026)
by: Wang, Dali, et al.
Published: (2026)
Understanding Before Recommendation: Semantic Aspect-Aware Review Exploitation via Large Language Models
by: Liu, Fan, et al.
Published: (2023)
by: Liu, Fan, et al.
Published: (2023)
Similar Items
-
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
by: Chen, Siran, et al.
Published: (2025) -
G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior Simulation
by: Chen, Boyu, et al.
Published: (2025) -
Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network
by: Chen, Zheyu, et al.
Published: (2024) -
Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design
by: Yang, Yuxuan, et al.
Published: (2026) -
MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
by: Yuan, Xiang, et al.
Published: (2026)