Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Jingjing, Si, Chongjie, Luo, Jun, Zhang, Hanwang, Ma, Chao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
von: Wang, Kangsheng, et al.
Veröffentlicht: (2025)
von: Wang, Kangsheng, et al.
Veröffentlicht: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
von: Ghaleb, Esam, et al.
Veröffentlicht: (2025)
von: Ghaleb, Esam, et al.
Veröffentlicht: (2025)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing
von: Mofayezi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Mofayezi, Mohammadreza, et al.
Veröffentlicht: (2024)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
Multi-task Prompt Words Learning for Social Media Content Generation
von: Xue, Haochen, et al.
Veröffentlicht: (2024)
von: Xue, Haochen, et al.
Veröffentlicht: (2024)
How Far Are We from Generating Missing Modalities with Foundation Models?
von: Ke, Guanzhou, et al.
Veröffentlicht: (2025)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2025)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
von: Wang, Bing, et al.
Veröffentlicht: (2024)
von: Wang, Bing, et al.
Veröffentlicht: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
von: An, Wenbin, et al.
Veröffentlicht: (2025)
von: An, Wenbin, et al.
Veröffentlicht: (2025)
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
von: Stamatakis, Markos, et al.
Veröffentlicht: (2025)
von: Stamatakis, Markos, et al.
Veröffentlicht: (2025)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
von: Masumura, Ryo, et al.
Veröffentlicht: (2025)
von: Masumura, Ryo, et al.
Veröffentlicht: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
von: Wang, Bing, et al.
Veröffentlicht: (2025)
von: Wang, Bing, et al.
Veröffentlicht: (2025)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
Can Multimodal Large Language Models Understand Spatial Relations?
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
von: Ku, Max, et al.
Veröffentlicht: (2025)
von: Ku, Max, et al.
Veröffentlicht: (2025)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024) -
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2022) -
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025) -
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025) -
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)