Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Jingjing, Si, Chongjie, Luo, Jun, Zhang, Hanwang, Ma, Chao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
por: Luo, Chuwei, et al.
Publicado: (2022)
por: Luo, Chuwei, et al.
Publicado: (2022)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
por: Zhang, Xueqiao, et al.
Publicado: (2025)
por: Zhang, Xueqiao, et al.
Publicado: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
por: Gan, Chengguang, et al.
Publicado: (2025)
por: Gan, Chengguang, et al.
Publicado: (2025)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
por: Han, Jiaming, et al.
Publicado: (2025)
por: Han, Jiaming, et al.
Publicado: (2025)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
por: Wang, Xiao, et al.
Publicado: (2024)
por: Wang, Xiao, et al.
Publicado: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
por: Qu, Leigang, et al.
Publicado: (2024)
por: Qu, Leigang, et al.
Publicado: (2024)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
por: Wang, Xiao, et al.
Publicado: (2025)
por: Wang, Xiao, et al.
Publicado: (2025)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
por: Wang, Kangsheng, et al.
Publicado: (2025)
por: Wang, Kangsheng, et al.
Publicado: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
por: Wang, Xiao, et al.
Publicado: (2025)
por: Wang, Xiao, et al.
Publicado: (2025)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
por: Ghaleb, Esam, et al.
Publicado: (2025)
por: Ghaleb, Esam, et al.
Publicado: (2025)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
por: Chow, Wei, et al.
Publicado: (2024)
por: Chow, Wei, et al.
Publicado: (2024)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
por: Chen, Yanzhe, et al.
Publicado: (2025)
por: Chen, Yanzhe, et al.
Publicado: (2025)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
por: Wang, Jiapeng, et al.
Publicado: (2024)
por: Wang, Jiapeng, et al.
Publicado: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
por: Jiang, Chaoya, et al.
Publicado: (2024)
por: Jiang, Chaoya, et al.
Publicado: (2024)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
por: Lu, Renjie, et al.
Publicado: (2026)
por: Lu, Renjie, et al.
Publicado: (2026)
M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing
por: Mofayezi, Mohammadreza, et al.
Publicado: (2024)
por: Mofayezi, Mohammadreza, et al.
Publicado: (2024)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
por: Cheng, Zebang, et al.
Publicado: (2024)
por: Cheng, Zebang, et al.
Publicado: (2024)
Multi-task Prompt Words Learning for Social Media Content Generation
por: Xue, Haochen, et al.
Publicado: (2024)
por: Xue, Haochen, et al.
Publicado: (2024)
How Far Are We from Generating Missing Modalities with Foundation Models?
por: Ke, Guanzhou, et al.
Publicado: (2025)
por: Ke, Guanzhou, et al.
Publicado: (2025)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
por: Chen, Shuang, et al.
Publicado: (2026)
por: Chen, Shuang, et al.
Publicado: (2026)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
por: Zhang, Dongxu, et al.
Publicado: (2026)
por: Zhang, Dongxu, et al.
Publicado: (2026)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
por: Liang, Zhengyang, et al.
Publicado: (2024)
por: Liang, Zhengyang, et al.
Publicado: (2024)
Towards Event-oriented Long Video Understanding
por: Du, Yifan, et al.
Publicado: (2024)
por: Du, Yifan, et al.
Publicado: (2024)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
por: Wang, Bing, et al.
Publicado: (2024)
por: Wang, Bing, et al.
Publicado: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
por: Bin, Yi, et al.
Publicado: (2024)
por: Bin, Yi, et al.
Publicado: (2024)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
por: Song, Shezheng, et al.
Publicado: (2023)
por: Song, Shezheng, et al.
Publicado: (2023)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
por: An, Wenbin, et al.
Publicado: (2025)
por: An, Wenbin, et al.
Publicado: (2025)
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
por: Stamatakis, Markos, et al.
Publicado: (2025)
por: Stamatakis, Markos, et al.
Publicado: (2025)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
por: Masumura, Ryo, et al.
Publicado: (2025)
por: Masumura, Ryo, et al.
Publicado: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
por: Zhao, Zhixian, et al.
Publicado: (2026)
por: Zhao, Zhixian, et al.
Publicado: (2026)
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
por: Xu, Jiaqi, et al.
Publicado: (2024)
por: Xu, Jiaqi, et al.
Publicado: (2024)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
por: Wang, Bing, et al.
Publicado: (2025)
por: Wang, Bing, et al.
Publicado: (2025)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
por: Chen, Qian, et al.
Publicado: (2026)
por: Chen, Qian, et al.
Publicado: (2026)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
por: Luo, Ziyang, et al.
Publicado: (2024)
por: Luo, Ziyang, et al.
Publicado: (2024)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
por: Wu, Jiaying, et al.
Publicado: (2025)
por: Wu, Jiaying, et al.
Publicado: (2025)
Can Multimodal Large Language Models Understand Spatial Relations?
por: Liu, Jingping, et al.
Publicado: (2025)
por: Liu, Jingping, et al.
Publicado: (2025)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
por: Ku, Max, et al.
Publicado: (2025)
por: Ku, Max, et al.
Publicado: (2025)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
por: Li, Jinyuan, et al.
Publicado: (2024)
por: Li, Jinyuan, et al.
Publicado: (2024)
Ejemplares similares
-
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
por: Li, Yunxin, et al.
Publicado: (2024) -
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
por: Luo, Chuwei, et al.
Publicado: (2022) -
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
por: Zhang, Xueqiao, et al.
Publicado: (2025) -
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
por: Gan, Chengguang, et al.
Publicado: (2025) -
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
por: Han, Jiaming, et al.
Publicado: (2025)