DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Renjie, Zhang, Xulong, Qu, Xiaoyang, Wang, Shangfei, Wang, Jianzong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
FakeBench: Probing Explainable Fake Image Detection via Large Multimodal Models
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
A Novel FACS-Aligned Anatomical Text Description Paradigm for Fine-Grained Facial Behavior Synthesis
von: Wang, Jiahe, et al.
Veröffentlicht: (2026)
von: Wang, Jiahe, et al.
Veröffentlicht: (2026)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
DIVA-VQA: Detecting Inter-frame Variations in UGC Video Quality
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
von: Chen, Sen, et al.
Veröffentlicht: (2022)
von: Chen, Sen, et al.
Veröffentlicht: (2022)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
von: Cai, Qi, et al.
Veröffentlicht: (2026)
von: Cai, Qi, et al.
Veröffentlicht: (2026)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
von: Masui, Kento, et al.
Veröffentlicht: (2024)
von: Masui, Kento, et al.
Veröffentlicht: (2024)
Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion
von: Jia, Tianzhi, et al.
Veröffentlicht: (2026)
von: Jia, Tianzhi, et al.
Veröffentlicht: (2026)
Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection
von: Yi, Kang, et al.
Veröffentlicht: (2025)
von: Yi, Kang, et al.
Veröffentlicht: (2025)
Predicting Satisfied User and Machine Ratio for Compressed Images: A Unified Approach
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
A Unified Non-Parametric and Interpretable Point Cloud Analysis via t-FCW Graph Representation
von: Lai, Haijian, et al.
Veröffentlicht: (2026)
von: Lai, Haijian, et al.
Veröffentlicht: (2026)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
von: Wang, Yuze, et al.
Veröffentlicht: (2025)
von: Wang, Yuze, et al.
Veröffentlicht: (2025)
EchoSR: Efficient Context Harnessing for Lightweight Image Super-Resolution
von: Zhao, Hanli, et al.
Veröffentlicht: (2026)
von: Zhao, Hanli, et al.
Veröffentlicht: (2026)
Detached and Interactive Multimodal Learning
von: Fan, Yunfeng, et al.
Veröffentlicht: (2024)
von: Fan, Yunfeng, et al.
Veröffentlicht: (2024)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating Assessment
von: Wang, Fengshun, et al.
Veröffentlicht: (2025)
von: Wang, Fengshun, et al.
Veröffentlicht: (2025)
PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
von: Zhang, Yongjian, et al.
Veröffentlicht: (2025)
von: Zhang, Yongjian, et al.
Veröffentlicht: (2025)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
von: Zhang, Chen-Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Chen-Lin, et al.
Veröffentlicht: (2025)
Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification
von: Wang, Yingquan, et al.
Veröffentlicht: (2024)
von: Wang, Yingquan, et al.
Veröffentlicht: (2024)
Self-supervised Photographic Image Layout Representation Learning
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)
Can Multimodal Large Language Models Understand Spatial Relations?
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
Unveiling Encoder-Free Vision-Language Models
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
von: Lu, Renjie, et al.
Veröffentlicht: (2026) -
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026) -
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025) -
FakeBench: Probing Explainable Fake Image Detection via Large Multimodal Models
von: Li, Yixuan, et al.
Veröffentlicht: (2024) -
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)