Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chenlong, Chen, Yuhang, Hu, Zhihan, Chen, Dongping, Chen, Wenhu, Wiegreffe, Sarah, Zhou, Tianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
von: Wang, Chenlong, et al.
Veröffentlicht: (2025)
von: Wang, Chenlong, et al.
Veröffentlicht: (2025)
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
von: Chen, Dongping, et al.
Veröffentlicht: (2026)
von: Chen, Dongping, et al.
Veröffentlicht: (2026)
Optimizing Length Compression in Large Reasoning Models
von: Cheng, Zhengxiang, et al.
Veröffentlicht: (2025)
von: Cheng, Zhengxiang, et al.
Veröffentlicht: (2025)
DataGen: Unified Synthetic Dataset Generation via Large Language Models
von: Huang, Yue, et al.
Veröffentlicht: (2024)
von: Huang, Yue, et al.
Veröffentlicht: (2024)
A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
Mechanistic?
von: Saphra, Naomi, et al.
Veröffentlicht: (2024)
von: Saphra, Naomi, et al.
Veröffentlicht: (2024)
Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion
von: Xu, Yuanfeng, et al.
Veröffentlicht: (2026)
von: Xu, Yuanfeng, et al.
Veröffentlicht: (2026)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
von: Chen, Xiaokang, et al.
Veröffentlicht: (2025)
von: Chen, Xiaokang, et al.
Veröffentlicht: (2025)
Reinforced Visual Perception with Tools
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
von: Ravisankar, Kartik, et al.
Veröffentlicht: (2025)
von: Ravisankar, Kartik, et al.
Veröffentlicht: (2025)
CODESYNC: Synchronizing Large Language Models with Dynamic Code Evolution at Scale
von: Wang, Chenlong, et al.
Veröffentlicht: (2025)
von: Wang, Chenlong, et al.
Veröffentlicht: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
von: Wang, Bin, et al.
Veröffentlicht: (2025)
von: Wang, Bin, et al.
Veröffentlicht: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
On Linear Representations and Pretraining Data Frequency in Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
von: Ruan, Chi, et al.
Veröffentlicht: (2025)
von: Ruan, Chi, et al.
Veröffentlicht: (2025)
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
von: Ku, Max, et al.
Veröffentlicht: (2025)
von: Ku, Max, et al.
Veröffentlicht: (2025)
On Fairness of Unified Multimodal Large Language Model for Image Generation
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
ChatSOS: Vector Database Augmented Generative Question Answering Assistant in Safety Engineering
von: Tang, Haiyang, et al.
Veröffentlicht: (2024)
von: Tang, Haiyang, et al.
Veröffentlicht: (2024)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
von: Song, Lin, et al.
Veröffentlicht: (2026)
von: Song, Lin, et al.
Veröffentlicht: (2026)
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
von: Wang, Yubo, et al.
Veröffentlicht: (2023)
von: Wang, Yubo, et al.
Veröffentlicht: (2023)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
A Survey of Multimodal Retrieval-Augmented Generation
von: Mei, Lang, et al.
Veröffentlicht: (2025)
von: Mei, Lang, et al.
Veröffentlicht: (2025)
Paper2Web: Let's Make Your Paper Alive!
von: Chen, Yuhang, et al.
Veröffentlicht: (2025)
von: Chen, Yuhang, et al.
Veröffentlicht: (2025)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
von: Ma, Yiyang, et al.
Veröffentlicht: (2024)
von: Ma, Yiyang, et al.
Veröffentlicht: (2024)
Multi-Objective Linguistic Control of Large Language Models
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection
von: Srikanth, Maya, et al.
Veröffentlicht: (2025)
von: Srikanth, Maya, et al.
Veröffentlicht: (2025)
URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
von: Shi, Yongxin, et al.
Veröffentlicht: (2025)
von: Shi, Yongxin, et al.
Veröffentlicht: (2025)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
von: Hase, Peter, et al.
Veröffentlicht: (2024)
von: Hase, Peter, et al.
Veröffentlicht: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
von: Wang, Chenlong, et al.
Veröffentlicht: (2025) -
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
von: Chen, Dongping, et al.
Veröffentlicht: (2026) -
Optimizing Length Compression in Large Reasoning Models
von: Cheng, Zhengxiang, et al.
Veröffentlicht: (2025) -
DataGen: Unified Synthetic Dataset Generation via Large Language Models
von: Huang, Yue, et al.
Veröffentlicht: (2024) -
A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
von: Xiao, Quan, et al.
Veröffentlicht: (2025)