From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Cheng, Shi, Chufan, Shui, Bo, Wu, Yaokang, Tao, Muzi, Wang, Huijuan, Lee, Ivan Yee, Liu, Yong, Ma, Xuezhe, Berg-Kirkpatrick, Taylor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026)
by: Shi, Chufan, et al.
Published: (2026)
Asymmetric Idiosyncrasies in Multimodal Models
by: Tao, Muzi, et al.
Published: (2026)
by: Tao, Muzi, et al.
Published: (2026)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
by: Gao, Xin, et al.
Published: (2026)
by: Gao, Xin, et al.
Published: (2026)
Optical Context Compression Is Just (Bad) Autoencoding
by: Lee, Ivan Yee, et al.
Published: (2025)
by: Lee, Ivan Yee, et al.
Published: (2025)
The Format Tax
by: Lee, Ivan Yee, et al.
Published: (2026)
by: Lee, Ivan Yee, et al.
Published: (2026)
LLM2: Let Large Language Models Harness System 2 Reasoning
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
ContextVis: Envision Contextual Learning and Interaction with Generative Models
by: Shui, Bo, et al.
Published: (2024)
by: Shui, Bo, et al.
Published: (2024)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
by: Lee, Ivan, et al.
Published: (2025)
by: Lee, Ivan, et al.
Published: (2025)
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
by: Chen, Danlu, et al.
Published: (2024)
by: Chen, Danlu, et al.
Published: (2024)
PixelLM: Pixel Reasoning with Large Multimodal Model
by: Ren, Zhongwei, et al.
Published: (2023)
by: Ren, Zhongwei, et al.
Published: (2023)
UEval: A Benchmark for Unified Multimodal Generation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
by: Sun, Kaiser, et al.
Published: (2026)
by: Sun, Kaiser, et al.
Published: (2026)
Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities
by: Wen, Nuan, et al.
Published: (2026)
by: Wen, Nuan, et al.
Published: (2026)
Alt-Text with Context: Improving Accessibility for Images on Twitter
by: Srivatsan, Nikita, et al.
Published: (2023)
by: Srivatsan, Nikita, et al.
Published: (2023)
Studying the Soupability of Documents in State Space Models
by: Jafari, Yasaman, et al.
Published: (2025)
by: Jafari, Yasaman, et al.
Published: (2025)
MORL-Prompt: An Empirical Analysis of Multi-Objective Reinforcement Learning for Discrete Prompt Optimization
by: Jafari, Yasaman, et al.
Published: (2024)
by: Jafari, Yasaman, et al.
Published: (2024)
MERBench: A Unified Evaluation Benchmark for Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2024)
by: Lian, Zheng, et al.
Published: (2024)
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
by: Luo, Ruilin, et al.
Published: (2026)
by: Luo, Ruilin, et al.
Published: (2026)
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
by: Wang, Zhenglin, et al.
Published: (2025)
by: Wang, Zhenglin, et al.
Published: (2025)
HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection
by: Zhou, Han, et al.
Published: (2026)
by: Zhou, Han, et al.
Published: (2026)
Conversational Alignment with Artificial Intelligence in Context
by: Sterken, Rachel Katharine, et al.
Published: (2025)
by: Sterken, Rachel Katharine, et al.
Published: (2025)
Bridging the Arithmetic Gap: The Cognitive Complexity Benchmark and Financial-PoT for Robust Financial Reasoning
by: Zhao, Boxiang, et al.
Published: (2026)
by: Zhao, Boxiang, et al.
Published: (2026)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises
by: Xu, Nan, et al.
Published: (2024)
by: Xu, Nan, et al.
Published: (2024)
From sunblock to softblock: Analyzing the correlates of neology in published writing and on social media
by: Ryskina, Maria, et al.
Published: (2026)
by: Ryskina, Maria, et al.
Published: (2026)
PixelBytes: Catching Unified Representation for Multimodal Generation
by: Furfaro, Fabien
Published: (2024)
by: Furfaro, Fabien
Published: (2024)
PixelBytes: Catching Unified Embedding for Multimodal Generation
by: Furfaro, Fabien
Published: (2024)
by: Furfaro, Fabien
Published: (2024)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
by: Ryan, Yuriel, et al.
Published: (2025)
by: Ryan, Yuriel, et al.
Published: (2025)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
by: Shi, Mengqi, et al.
Published: (2026)
by: Shi, Mengqi, et al.
Published: (2026)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems
by: Xu, Nan, et al.
Published: (2024)
by: Xu, Nan, et al.
Published: (2024)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
by: Jiang, Houcheng, et al.
Published: (2026)
by: Jiang, Houcheng, et al.
Published: (2026)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
Constrained Adaptive Rejection Sampling
by: Parys, Paweł, et al.
Published: (2025)
by: Parys, Paweł, et al.
Published: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
EMO: Frustratingly Easy Progressive Training of Extendable MoE
by: Jin, Linghao, et al.
Published: (2026)
by: Jin, Linghao, et al.
Published: (2026)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
by: Xia, Yinan, et al.
Published: (2025)
by: Xia, Yinan, et al.
Published: (2025)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
by: Stogiannidis, Ilias, et al.
Published: (2025)
by: Stogiannidis, Ilias, et al.
Published: (2025)
Similar Items
-
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026) -
Asymmetric Idiosyncrasies in Multimodal Models
by: Tao, Muzi, et al.
Published: (2026) -
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
by: Gao, Xin, et al.
Published: (2026) -
Optical Context Compression Is Just (Bad) Autoencoding
by: Lee, Ivan Yee, et al.
Published: (2025) -
The Format Tax
by: Lee, Ivan Yee, et al.
Published: (2026)