From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Cheng, Shi, Chufan, Shui, Bo, Wu, Yaokang, Tao, Muzi, Wang, Huijuan, Lee, Ivan Yee, Liu, Yong, Ma, Xuezhe, Berg-Kirkpatrick, Taylor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
von: Shi, Chufan, et al.
Veröffentlicht: (2026)
von: Shi, Chufan, et al.
Veröffentlicht: (2026)
Asymmetric Idiosyncrasies in Multimodal Models
von: Tao, Muzi, et al.
Veröffentlicht: (2026)
von: Tao, Muzi, et al.
Veröffentlicht: (2026)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
von: Gao, Xin, et al.
Veröffentlicht: (2026)
von: Gao, Xin, et al.
Veröffentlicht: (2026)
Optical Context Compression Is Just (Bad) Autoencoding
von: Lee, Ivan Yee, et al.
Veröffentlicht: (2025)
von: Lee, Ivan Yee, et al.
Veröffentlicht: (2025)
The Format Tax
von: Lee, Ivan Yee, et al.
Veröffentlicht: (2026)
von: Lee, Ivan Yee, et al.
Veröffentlicht: (2026)
LLM2: Let Large Language Models Harness System 2 Reasoning
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
ContextVis: Envision Contextual Learning and Interaction with Generative Models
von: Shui, Bo, et al.
Veröffentlicht: (2024)
von: Shui, Bo, et al.
Veröffentlicht: (2024)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
von: Chen, Danlu, et al.
Veröffentlicht: (2024)
von: Chen, Danlu, et al.
Veröffentlicht: (2024)
PixelLM: Pixel Reasoning with Large Multimodal Model
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
UEval: A Benchmark for Unified Multimodal Generation
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities
von: Wen, Nuan, et al.
Veröffentlicht: (2026)
von: Wen, Nuan, et al.
Veröffentlicht: (2026)
Alt-Text with Context: Improving Accessibility for Images on Twitter
von: Srivatsan, Nikita, et al.
Veröffentlicht: (2023)
von: Srivatsan, Nikita, et al.
Veröffentlicht: (2023)
Studying the Soupability of Documents in State Space Models
von: Jafari, Yasaman, et al.
Veröffentlicht: (2025)
von: Jafari, Yasaman, et al.
Veröffentlicht: (2025)
MORL-Prompt: An Empirical Analysis of Multi-Objective Reinforcement Learning for Discrete Prompt Optimization
von: Jafari, Yasaman, et al.
Veröffentlicht: (2024)
von: Jafari, Yasaman, et al.
Veröffentlicht: (2024)
MERBench: A Unified Evaluation Benchmark for Multimodal Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2024)
von: Lian, Zheng, et al.
Veröffentlicht: (2024)
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025)
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025)
HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection
von: Zhou, Han, et al.
Veröffentlicht: (2026)
von: Zhou, Han, et al.
Veröffentlicht: (2026)
Conversational Alignment with Artificial Intelligence in Context
von: Sterken, Rachel Katharine, et al.
Veröffentlicht: (2025)
von: Sterken, Rachel Katharine, et al.
Veröffentlicht: (2025)
Bridging the Arithmetic Gap: The Cognitive Complexity Benchmark and Financial-PoT for Robust Financial Reasoning
von: Zhao, Boxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Boxiang, et al.
Veröffentlicht: (2026)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises
von: Xu, Nan, et al.
Veröffentlicht: (2024)
von: Xu, Nan, et al.
Veröffentlicht: (2024)
From sunblock to softblock: Analyzing the correlates of neology in published writing and on social media
von: Ryskina, Maria, et al.
Veröffentlicht: (2026)
von: Ryskina, Maria, et al.
Veröffentlicht: (2026)
PixelBytes: Catching Unified Representation for Multimodal Generation
von: Furfaro, Fabien
Veröffentlicht: (2024)
von: Furfaro, Fabien
Veröffentlicht: (2024)
PixelBytes: Catching Unified Embedding for Multimodal Generation
von: Furfaro, Fabien
Veröffentlicht: (2024)
von: Furfaro, Fabien
Veröffentlicht: (2024)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
von: Shi, Mengqi, et al.
Veröffentlicht: (2026)
von: Shi, Mengqi, et al.
Veröffentlicht: (2026)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems
von: Xu, Nan, et al.
Veröffentlicht: (2024)
von: Xu, Nan, et al.
Veröffentlicht: (2024)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
Constrained Adaptive Rejection Sampling
von: Parys, Paweł, et al.
Veröffentlicht: (2025)
von: Parys, Paweł, et al.
Veröffentlicht: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
EMO: Frustratingly Easy Progressive Training of Extendable MoE
von: Jin, Linghao, et al.
Veröffentlicht: (2026)
von: Jin, Linghao, et al.
Veröffentlicht: (2026)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
von: Shi, Chufan, et al.
Veröffentlicht: (2026) -
Asymmetric Idiosyncrasies in Multimodal Models
von: Tao, Muzi, et al.
Veröffentlicht: (2026) -
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
von: Gao, Xin, et al.
Veröffentlicht: (2026) -
Optical Context Compression Is Just (Bad) Autoencoding
von: Lee, Ivan Yee, et al.
Veröffentlicht: (2025) -
The Format Tax
von: Lee, Ivan Yee, et al.
Veröffentlicht: (2026)