Reconstruction Alignment Improves Unified Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Ji, Darrell, Trevor, Zettlemoyer, Luke, Wang, XuDong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025)
by: Yu, Junwei, et al.
Published: (2025)
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)
by: Tang, Zineng, et al.
Published: (2025)
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025)
by: Kutscher, Declan, et al.
Published: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023)
by: Goel, Vidit, et al.
Published: (2023)
TV2TV: A Unified Framework for Interleaved Language and Video Generation
by: Han, Xiaochuang, et al.
Published: (2025)
by: Han, Xiaochuang, et al.
Published: (2025)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
by: Shi, Weijia, et al.
Published: (2024)
by: Shi, Weijia, et al.
Published: (2024)
Lifting Embodied World Models for Planning and Control
by: Wang, Alex N., et al.
Published: (2026)
by: Wang, Alex N., et al.
Published: (2026)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024)
by: Huang, Brandon, et al.
Published: (2024)
Hidden in plain sight: VLMs overlook their visual representations
by: Fu, Stephanie, et al.
Published: (2025)
by: Fu, Stephanie, et al.
Published: (2025)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
by: Yasunaga, Michihiro, et al.
Published: (2025)
by: Yasunaga, Michihiro, et al.
Published: (2025)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
SegLLM: Multi-round Reasoning Segmentation
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
MCLR: Improving Conditional Modeling via Inter-Class Likelihood-Ratio Maximization and Unifying Classifier-Free Guidance with Alignment Objectives
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Shape-Guided Diffusion with Inside-Outside Attention
by: Park, Dong Huk, et al.
Published: (2022)
by: Park, Dong Huk, et al.
Published: (2022)
UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
by: Su, Zhaolong, et al.
Published: (2025)
by: Su, Zhaolong, et al.
Published: (2025)
Simulating the Real World: A Unified Survey of Multimodal Generative Models
by: Hu, Yuqi, et al.
Published: (2025)
by: Hu, Yuqi, et al.
Published: (2025)
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
by: Shang, Chuyi, et al.
Published: (2024)
by: Shang, Chuyi, et al.
Published: (2024)
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
Precipitation Nowcasting Using Diffusion Transformer with Causal Attention
by: Li, ChaoRong, et al.
Published: (2024)
by: Li, ChaoRong, et al.
Published: (2024)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026)
by: Chen, Leon Liangyu, et al.
Published: (2026)
Gramian Multimodal Representation Learning and Alignment
by: Cicchetti, Giordano, et al.
Published: (2024)
by: Cicchetti, Giordano, et al.
Published: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
by: Miao, Yanting, et al.
Published: (2026)
by: Miao, Yanting, et al.
Published: (2026)
Visually Prompted Benchmarks Are Surprisingly Fragile
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024)
by: Lu, Shiyin, et al.
Published: (2024)
MoDE: CLIP Data Experts via Clustering
by: Ma, Jiawei, et al.
Published: (2024)
by: Ma, Jiawei, et al.
Published: (2024)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
by: Zhan, Jun, et al.
Published: (2024)
by: Zhan, Jun, et al.
Published: (2024)
ALOHa: A New Measure for Hallucination in Captioning Models
by: Petryk, Suzanne, et al.
Published: (2024)
by: Petryk, Suzanne, et al.
Published: (2024)
Fantastic Copyrighted Beasts and How (Not) to Generate Them
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
Unified Vision-Language Modeling via Concept Space Alignment
by: Qiu, Yifu, et al.
Published: (2026)
by: Qiu, Yifu, et al.
Published: (2026)
Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models
by: Rao, Abinav, et al.
Published: (2026)
by: Rao, Abinav, et al.
Published: (2026)
Similar Items
-
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025) -
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024) -
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025) -
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024) -
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)