HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiang, Zhang, Zhifei, Zhang, He, Lin, Zhe, Zhou, Yuqian, Liu, Qing, Zhang, Shiwei, Li, Yijun, Liu, Shaoteng, Zheng, Haitian, Kuen, Jason, Wang, Yuehuan, Gao, Changxin, Sang, Nong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915637552480256
author Wang, Xiang
Zhang, Zhifei
Zhang, He
Lin, Zhe
Zhou, Yuqian
Liu, Qing
Zhang, Shiwei
Li, Yijun
Liu, Shaoteng
Zheng, Haitian
Kuen, Jason
Wang, Yuehuan
Gao, Changxin
Sang, Nong
author_facet Wang, Xiang
Zhang, Zhifei
Zhang, He
Lin, Zhe
Zhou, Yuqian
Liu, Qing
Zhang, Shiwei
Li, Yijun
Liu, Shaoteng
Zheng, Haitian
Kuen, Jason
Wang, Yuehuan
Gao, Changxin
Sang, Nong
contents Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanced methods such as BAGEL and LMFusion follow the Mixture-of-Transformers (MoT) paradigm, adopting a symmetric design that mirrors one expert to another for convenient initialization and fusion, which remains suboptimal due to inherent modality discrepancies. In this work, we propose HBridge, an asymmetric H-shaped architecture that enables heterogeneous experts to optimally leverage pretrained priors from their respective modality domains. Unlike prior dense fusion strategies that straightforwardly connect all layers between experts via shared attention, HBridge selectively bridges intermediate layers, reducing over 40% attention sharing, which improves efficiency and enhances generation quality. Shallow and deep layers, which capture modality-specific representations, are decoupled, while mid-layer bridging promotes semantic alignment. To further strengthen cross-modal coherence, we introduce semantic reconstruction tokens that explicitly guide the generative expert to reconstruct visual semantic tokens of the target image. Extensive experiments across multiple benchmarks demonstrate the effectiveness and superior performance of HBridge, establishing a new paradigm for unified multimodal generation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20520
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
Wang, Xiang
Zhang, Zhifei
Zhang, He
Lin, Zhe
Zhou, Yuqian
Liu, Qing
Zhang, Shiwei
Li, Yijun
Liu, Shaoteng
Zheng, Haitian
Kuen, Jason
Wang, Yuehuan
Gao, Changxin
Sang, Nong
Computer Vision and Pattern Recognition
Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanced methods such as BAGEL and LMFusion follow the Mixture-of-Transformers (MoT) paradigm, adopting a symmetric design that mirrors one expert to another for convenient initialization and fusion, which remains suboptimal due to inherent modality discrepancies. In this work, we propose HBridge, an asymmetric H-shaped architecture that enables heterogeneous experts to optimally leverage pretrained priors from their respective modality domains. Unlike prior dense fusion strategies that straightforwardly connect all layers between experts via shared attention, HBridge selectively bridges intermediate layers, reducing over 40% attention sharing, which improves efficiency and enhances generation quality. Shallow and deep layers, which capture modality-specific representations, are decoupled, while mid-layer bridging promotes semantic alignment. To further strengthen cross-modal coherence, we introduce semantic reconstruction tokens that explicitly guide the generative expert to reconstruct visual semantic tokens of the target image. Extensive experiments across multiple benchmarks demonstrate the effectiveness and superior performance of HBridge, establishing a new paradigm for unified multimodal generation.
title HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20520