HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915637552480256 |
|---|---|
| author | Wang, Xiang Zhang, Zhifei Zhang, He Lin, Zhe Zhou, Yuqian Liu, Qing Zhang, Shiwei Li, Yijun Liu, Shaoteng Zheng, Haitian Kuen, Jason Wang, Yuehuan Gao, Changxin Sang, Nong |
| author_facet | Wang, Xiang Zhang, Zhifei Zhang, He Lin, Zhe Zhou, Yuqian Liu, Qing Zhang, Shiwei Li, Yijun Liu, Shaoteng Zheng, Haitian Kuen, Jason Wang, Yuehuan Gao, Changxin Sang, Nong |
| contents | Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanced methods such as BAGEL and LMFusion follow the Mixture-of-Transformers (MoT) paradigm, adopting a symmetric design that mirrors one expert to another for convenient initialization and fusion, which remains suboptimal due to inherent modality discrepancies. In this work, we propose HBridge, an asymmetric H-shaped architecture that enables heterogeneous experts to optimally leverage pretrained priors from their respective modality domains. Unlike prior dense fusion strategies that straightforwardly connect all layers between experts via shared attention, HBridge selectively bridges intermediate layers, reducing over 40% attention sharing, which improves efficiency and enhances generation quality. Shallow and deep layers, which capture modality-specific representations, are decoupled, while mid-layer bridging promotes semantic alignment. To further strengthen cross-modal coherence, we introduce semantic reconstruction tokens that explicitly guide the generative expert to reconstruct visual semantic tokens of the target image. Extensive experiments across multiple benchmarks demonstrate the effectiveness and superior performance of HBridge, establishing a new paradigm for unified multimodal generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_20520 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation Wang, Xiang Zhang, Zhifei Zhang, He Lin, Zhe Zhou, Yuqian Liu, Qing Zhang, Shiwei Li, Yijun Liu, Shaoteng Zheng, Haitian Kuen, Jason Wang, Yuehuan Gao, Changxin Sang, Nong Computer Vision and Pattern Recognition Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanced methods such as BAGEL and LMFusion follow the Mixture-of-Transformers (MoT) paradigm, adopting a symmetric design that mirrors one expert to another for convenient initialization and fusion, which remains suboptimal due to inherent modality discrepancies. In this work, we propose HBridge, an asymmetric H-shaped architecture that enables heterogeneous experts to optimally leverage pretrained priors from their respective modality domains. Unlike prior dense fusion strategies that straightforwardly connect all layers between experts via shared attention, HBridge selectively bridges intermediate layers, reducing over 40% attention sharing, which improves efficiency and enhances generation quality. Shallow and deep layers, which capture modality-specific representations, are decoupled, while mid-layer bridging promotes semantic alignment. To further strengthen cross-modal coherence, we introduce semantic reconstruction tokens that explicitly guide the generative expert to reconstruct visual semantic tokens of the target image. Extensive experiments across multiple benchmarks demonstrate the effectiveness and superior performance of HBridge, establishing a new paradigm for unified multimodal generation. |
| title | HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.20520 |