AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Dian, Zhang, Manyuan, Li, Hongyu, Zou, Kai, Liu, Hongbo, Guo, Ziyu, Feng, Kaituo, Liu, Yexin, Luo, Ying, Li, Hongsheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
di: Zheng, Dian, et al.
Pubblicazione: (2026)
di: Zheng, Dian, et al.
Pubblicazione: (2026)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
di: Li, Hongyu, et al.
Pubblicazione: (2025)
di: Li, Hongyu, et al.
Pubblicazione: (2025)
Gen-Searcher: Reinforcing Agentic Search for Image Generation
di: Feng, Kaituo, et al.
Pubblicazione: (2026)
di: Feng, Kaituo, et al.
Pubblicazione: (2026)
EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding
di: Zou, Kai, et al.
Pubblicazione: (2026)
di: Zou, Kai, et al.
Pubblicazione: (2026)
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
di: Liu, Yexin, et al.
Pubblicazione: (2025)
di: Liu, Yexin, et al.
Pubblicazione: (2025)
Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark
di: Zou, Kai, et al.
Pubblicazione: (2025)
di: Zou, Kai, et al.
Pubblicazione: (2025)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
di: Zou, Kai, et al.
Pubblicazione: (2026)
di: Zou, Kai, et al.
Pubblicazione: (2026)
OneThinker: All-in-one Reasoning Model for Image and Video
di: Feng, Kaituo, et al.
Pubblicazione: (2025)
di: Feng, Kaituo, et al.
Pubblicazione: (2025)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
di: Chen, Zeyu, et al.
Pubblicazione: (2026)
di: Chen, Zeyu, et al.
Pubblicazione: (2026)
Learning A Multi-Task Transformer Via Unified And Customized Instruction Tuning For Chest Radiograph Interpretation
di: Xu, Lijian, et al.
Pubblicazione: (2023)
di: Xu, Lijian, et al.
Pubblicazione: (2023)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
di: Chen, Shuang, et al.
Pubblicazione: (2026)
di: Chen, Shuang, et al.
Pubblicazione: (2026)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
di: Liu, Xianjie, et al.
Pubblicazione: (2025)
di: Liu, Xianjie, et al.
Pubblicazione: (2025)
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
di: AI, Inclusion, et al.
Pubblicazione: (2025)
di: AI, Inclusion, et al.
Pubblicazione: (2025)
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
Rethinking Cell Counting Methods: Decoupling Counting and Localization
di: Zheng, Zixuan, et al.
Pubblicazione: (2025)
di: Zheng, Zixuan, et al.
Pubblicazione: (2025)
Rethinking VLM Representation for VLA Initialization
di: Lin, Weifeng, et al.
Pubblicazione: (2026)
di: Lin, Weifeng, et al.
Pubblicazione: (2026)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
di: Liu, Yexin, et al.
Pubblicazione: (2025)
di: Liu, Yexin, et al.
Pubblicazione: (2025)
Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models
di: Wu, Xiaoshi, et al.
Pubblicazione: (2024)
di: Wu, Xiaoshi, et al.
Pubblicazione: (2024)
Rethinking Normalization Strategies and Convolutional Kernels for Multimodal Image Fusion
di: He, Dan, et al.
Pubblicazione: (2024)
di: He, Dan, et al.
Pubblicazione: (2024)
AdaTooler-V: Adaptive Tool-Use for Images and Videos
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
Rethinking Early-Fusion Strategies for Improved Multimodal Image Segmentation
di: Shen, Zhengwen, et al.
Pubblicazione: (2025)
di: Shen, Zhengwen, et al.
Pubblicazione: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
di: Tian, Changyao, et al.
Pubblicazione: (2025)
di: Tian, Changyao, et al.
Pubblicazione: (2025)
AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing
di: Liu, Tianyu, et al.
Pubblicazione: (2026)
di: Liu, Tianyu, et al.
Pubblicazione: (2026)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
di: Liu, Jihao, et al.
Pubblicazione: (2024)
di: Liu, Jihao, et al.
Pubblicazione: (2024)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
di: Zhang, Wanpeng, et al.
Pubblicazione: (2025)
di: Zhang, Wanpeng, et al.
Pubblicazione: (2025)
PIDNet: Progressive Implicit Decouple Network for Multimodal Action Quality Assessment
di: Li, Qiqi, et al.
Pubblicazione: (2026)
di: Li, Qiqi, et al.
Pubblicazione: (2026)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
di: Hou, Ruibing, et al.
Pubblicazione: (2025)
di: Hou, Ruibing, et al.
Pubblicazione: (2025)
Towards Unified Modeling in Federated Multi-Task Learning via Subspace Decoupling
di: Wei, Yipan, et al.
Pubblicazione: (2025)
di: Wei, Yipan, et al.
Pubblicazione: (2025)
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
di: Liu, Dongyang, et al.
Pubblicazione: (2025)
di: Liu, Dongyang, et al.
Pubblicazione: (2025)
Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
di: Lin, Ziyue, et al.
Pubblicazione: (2026)
di: Lin, Ziyue, et al.
Pubblicazione: (2026)
Temporal Adaptive RGBT Tracking with Modality Prompt
di: Wang, Hongyu, et al.
Pubblicazione: (2024)
di: Wang, Hongyu, et al.
Pubblicazione: (2024)
LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
di: AI, Inclusion, et al.
Pubblicazione: (2025)
di: AI, Inclusion, et al.
Pubblicazione: (2025)
GoodSAM++: Bridging Domain and Capacity Gaps via Segment Anything Model for Panoramic Semantic Segmentation
di: Zhang, Weiming, et al.
Pubblicazione: (2024)
di: Zhang, Weiming, et al.
Pubblicazione: (2024)
GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-aware Panoramic Semantic Segmentation
di: Zhang, Weiming, et al.
Pubblicazione: (2024)
di: Zhang, Weiming, et al.
Pubblicazione: (2024)
Rethinking Decoupled Knowledge Distillation: A Predictive Distribution Perspective
di: Zheng, Bowen, et al.
Pubblicazione: (2025)
di: Zheng, Bowen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
di: Zheng, Dian, et al.
Pubblicazione: (2026) -
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
di: Li, Hongyu, et al.
Pubblicazione: (2025) -
Gen-Searcher: Reinforcing Agentic Search for Image Generation
di: Feng, Kaituo, et al.
Pubblicazione: (2026) -
EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding
di: Zou, Kai, et al.
Pubblicazione: (2026) -
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
di: Liu, Yexin, et al.
Pubblicazione: (2025)