UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Chenkai, Wang, Xu, Liao, Zhenyi, Li, Yishun, Hou, Tianqi, Deng, Zhijie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TLCM: Training-efficient Latent Consistency Model for Image Generation with 2-8 Steps
di: Xie, Qingsong, et al.
Pubblicazione: (2024)
di: Xie, Qingsong, et al.
Pubblicazione: (2024)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
di: Li, Yi, et al.
Pubblicazione: (2025)
di: Li, Yi, et al.
Pubblicazione: (2025)
UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
di: Li, Jinke, et al.
Pubblicazione: (2025)
di: Li, Jinke, et al.
Pubblicazione: (2025)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
di: Jiao, Yang, et al.
Pubblicazione: (2025)
di: Jiao, Yang, et al.
Pubblicazione: (2025)
Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models
di: Lu, Yishun, et al.
Pubblicazione: (2026)
di: Lu, Yishun, et al.
Pubblicazione: (2026)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
di: Li, Yiheng, et al.
Pubblicazione: (2024)
di: Li, Yiheng, et al.
Pubblicazione: (2024)
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
di: Wen, Zimo, et al.
Pubblicazione: (2026)
di: Wen, Zimo, et al.
Pubblicazione: (2026)
Understanding and Harnessing Sparsity in Unified Multimodal Models
di: He, Shwai, et al.
Pubblicazione: (2025)
di: He, Shwai, et al.
Pubblicazione: (2025)
Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
di: Qin, Luozheng, et al.
Pubblicazione: (2026)
di: Qin, Luozheng, et al.
Pubblicazione: (2026)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
di: Qu, Liao, et al.
Pubblicazione: (2024)
di: Qu, Liao, et al.
Pubblicazione: (2024)
Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing
di: Zhang, Weiyu, et al.
Pubblicazione: (2026)
di: Zhang, Weiyu, et al.
Pubblicazione: (2026)
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
di: Han, Ruiyan, et al.
Pubblicazione: (2026)
di: Han, Ruiyan, et al.
Pubblicazione: (2026)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
di: Zhang, Huichao, et al.
Pubblicazione: (2026)
di: Zhang, Huichao, et al.
Pubblicazione: (2026)
UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
di: Zhao, Xiangyu, et al.
Pubblicazione: (2024)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2024)
UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
di: Zhuang, Shaobin, et al.
Pubblicazione: (2026)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2026)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
di: Xiao, Yicheng, et al.
Pubblicazione: (2025)
di: Xiao, Yicheng, et al.
Pubblicazione: (2025)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
di: Liao, Zhenyi, et al.
Pubblicazione: (2025)
di: Liao, Zhenyi, et al.
Pubblicazione: (2025)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
di: Li, Junzhe, et al.
Pubblicazione: (2025)
di: Li, Junzhe, et al.
Pubblicazione: (2025)
Towards Understanding Deep Learning Model in Image Recognition via Coverage Test
di: Li, Wenkai, et al.
Pubblicazione: (2025)
di: Li, Wenkai, et al.
Pubblicazione: (2025)
UniPAR: A Unified Framework for Pedestrian Attribute Recognition
di: Xu, Minghe, et al.
Pubblicazione: (2026)
di: Xu, Minghe, et al.
Pubblicazione: (2026)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
di: Liu, Zeyu, et al.
Pubblicazione: (2026)
di: Liu, Zeyu, et al.
Pubblicazione: (2026)
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
di: Shi, Jin, et al.
Pubblicazione: (2026)
di: Shi, Jin, et al.
Pubblicazione: (2026)
LOVECon: Text-driven Training-Free Long Video Editing with ControlNet
di: Liao, Zhenyi, et al.
Pubblicazione: (2023)
di: Liao, Zhenyi, et al.
Pubblicazione: (2023)
Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
di: Fang, Chengyu, et al.
Pubblicazione: (2026)
di: Fang, Chengyu, et al.
Pubblicazione: (2026)
UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
di: Yang, Panqi, et al.
Pubblicazione: (2025)
di: Yang, Panqi, et al.
Pubblicazione: (2025)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
di: Chen, Leon Liangyu, et al.
Pubblicazione: (2026)
di: Chen, Leon Liangyu, et al.
Pubblicazione: (2026)
A Unified and General Framework for Continual Learning
di: Wang, Zhenyi, et al.
Pubblicazione: (2024)
di: Wang, Zhenyi, et al.
Pubblicazione: (2024)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
di: LASA Team, et al.
Pubblicazione: (2025)
di: LASA Team, et al.
Pubblicazione: (2025)
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
di: Chen, Zisheng, et al.
Pubblicazione: (2025)
di: Chen, Zisheng, et al.
Pubblicazione: (2025)
UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
di: Su, Zhaolong, et al.
Pubblicazione: (2025)
di: Su, Zhaolong, et al.
Pubblicazione: (2025)
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
di: Jiang, Yibo, et al.
Pubblicazione: (2026)
di: Jiang, Yibo, et al.
Pubblicazione: (2026)
AdvDMD: Adversarial Reward Meets DMD For High-Quality Few-Step Generation
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Understanding vs. Generation: Navigating Optimization Dilemma in Multimodal Models
di: Ye, Sen, et al.
Pubblicazione: (2026)
di: Ye, Sen, et al.
Pubblicazione: (2026)
Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
di: Hao, Jitai, et al.
Pubblicazione: (2025)
di: Hao, Jitai, et al.
Pubblicazione: (2025)
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
di: Mao, Jiawei, et al.
Pubblicazione: (2025)
di: Mao, Jiawei, et al.
Pubblicazione: (2025)
Semantic Generative Tuning for Unified Multimodal Models
di: Yu, Songsong, et al.
Pubblicazione: (2026)
di: Yu, Songsong, et al.
Pubblicazione: (2026)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
di: Lin, Bin, et al.
Pubblicazione: (2025)
di: Lin, Bin, et al.
Pubblicazione: (2025)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
di: Zhang, Jihai, et al.
Pubblicazione: (2025)
di: Zhang, Jihai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TLCM: Training-efficient Latent Consistency Model for Image Generation with 2-8 Steps
di: Xie, Qingsong, et al.
Pubblicazione: (2024) -
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
di: Li, Yi, et al.
Pubblicazione: (2025) -
UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
di: Li, Jinke, et al.
Pubblicazione: (2025) -
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
di: Jiao, Yang, et al.
Pubblicazione: (2025) -
Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models
di: Lu, Yishun, et al.
Pubblicazione: (2026)