Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Ziyuan, Zheng, DanDan, Zou, Cheng, Liu, Rui, Wang, Xiaolong, Ji, Kaixiang, Chai, Weilong, Sun, Jianxin, Wang, Libin, Lv, Yongjie, Huang, Taozhi, Liu, Jiajia, Guo, Qingpei, Yang, Ming, Chen, Jingdong, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
by: Yan, Canxiang, et al.
Published: (2025)
by: Yan, Canxiang, et al.
Published: (2025)
UniVision: A Unified Framework for Vision-Centric 3D Perception
by: Hong, Yu, et al.
Published: (2024)
by: Hong, Yu, et al.
Published: (2024)
Ming-Omni: A Unified Multimodal Model for Perception and Generation
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models
by: Chai, Weilong, et al.
Published: (2023)
by: Chai, Weilong, et al.
Published: (2023)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
6Diffusion: IPv6 Target Generation Using a Diffusion Model with Global-Local Attention Mechanisms for Internet-wide IPv6 Scanning
by: He, Nabo, et al.
Published: (2024)
by: He, Nabo, et al.
Published: (2024)
UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
by: Song, Xinyang, et al.
Published: (2025)
by: Song, Xinyang, et al.
Published: (2025)
ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
by: Wang, Xiaolong, et al.
Published: (2025)
by: Wang, Xiaolong, et al.
Published: (2025)
VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
by: Jiang, Longteng, et al.
Published: (2026)
by: Jiang, Longteng, et al.
Published: (2026)
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
by: Zeng, Weili, et al.
Published: (2025)
by: Zeng, Weili, et al.
Published: (2025)
The evolution of commercial finance in Ming-Qing China: 16th to Early-20th Centuries
by: Kaixiang Peng
Published: (2023)
by: Kaixiang Peng
Published: (2023)
The Safety Analysis of Live Working on 220‐kV Double‐Circuit Line on the Drilling and Spanning Tower
by: Xiang Cai, et al.
Published: (2025)
by: Xiang Cai, et al.
Published: (2025)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations
by: Qiao, Qianqian, et al.
Published: (2025)
by: Qiao, Qianqian, et al.
Published: (2025)
Dynamics for a diffusive epidemic model with a free boundary: spreading-vanishing dichotomy
by: Li, Xueping, et al.
Published: (2024)
by: Li, Xueping, et al.
Published: (2024)
UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
by: Liu, Zhe, et al.
Published: (2025)
by: Liu, Zhe, et al.
Published: (2025)
GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
UniGeo: A Unified 3D Indoor Object Detection Framework Integrating Geometry-Aware Learning and Dynamic Channel Gating
by: Yi, Xing, et al.
Published: (2026)
by: Yi, Xing, et al.
Published: (2026)
3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
by: Song, Xinyang, et al.
Published: (2025)
by: Song, Xinyang, et al.
Published: (2025)
Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense Predictions
by: Zhang, Jingdong, et al.
Published: (2024)
by: Zhang, Jingdong, et al.
Published: (2024)
RLSLM: A Hybrid Reinforcement Learning Framework Aligning Rule-Based Social Locomotion Model with Human Social Norms
by: Kou, Yitian, et al.
Published: (2025)
by: Kou, Yitian, et al.
Published: (2025)
PPS-QMIX: Periodically Parameter Sharing for Accelerating Convergence of Multi-Agent Reinforcement Learning
by: Zhang, Ke, et al.
Published: (2024)
by: Zhang, Ke, et al.
Published: (2024)
Rational Fabrication of One‐Dimensional TiO 2 Nanowires for Enhanced Supercapacitor and Rechargeable Lithium Ion Battery
by: Xinyi Li, et al.
Published: (2025)
by: Xinyi Li, et al.
Published: (2025)
CD74 Affects Ferroptosis in Traumatic Brain Injury by Modulating the Nrf2/HO‐1 Signaling Pathway
by: GuangWei Sun, et al.
Published: (2026)
by: GuangWei Sun, et al.
Published: (2026)
UniShare: A Unified Framework for Joint Video and Receiver Recommendation in Social Sharing
by: Wang, Caimeng, et al.
Published: (2026)
by: Wang, Caimeng, et al.
Published: (2026)
Development of a PCR ‐Cas12a‐ LFD visual detection system for highly sensitive and specific detection of Ralstonia sp. , Phytophthora sp. , Alternaria sp. , and Pseudomonas sp. in tobacco
by: Chenqi Niu, et al.
Published: (2026)
by: Chenqi Niu, et al.
Published: (2026)
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
by: Huang, Ziyuan, et al.
Published: (2024)
by: Huang, Ziyuan, et al.
Published: (2024)
Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark
by: Zou, Kai, et al.
Published: (2025)
by: Zou, Kai, et al.
Published: (2025)
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
by: Yue, Zhengrong, et al.
Published: (2025)
by: Yue, Zhengrong, et al.
Published: (2025)
UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation
by: Lin, Jiaying, et al.
Published: (2026)
by: Lin, Jiaying, et al.
Published: (2026)
Genome identification of C‐type lectins in Bactrocera dorsalis and functional characterization of BdCTL‐S12 with a broad‐spectrum pathogen recognition
by: Wei Zhao, et al.
Published: (2025)
by: Wei Zhao, et al.
Published: (2025)
A Highly Compatible Deep Eutectic Solvent‐Based Poly(ethylene) Oxide Polymer Electrolyte to Enable the Stable Operation of 4.5 V Lithium Metal Batteries
by: Qi Liu, et al.
Published: (2024)
by: Qi Liu, et al.
Published: (2024)
UniARM: Towards a Unified Autoregressive Reward Model for Multi-Objective Test-Time Alignment
by: Xie, Hongyan, et al.
Published: (2026)
by: Xie, Hongyan, et al.
Published: (2026)
UniDex: Rethinking Search Inverted Indexing with Unified Semantic Modeling
by: Li, Zan, et al.
Published: (2025)
by: Li, Zan, et al.
Published: (2025)
UniMoT: Unified Molecule-Text Language Model with Discrete Token Representation
by: Guo, Shuhan, et al.
Published: (2024)
by: Guo, Shuhan, et al.
Published: (2024)
UniSER: A Foundation Model for Unified Soft Effects Removal
by: Zhang, Jingdong, et al.
Published: (2025)
by: Zhang, Jingdong, et al.
Published: (2025)
Similar Items
-
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
by: AI, Inclusion, et al.
Published: (2025) -
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025) -
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
by: Yan, Canxiang, et al.
Published: (2025) -
UniVision: A Unified Framework for Vision-Centric 3D Perception
by: Hong, Yu, et al.
Published: (2024) -
Ming-Omni: A Unified Multimodal Model for Perception and Generation
by: AI, Inclusion, et al.
Published: (2025)