OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Size, Wu, Zhonghua, Gong, Zerui, Tao, Qingyi, Jin, Sheng, Li, Qinyue, Li, Wei, Loy, Chen Change |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
by: Gong, Zerui, et al.
Published: (2025)
by: Gong, Zerui, et al.
Published: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
by: Liao, Kang, et al.
Published: (2025)
by: Liao, Kang, et al.
Published: (2025)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023)
by: Xu, Shilin, et al.
Published: (2023)
F-LMM: Grounding Frozen Large Multimodal Models
by: Wu, Size, et al.
Published: (2024)
by: Wu, Size, et al.
Published: (2024)
Controllable Human-centric Keyframe Interpolation with Generative Prior
by: Guo, Zujin, et al.
Published: (2025)
by: Guo, Zujin, et al.
Published: (2025)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
by: Wu, Size, et al.
Published: (2023)
by: Wu, Size, et al.
Published: (2023)
Next Visual Granularity Generation
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
Learning Inclusion Matching for Animation Paint Bucket Colorization
by: Dai, Yuekun, et al.
Published: (2024)
by: Dai, Yuekun, et al.
Published: (2024)
MOWA: Multiple-in-One Image Warping Model
by: Liao, Kang, et al.
Published: (2024)
by: Liao, Kang, et al.
Published: (2024)
Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation
by: Zhao, Ran, et al.
Published: (2026)
by: Zhao, Ran, et al.
Published: (2026)
Paint Bucket Colorization Using Anime Character Color Design Sheets
by: Dai, Yuekun, et al.
Published: (2024)
by: Dai, Yuekun, et al.
Published: (2024)
MatAnyone: Stable Video Matting with Consistent Memory Propagation
by: Yang, Peiqing, et al.
Published: (2025)
by: Yang, Peiqing, et al.
Published: (2025)
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024)
by: Li, Xiangtai, et al.
Published: (2024)
AITTI: Learning Adaptive Inclusive Token for Text-to-Image Generation
by: Hou, Xinyu, et al.
Published: (2024)
by: Hou, Xinyu, et al.
Published: (2024)
Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
by: Li, Xiaoming, et al.
Published: (2025)
by: Li, Xiaoming, et al.
Published: (2025)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
by: Guo, Zujin, et al.
Published: (2024)
by: Guo, Zujin, et al.
Published: (2024)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
by: Li, Teng, et al.
Published: (2025)
by: Li, Teng, et al.
Published: (2025)
Control Color: Multimodal Diffusion-based Interactive Image Colorization
by: Liang, Zhexin, et al.
Published: (2024)
by: Liang, Zhexin, et al.
Published: (2024)
Contextual Object Detection with Multimodal Large Language Models
by: Zang, Yuhang, et al.
Published: (2023)
by: Zang, Yuhang, et al.
Published: (2023)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Kalman-Inspired Feature Propagation for Video Face Super-Resolution
by: Feng, Ruicheng, et al.
Published: (2024)
by: Feng, Ruicheng, et al.
Published: (2024)
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
by: AI, Inclusion, et al.
Published: (2026)
by: AI, Inclusion, et al.
Published: (2026)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Generative Photographic Control for Scene-Consistent Video Cinematic Editing
by: Sun, Huiqiang, et al.
Published: (2025)
by: Sun, Huiqiang, et al.
Published: (2025)
Uni-Sign: Toward Unified Sign Language Understanding at Scale
by: Li, Zecheng, et al.
Published: (2025)
by: Li, Zecheng, et al.
Published: (2025)
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
by: Yuan, Haobo, et al.
Published: (2024)
by: Yuan, Haobo, et al.
Published: (2024)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
by: Wei, Hongyang, et al.
Published: (2025)
by: Wei, Hongyang, et al.
Published: (2025)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
DifFace: Blind Face Restoration with Diffused Error Contraction
by: Yue, Zongsheng, et al.
Published: (2022)
by: Yue, Zongsheng, et al.
Published: (2022)
MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
by: Huang, Xiaoke, et al.
Published: (2025)
by: Huang, Xiaoke, et al.
Published: (2025)
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
by: Wang, Peiyu, et al.
Published: (2025)
by: Wang, Peiyu, et al.
Published: (2025)
UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
by: Wang, Dianyi, et al.
Published: (2026)
by: Wang, Dianyi, et al.
Published: (2026)
UniVideo: Unified Understanding, Generation, and Editing for Videos
by: Wei, Cong, et al.
Published: (2025)
by: Wei, Cong, et al.
Published: (2025)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
Similar Items
-
SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
by: Gong, Zerui, et al.
Published: (2025) -
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025) -
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
by: Liao, Kang, et al.
Published: (2025) -
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023) -
F-LMM: Grounding Frozen Large Multimodal Models
by: Wu, Size, et al.
Published: (2024)