Saved in:
| Main Authors: | Niu, Yuwei, Jin, Weiyang, Liao, Jiaqi, Feng, Chaoran, Jin, Peng, Lin, Bin, Li, Zongjian, Zhu, Bin, Yu, Weihao, Yuan, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.20561 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
by: Niu, Yuwei, et al.
Published: (2025)
by: Niu, Yuwei, et al.
Published: (2025)
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
by: Jin, Weiyang, et al.
Published: (2025)
by: Jin, Weiyang, et al.
Published: (2025)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025)
by: Lin, Bin, et al.
Published: (2025)
Unified Reward Model for Multimodal Understanding and Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
ImgEdit: A Unified Image Editing Dataset and Benchmark
by: Ye, Yang, et al.
Published: (2025)
by: Ye, Yang, et al.
Published: (2025)
Unified Multimodal Models as Auto-Encoders
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024)
by: Chen, Liuhan, et al.
Published: (2024)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
by: Li, Zongjian, et al.
Published: (2025)
by: Li, Zongjian, et al.
Published: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
iFSQ: Improving FSQ for Image Generation with 1 Line of Code
by: Lin, Bin, et al.
Published: (2026)
by: Lin, Bin, et al.
Published: (2026)
LLMBind: A Unified Modality-Task Integration Framework
by: Zhu, Bin, et al.
Published: (2024)
by: Zhu, Bin, et al.
Published: (2024)
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
by: Lin, Bin, et al.
Published: (2023)
by: Lin, Bin, et al.
Published: (2023)
UniNote: A Unified Embedding Model for Multimodal Representation and Ranking
by: Zhao, Jinghan, et al.
Published: (2026)
by: Zhao, Jinghan, et al.
Published: (2026)
Chatlaw: A Multi-Agent Collaborative Legal Assistant with Knowledge Graph Enhanced Mixture-of-Experts Large Language Model
by: Cui, Jiaxi, et al.
Published: (2023)
by: Cui, Jiaxi, et al.
Published: (2023)
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
by: Liao, Kang, et al.
Published: (2025)
by: Liao, Kang, et al.
Published: (2025)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
by: Jin, Peng, et al.
Published: (2023)
by: Jin, Peng, et al.
Published: (2023)
Emerging Properties in Unified Multimodal Pretraining
by: Deng, Chaorui, et al.
Published: (2025)
by: Deng, Chaorui, et al.
Published: (2025)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
by: Shi, Yongxin, et al.
Published: (2025)
by: Shi, Yongxin, et al.
Published: (2025)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
UMBRAE: Unified Multimodal Brain Decoding
by: Xia, Weihao, et al.
Published: (2024)
by: Xia, Weihao, et al.
Published: (2024)
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
by: Fu, Teng, et al.
Published: (2025)
by: Fu, Teng, et al.
Published: (2025)
Helios: Real Real-Time Long Video Generation Model
by: Yuan, Shenghai, et al.
Published: (2026)
by: Yuan, Shenghai, et al.
Published: (2026)
Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker
by: Wu, Zongjian, et al.
Published: (2026)
by: Wu, Zongjian, et al.
Published: (2026)
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving
by: Hu, Bin, et al.
Published: (2025)
by: Hu, Bin, et al.
Published: (2025)
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
by: Tang, Zhenyu, et al.
Published: (2024)
by: Tang, Zhenyu, et al.
Published: (2024)
RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba
by: Lu, Andong, et al.
Published: (2024)
by: Lu, Andong, et al.
Published: (2024)
UMIT: Unifying Medical Imaging Tasks via Vision-Language Models
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
DeblurNVS: Geometric Latent Diffusion for Novel View Synthesis from Sparse Motion-Blurred Images
by: Shi, Changyue, et al.
Published: (2026)
by: Shi, Changyue, et al.
Published: (2026)
UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass
by: Li, Mengfei, et al.
Published: (2026)
by: Li, Mengfei, et al.
Published: (2026)
LangBridge: Interpreting Image as a Combination of Language Embeddings
by: Liao, Jiaqi, et al.
Published: (2025)
by: Liao, Jiaqi, et al.
Published: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
by: Cheng, Zihui, et al.
Published: (2025)
by: Cheng, Zihui, et al.
Published: (2025)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
by: Xie, Jinheng, et al.
Published: (2024)
by: Xie, Jinheng, et al.
Published: (2024)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
Video Understanding: From Geometry and Semantics to Unified Models
by: An, Zhaochong, et al.
Published: (2026)
by: An, Zhaochong, et al.
Published: (2026)
Similar Items
-
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
by: Niu, Yuwei, et al.
Published: (2025) -
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
by: Jin, Weiyang, et al.
Published: (2025) -
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025) -
Unified Reward Model for Multimodal Understanding and Generation
by: Wang, Yibin, et al.
Published: (2025) -
ImgEdit: A Unified Image Editing Dataset and Benchmark
by: Ye, Yang, et al.
Published: (2025)