Unified Multimodal Models as Auto-Encoders
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Zhiyuan, Lin, Kaiqing, Li, Zongjian, Ye, Junyan, Han, Hui, Wang, Haochen, Wang, Zhendong, Lin, Bin, Li, Hao, Xiao, Xinyan, Wang, Jingdong, Wang, Haifeng, Yuan, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ImgEdit: A Unified Image Editing Dataset and Benchmark
by: Ye, Yang, et al.
Published: (2025)
by: Ye, Yang, et al.
Published: (2025)
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
by: Zhu, Leqi, et al.
Published: (2026)
by: Zhu, Leqi, et al.
Published: (2026)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025)
by: Lin, Bin, et al.
Published: (2025)
Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection
by: Lin, Kaiqing, et al.
Published: (2025)
by: Lin, Kaiqing, et al.
Published: (2025)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
by: Niu, Yuwei, et al.
Published: (2025)
by: Niu, Yuwei, et al.
Published: (2025)
CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
by: Du, Yuxuan, et al.
Published: (2025)
by: Du, Yuxuan, et al.
Published: (2025)
Distribution Matching Variational AutoEncoder
by: Ye, Sen, et al.
Published: (2025)
by: Ye, Sen, et al.
Published: (2025)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
by: Song, Yuxin, et al.
Published: (2025)
by: Song, Yuxin, et al.
Published: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
by: Li, Zongjian, et al.
Published: (2025)
by: Li, Zongjian, et al.
Published: (2025)
Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection
by: Lin, Kaiqing, et al.
Published: (2024)
by: Lin, Kaiqing, et al.
Published: (2024)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024)
by: Chen, Liuhan, et al.
Published: (2024)
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
by: Liu, Jinlai, et al.
Published: (2025)
by: Liu, Jinlai, et al.
Published: (2025)
Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes
by: Lin, Kaiqing, et al.
Published: (2025)
by: Lin, Kaiqing, et al.
Published: (2025)
UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
by: Tang, Hongxuan, et al.
Published: (2025)
by: Tang, Hongxuan, et al.
Published: (2025)
UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
An Efficient Watermarking Method for Latent Diffusion Models via Low-Rank Adaptation and Dynamic Loss Weighting
by: Lin, Dongdong, et al.
Published: (2024)
by: Lin, Dongdong, et al.
Published: (2024)
UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
by: Zhuang, Shaobin, et al.
Published: (2026)
by: Zhuang, Shaobin, et al.
Published: (2026)
Unified Reward Model for Multimodal Understanding and Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
Helios: Real Real-Time Long Video Generation Model
by: Yuan, Shenghai, et al.
Published: (2026)
by: Yuan, Shenghai, et al.
Published: (2026)
Unified Map Prior Encoder for Mapping and Planning
by: Zhang, Zongzheng, et al.
Published: (2026)
by: Zhang, Zongzheng, et al.
Published: (2026)
Inference-Time Scaling for Visual AutoRegressive modeling by Searching Representative Samples
by: Tang, Weidong, et al.
Published: (2026)
by: Tang, Weidong, et al.
Published: (2026)
ConsistCompose: Unified Multimodal Layout Control for Image Composition
by: Shi, Xuanke, et al.
Published: (2025)
by: Shi, Xuanke, et al.
Published: (2025)
MVAR: MultiVariate AutoRegressive Air Pollutants Forecasting Model
by: Fan, Xu, et al.
Published: (2025)
by: Fan, Xu, et al.
Published: (2025)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
AMC: AutoML for Model Compression and Acceleration on Mobile Devices
by: He, Yihui, et al.
Published: (2018)
by: He, Yihui, et al.
Published: (2018)
Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
by: Han, Haochen, et al.
Published: (2024)
by: Han, Haochen, et al.
Published: (2024)
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
by: Li, Junxian, et al.
Published: (2026)
by: Li, Junxian, et al.
Published: (2026)
TIER: Text-Image Encoder-based Regression for AIGC Image Quality Assessment
by: Yuan, Jiquan, et al.
Published: (2024)
by: Yuan, Jiquan, et al.
Published: (2024)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
by: Xu, Linrui, et al.
Published: (2024)
by: Xu, Linrui, et al.
Published: (2024)
Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
by: Hao, Jitai, et al.
Published: (2025)
by: Hao, Jitai, et al.
Published: (2025)
TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP
by: Li, Fan, et al.
Published: (2025)
by: Li, Fan, et al.
Published: (2025)
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
by: Qin, Jie, et al.
Published: (2025)
by: Qin, Jie, et al.
Published: (2025)
EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
by: Dong, Ruixiao, et al.
Published: (2025)
by: Dong, Ruixiao, et al.
Published: (2025)
Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection
by: Li, Jiaming, et al.
Published: (2024)
by: Li, Jiaming, et al.
Published: (2024)
Representation Forcing for Bottleneck-Free Unified Multimodal Models
by: Wang, Yuqing, et al.
Published: (2026)
by: Wang, Yuqing, et al.
Published: (2026)
3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud Pretraining
by: Yan, Siming, et al.
Published: (2023)
by: Yan, Siming, et al.
Published: (2023)
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
by: Guo, Qingpei, et al.
Published: (2024)
by: Guo, Qingpei, et al.
Published: (2024)
Similar Items
-
ImgEdit: A Unified Image Editing Dataset and Benchmark
by: Ye, Yang, et al.
Published: (2025) -
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
by: Zhu, Leqi, et al.
Published: (2026) -
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025) -
Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection
by: Lin, Kaiqing, et al.
Published: (2025) -
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
by: Yan, Zhiyuan, et al.
Published: (2025)