UniAPO: Unified Multimodal Automated Prompt Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Qipeng, Chen, Yanzhe, Zhong, Huasong, Li, Yan, Chen, Jie, Zhang, Zhixin, Zhang, Junping, Yang, Zhenheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification
by: Qu, Xiangyan, et al.
Published: (2025)
by: Qu, Xiangyan, et al.
Published: (2025)
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
by: Zhao, Yaqi, et al.
Published: (2026)
by: Zhao, Yaqi, et al.
Published: (2026)
Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers
by: Zhu, Jingyuan, et al.
Published: (2026)
by: Zhu, Jingyuan, et al.
Published: (2026)
Show-o2: Improved Native Unified Multimodal Models
by: Xie, Jinheng, et al.
Published: (2025)
by: Xie, Jinheng, et al.
Published: (2025)
UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
by: Zhuang, Shaobin, et al.
Published: (2026)
by: Zhuang, Shaobin, et al.
Published: (2026)
UniVS: Unified and Universal Video Segmentation with Prompts as Queries
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
by: Yao, Yuxuan, et al.
Published: (2026)
by: Yao, Yuxuan, et al.
Published: (2026)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
by: Li, Teng, et al.
Published: (2025)
by: Li, Teng, et al.
Published: (2025)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
by: Jiao, Yang, et al.
Published: (2025)
by: Jiao, Yang, et al.
Published: (2025)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning
by: Lyu, Mengyao, et al.
Published: (2025)
by: Lyu, Mengyao, et al.
Published: (2025)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
UniVBench: Towards Unified Evaluation for Video Foundation Models
by: Wei, Jianhui, et al.
Published: (2026)
by: Wei, Jianhui, et al.
Published: (2026)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
by: Chen, Houyuan, et al.
Published: (2026)
by: Chen, Houyuan, et al.
Published: (2026)
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
by: Hong, Minjie, et al.
Published: (2025)
by: Hong, Minjie, et al.
Published: (2025)
MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception
by: Liu, Wenzhuo, et al.
Published: (2025)
by: Liu, Wenzhuo, et al.
Published: (2025)
ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
by: Tan, Jiangtong, et al.
Published: (2025)
by: Tan, Jiangtong, et al.
Published: (2025)
UniABG: Unified Adversarial View Bridging and Graph Correspondence for Unsupervised Cross-View Geo-Localization
by: Chen, Cuiqun, et al.
Published: (2025)
by: Chen, Cuiqun, et al.
Published: (2025)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
by: Xie, Jinheng, et al.
Published: (2024)
by: Xie, Jinheng, et al.
Published: (2024)
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
by: Wen, Zimo, et al.
Published: (2026)
by: Wen, Zimo, et al.
Published: (2026)
UniScene: Unified Occupancy-centric Driving Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
UniMesh: Unifying 3D Mesh Understanding and Generation
by: Huang, Peng, et al.
Published: (2026)
by: Huang, Peng, et al.
Published: (2026)
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
by: Li, Junxian, et al.
Published: (2026)
by: Li, Junxian, et al.
Published: (2026)
UniChange: Unifying Change Detection with Multimodal Large Language Model
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
UniCal: Unified Neural Sensor Calibration
by: Yang, Ze, et al.
Published: (2024)
by: Yang, Ze, et al.
Published: (2024)
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
UniRSCD: A Unified Novel Architectural Paradigm for Remote Sensing Change Detection
by: Qu, Yuan, et al.
Published: (2025)
by: Qu, Yuan, et al.
Published: (2025)
Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion
by: Chen, Zhifei, et al.
Published: (2024)
by: Chen, Zhifei, et al.
Published: (2024)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
by: Xu, Yueming, et al.
Published: (2025)
by: Xu, Yueming, et al.
Published: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Implicit Neural Representation Facilitates Unified Universal Vision Encoding
by: Gwilliam, Matthew, et al.
Published: (2026)
by: Gwilliam, Matthew, et al.
Published: (2026)
UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
by: Ni, Shuo, et al.
Published: (2025)
by: Ni, Shuo, et al.
Published: (2025)
Similar Items
-
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
by: Chen, Yanzhe, et al.
Published: (2025) -
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
by: Mao, Weijia, et al.
Published: (2025) -
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
by: Mao, Weijia, et al.
Published: (2025) -
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification
by: Qu, Xiangyan, et al.
Published: (2025) -
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
by: Zhao, Yaqi, et al.
Published: (2026)