Gespeichert in:
| Hauptverfasser: | Zeng, Yanbing, Wang, Jia, Ma, Hanghang, Wu, Junqiang, Zhu, Jie, Wei, Xiaoming, Hu, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.04706 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
von: Wang, Jia, et al.
Veröffentlicht: (2025)
von: Wang, Jia, et al.
Veröffentlicht: (2025)
LongCat-Image Technical Report
von: Meituan LongCat Team, et al.
Veröffentlicht: (2025)
von: Meituan LongCat Team, et al.
Veröffentlicht: (2025)
Rascene: High-Fidelity 3D Scene Imaging with mmWave Communication Signals
von: Song, Kunzhe, et al.
Veröffentlicht: (2026)
von: Song, Kunzhe, et al.
Veröffentlicht: (2026)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
Enhancing Image Generation Fidelity via Progressive Prompts
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
von: Wu, Xiaoping, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoping, et al.
Veröffentlicht: (2024)
Object Fidelity Diffusion for Remote Sensing Image Generation
von: Ye, Ziqi, et al.
Veröffentlicht: (2025)
von: Ye, Ziqi, et al.
Veröffentlicht: (2025)
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
von: Zhou, Chao, et al.
Veröffentlicht: (2025)
von: Zhou, Chao, et al.
Veröffentlicht: (2025)
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
von: Tan, Jiangtong, et al.
Veröffentlicht: (2025)
von: Tan, Jiangtong, et al.
Veröffentlicht: (2025)
An Analysis for Image-to-Image Translation and Style Transfer
von: Yu, Xiaoming, et al.
Veröffentlicht: (2024)
von: Yu, Xiaoming, et al.
Veröffentlicht: (2024)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
Multimodal Generation of Animatable 3D Human Models with AvatarForge
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
von: Deng, Xiang, et al.
Veröffentlicht: (2026)
von: Deng, Xiang, et al.
Veröffentlicht: (2026)
A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
von: Wu, Tianhe, et al.
Veröffentlicht: (2024)
von: Wu, Tianhe, et al.
Veröffentlicht: (2024)
On the Holistic Approach for Detecting Human Image Forgery
von: Guo, Xiao, et al.
Veröffentlicht: (2026)
von: Guo, Xiao, et al.
Veröffentlicht: (2026)
DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design
von: Hu, Xiwei, et al.
Veröffentlicht: (2025)
von: Hu, Xiwei, et al.
Veröffentlicht: (2025)
Towards Generalized Multi-Image Editing for Unified Multimodal Models
von: Xu, Pengcheng, et al.
Veröffentlicht: (2026)
von: Xu, Pengcheng, et al.
Veröffentlicht: (2026)
Denoising with a Joint-Embedding Predictive Architecture
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
Fine-gained Zero-shot Video Sampling
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
High-Resolution Image Synthesis via Next-Token Prediction
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
von: Han, Ruiyan, et al.
Veröffentlicht: (2026)
von: Han, Ruiyan, et al.
Veröffentlicht: (2026)
Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2025)
Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
von: Shao, Jie, et al.
Veröffentlicht: (2025)
von: Shao, Jie, et al.
Veröffentlicht: (2025)
Large Language Models for Multimodal Deformable Image Registration
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
PositionIC: Unified Position and Identity Consistency for Image Customization
von: Hu, Junjie, et al.
Veröffentlicht: (2025)
von: Hu, Junjie, et al.
Veröffentlicht: (2025)
DF-SLAM: Dictionary Factors Representation for High-Fidelity Neural Implicit Dense Visual SLAM System
von: Wei, Weifeng, et al.
Veröffentlicht: (2024)
von: Wei, Weifeng, et al.
Veröffentlicht: (2024)
FlightForge: Advancing UAV Research with Procedural Generation of High-Fidelity Simulation and Integrated Autonomy
von: Čapek, David, et al.
Veröffentlicht: (2025)
von: Čapek, David, et al.
Veröffentlicht: (2025)
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
von: Qin, Jie, et al.
Veröffentlicht: (2025)
von: Qin, Jie, et al.
Veröffentlicht: (2025)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
von: Liang, Qian, et al.
Veröffentlicht: (2025)
von: Liang, Qian, et al.
Veröffentlicht: (2025)
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation
von: Wang, Haonan, et al.
Veröffentlicht: (2024)
von: Wang, Haonan, et al.
Veröffentlicht: (2024)
DiffusionAgent: Navigating Expert Models for Agentic Image Generation
von: Qin, Jie, et al.
Veröffentlicht: (2024)
von: Qin, Jie, et al.
Veröffentlicht: (2024)
Multimodal Point Cloud Semantic Segmentation With Virtual Point Enhancement
von: Duan, Zaipeng, et al.
Veröffentlicht: (2025)
von: Duan, Zaipeng, et al.
Veröffentlicht: (2025)
Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector
von: Guo, Xiao, et al.
Veröffentlicht: (2025)
von: Guo, Xiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction
von: Zhu, Jie, et al.
Veröffentlicht: (2026) -
MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
von: Wang, Jia, et al.
Veröffentlicht: (2025) -
LongCat-Image Technical Report
von: Meituan LongCat Team, et al.
Veröffentlicht: (2025) -
Rascene: High-Fidelity 3D Scene Imaging with mmWave Communication Signals
von: Song, Kunzhe, et al.
Veröffentlicht: (2026) -
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)