Growing Visual Generative Capacity for Pre-Trained MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hanyu, Han, Jiaming, Yang, Ziyan, Zhao, Qi, Lin, Shanchuan, Yue, Xiangyu, Shrivastava, Abhinav, Yang, Zhenheng, Chen, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
von: Ai, Yuang, et al.
Veröffentlicht: (2026)
von: Ai, Yuang, et al.
Veröffentlicht: (2026)
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
von: Wang, Hanyu, et al.
Veröffentlicht: (2024)
von: Wang, Hanyu, et al.
Veröffentlicht: (2024)
NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis
von: Ren, Yixuan, et al.
Veröffentlicht: (2025)
von: Ren, Yixuan, et al.
Veröffentlicht: (2025)
Long Context Tuning for Video Generation
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
AnimateDiff-Lightning: Cross-Model Diffusion Distillation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2024)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2024)
Efficient Continuous Video Flow Model for Video Prediction
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
von: Zhu, Zifeng, et al.
Veröffentlicht: (2026)
von: Zhu, Zifeng, et al.
Veröffentlicht: (2026)
UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
von: Zhuang, Shaobin, et al.
Veröffentlicht: (2026)
von: Zhuang, Shaobin, et al.
Veröffentlicht: (2026)
Diffusion Model with Perceptual Loss
von: Lin, Shanchuan, et al.
Veröffentlicht: (2023)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2023)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
von: Lou, Haoran, et al.
Veröffentlicht: (2025)
von: Lou, Haoran, et al.
Veröffentlicht: (2025)
On the Generalization Capacities of MLLMs for Spatial Intelligence
von: Zhang, Gongjie, et al.
Veröffentlicht: (2026)
von: Zhang, Gongjie, et al.
Veröffentlicht: (2026)
Is Your Text-to-Image Model Robust to Caption Noise?
von: Yu, Weichen, et al.
Veröffentlicht: (2024)
von: Yu, Weichen, et al.
Veröffentlicht: (2024)
Parallelized Autoregressive Visual Generation
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
von: Li, Bozhou, et al.
Veröffentlicht: (2025)
von: Li, Bozhou, et al.
Veröffentlicht: (2025)
Common Diffusion Noise Schedules and Sample Steps are Flawed
von: Lin, Shanchuan, et al.
Veröffentlicht: (2023)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2023)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
Visual Jigsaw Post-Training Improves MLLMs
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
Diffusion Adversarial Post-Training for One-Step Video Generation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
Continuous Adversarial Flow Models
von: Lin, Shanchuan, et al.
Veröffentlicht: (2026)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2026)
Adversarial Flow Models
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
von: Wang, Jianyi, et al.
Veröffentlicht: (2025)
von: Wang, Jianyi, et al.
Veröffentlicht: (2025)
The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
von: Mao, Weijia, et al.
Veröffentlicht: (2025)
von: Mao, Weijia, et al.
Veröffentlicht: (2025)
FreeRet: MLLMs as Training-Free Retrievers
von: Zhu, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2025)
ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
von: Jiang, Yilei, et al.
Veröffentlicht: (2025)
von: Jiang, Yilei, et al.
Veröffentlicht: (2025)
SDXL-Lightning: Progressive Adversarial Diffusion Distillation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2024)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2024)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
von: Guan, Tongkun, et al.
Veröffentlicht: (2026)
von: Guan, Tongkun, et al.
Veröffentlicht: (2026)
How to Design and Train Your Implicit Neural Representation for Video Compression
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2025)
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2025)
V-VIPE: Variational View Invariant Pose Embedding
von: Levy, Mara, et al.
Veröffentlicht: (2024)
von: Levy, Mara, et al.
Veröffentlicht: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs
von: Lou, Haoran, et al.
Veröffentlicht: (2026)
von: Lou, Haoran, et al.
Veröffentlicht: (2026)
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025)
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
von: Li, Qi, et al.
Veröffentlicht: (2026)
von: Li, Qi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025) -
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
von: Ai, Yuang, et al.
Veröffentlicht: (2026) -
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
von: Wang, Hanyu, et al.
Veröffentlicht: (2024) -
NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis
von: Ren, Yixuan, et al.
Veröffentlicht: (2025) -
Long Context Tuning for Video Generation
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)