Exploring MLLM-Diffusion Information Transfer with MetaCanvas
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Han, Pan, Xichen, Huang, Ziqi, Hou, Ji, Wang, Jialiang, Chen, Weifeng, He, Zecheng, Juefei-Xu, Felix, Sun, Junzhe, Fan, Zhipeng, Thabet, Ali, Bansal, Mohit, Wang, Chu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transfer between Modalities with MetaQueries
by: Pan, Xichen, et al.
Published: (2025)
by: Pan, Xichen, et al.
Published: (2025)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
by: Cai, Yuanhao, et al.
Published: (2025)
by: Cai, Yuanhao, et al.
Published: (2025)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026)
by: Chen, Leon Liangyu, et al.
Published: (2026)
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
by: Lin, Han, et al.
Published: (2026)
by: Lin, Han, et al.
Published: (2026)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
Pixel-Space Post-Training of Latent Diffusion Models
by: Zhang, Christina, et al.
Published: (2024)
by: Zhang, Christina, et al.
Published: (2024)
IntuiTF: MLLM-Guided Transfer Function Optimization for Direct Volume Rendering
by: Wang, Yiyao, et al.
Published: (2025)
by: Wang, Yiyao, et al.
Published: (2025)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
by: Gu, Zeqi, et al.
Published: (2025)
by: Gu, Zeqi, et al.
Published: (2025)
Populate-A-Scene: Affordance-Aware Human Video Generation
by: Shan, Mengyi, et al.
Published: (2025)
by: Shan, Mengyi, et al.
Published: (2025)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025)
by: Tang, Bingda, et al.
Published: (2025)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters
by: Hansen-Estruch, Philippe, et al.
Published: (2026)
by: Hansen-Estruch, Philippe, et al.
Published: (2026)
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
by: Li, Chenyu, et al.
Published: (2025)
by: Li, Chenyu, et al.
Published: (2025)
Surveying the MLLM Landscape: A Meta-Review of Current Surveys
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
StreamDiT: Real-Time Streaming Text-to-Video Generation
by: Kodaira, Akio, et al.
Published: (2025)
by: Kodaira, Akio, et al.
Published: (2025)
MLLM-as-a-Judge for Image Safety without Human Labeling
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
by: Song, Kunpeng, et al.
Published: (2024)
by: Song, Kunpeng, et al.
Published: (2024)
Efficient Distributed MLLM Training with Cornstarch
by: Jang, Insu, et al.
Published: (2025)
by: Jang, Insu, et al.
Published: (2025)
Elysium: Exploring Object-level Perception in Videos via MLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
by: Liang, Feng, et al.
Published: (2025)
by: Liang, Feng, et al.
Published: (2025)
MoCha: Towards Movie-Grade Talking Character Synthesis
by: Wei, Cong, et al.
Published: (2025)
by: Wei, Cong, et al.
Published: (2025)
PER-DPP Sampling Framework and Its Application in Path Planning
by: Wang, Junzhe
Published: (2025)
by: Wang, Junzhe
Published: (2025)
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
by: Wang, Hongjie, et al.
Published: (2024)
by: Wang, Hongjie, et al.
Published: (2024)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
by: Huang, Runhui, et al.
Published: (2025)
by: Huang, Runhui, et al.
Published: (2025)
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)
by: Zhao, Chen, et al.
Published: (2020)
The Listening Canvas
by: Perc, Luciana
Published: (2025)
by: Perc, Luciana
Published: (2025)
CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
by: Li, Zian, et al.
Published: (2025)
by: Li, Zian, et al.
Published: (2025)
ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering
by: Wang, Xinyue, et al.
Published: (2025)
by: Wang, Xinyue, et al.
Published: (2025)
Transfer Learning for Meta-analysis Under Covariate Shift
by: Wang, Zilong, et al.
Published: (2026)
by: Wang, Zilong, et al.
Published: (2026)
SARD: Segmentation-Aware Anomaly Synthesis via Region-Constrained Diffusion with Discriminative Mask Guidance
by: Wang, Yanshu, et al.
Published: (2025)
by: Wang, Yanshu, et al.
Published: (2025)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Myrtus Communis Extracts as Biocontrol Agents against Tribolium Castaneum
by: Thabet Mudheher Khalaf
Published: (2025)
by: Thabet Mudheher Khalaf
Published: (2025)
Compositions of Resolvents: Fixed Points Sets and Set of Cycles
by: Alwadani, Salihah Thabet
Published: (2024)
by: Alwadani, Salihah Thabet
Published: (2024)
More on Arago'n Artacho -- Campoy's Algorithm Operators
by: Alwadani, Salihah Thabet
Published: (2024)
by: Alwadani, Salihah Thabet
Published: (2024)
Additional Studies on Displacement Mapping with Restrictions
by: Alwadani, Salihah Thabet
Published: (2024)
by: Alwadani, Salihah Thabet
Published: (2024)
The effectiveness of immersive virtual reality applications (human anatomy) on self‐directed learning competencies among undergraduate nursing students: A cross‐sectional study
by: Samar Thabet Jallad
Published: (2024)
by: Samar Thabet Jallad
Published: (2024)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
by: Lin, Han, et al.
Published: (2025)
by: Lin, Han, et al.
Published: (2025)
Primitive-Planner: An Ultra Lightweight Quadrotor Planner with Time-optimal Primitives
by: Hou, Jialiang, et al.
Published: (2025)
by: Hou, Jialiang, et al.
Published: (2025)
Exploring the Influence of Metabolic Changes in Fibrotic Lung Diseases
by: Swati Kumari, et al.
Published: (2025)
by: Swati Kumari, et al.
Published: (2025)
Similar Items
-
Transfer between Modalities with MetaQueries
by: Pan, Xichen, et al.
Published: (2025) -
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
by: Cai, Yuanhao, et al.
Published: (2025) -
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026) -
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
by: Lin, Han, et al.
Published: (2026) -
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)