Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Xing, Jinbo, Jiang, Zeyinzi, Tuo, Yuxiang, Mao, Chaojie, Gai, Xiaotang, Chen, Xi, Zhang, Jingfeng, Pan, Yulin, Han, Zhen, Xiao, Jie, Yan, Keyu, Xie, Chenwei, Zhong, Chongyang, Zhu, Kai, Shen, Tong, Huang, Lianghua, Liu, Yu, Yang, Yujiu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StyleBooth: Image Style Editing with Multimodal Instruction
di: Han, Zhen, et al.
Pubblicazione: (2024)
di: Han, Zhen, et al.
Pubblicazione: (2024)
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer
di: Han, Zhen, et al.
Pubblicazione: (2024)
di: Han, Zhen, et al.
Pubblicazione: (2024)
VACE: All-in-One Video Creation and Editing
di: Jiang, Zeyinzi, et al.
Pubblicazione: (2025)
di: Jiang, Zeyinzi, et al.
Pubblicazione: (2025)
Locate, Assign, Refine: Taming Customized Promptable Image Inpainting
di: Pan, Yulin, et al.
Pubblicazione: (2024)
di: Pan, Yulin, et al.
Pubblicazione: (2024)
ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing
di: Pan, Yulin, et al.
Pubblicazione: (2025)
di: Pan, Yulin, et al.
Pubblicazione: (2025)
ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling
di: Mao, Chaojie, et al.
Pubblicazione: (2025)
di: Mao, Chaojie, et al.
Pubblicazione: (2025)
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
di: Mao, Chaojie, et al.
Pubblicazione: (2026)
di: Mao, Chaojie, et al.
Pubblicazione: (2026)
Wan: Open and Advanced Large-Scale Video Generative Models
di: Wan, Team, et al.
Pubblicazione: (2025)
di: Wan, Team, et al.
Pubblicazione: (2025)
AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation
di: He, Junjie, et al.
Pubblicazione: (2025)
di: He, Junjie, et al.
Pubblicazione: (2025)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
di: Shi, Yudi, et al.
Pubblicazione: (2026)
di: Shi, Yudi, et al.
Pubblicazione: (2026)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
di: Tian, Changyao, et al.
Pubblicazione: (2024)
di: Tian, Changyao, et al.
Pubblicazione: (2024)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
di: Chen, Junyi, et al.
Pubblicazione: (2026)
di: Chen, Junyi, et al.
Pubblicazione: (2026)
Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
di: Zhao, Shihao, et al.
Pubblicazione: (2025)
di: Zhao, Shihao, et al.
Pubblicazione: (2025)
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
di: Chu, Ruihang, et al.
Pubblicazione: (2025)
di: Chu, Ruihang, et al.
Pubblicazione: (2025)
Semi‐Analytical Solution for Consolidation of Unsaturated Soil Composite Foundation Improved by Local Permeable Pile With Finite Hankel Transform
di: Hongping Meng, et al.
Pubblicazione: (2025)
di: Hongping Meng, et al.
Pubblicazione: (2025)
Analytical solution for consolidation of unsaturated composite foundation improved by permeable and impermeable columns considering depth‐dependent initial stress
di: Aifang Qin, et al.
Pubblicazione: (2024)
di: Aifang Qin, et al.
Pubblicazione: (2024)
Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
di: Wang, Junjie, et al.
Pubblicazione: (2025)
di: Wang, Junjie, et al.
Pubblicazione: (2025)
Graph Triple Attention Network: A Decoupled Perspective
di: Wang, Xiaotang, et al.
Pubblicazione: (2024)
di: Wang, Xiaotang, et al.
Pubblicazione: (2024)
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
On a Fiber Conjecture of Wan
di: Schmidt, Matthew
Pubblicazione: (2023)
di: Schmidt, Matthew
Pubblicazione: (2023)
TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild
di: Li, Huayang, et al.
Pubblicazione: (2023)
di: Li, Huayang, et al.
Pubblicazione: (2023)
La demanda de los Wan
di: Liliana Arsovska
Pubblicazione: (2003)
di: Liliana Arsovska
Pubblicazione: (2003)
Coastal current in Sagami-Wan (1).
di: Odamaki, Minoru, et al.
Pubblicazione: (1987)
di: Odamaki, Minoru, et al.
Pubblicazione: (1987)
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
di: Deng, Haoyou, et al.
Pubblicazione: (2026)
di: Deng, Haoyou, et al.
Pubblicazione: (2026)
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
di: Zhang, Yizhen, et al.
Pubblicazione: (2025)
di: Zhang, Yizhen, et al.
Pubblicazione: (2025)
MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
di: Wei, Yuxiang, et al.
Pubblicazione: (2024)
di: Wei, Yuxiang, et al.
Pubblicazione: (2024)
Research on Photovoltaic Grid‐Connected Inverter Based on Interleaved Parallel Decoupling
di: Yu Wang, et al.
Pubblicazione: (2025)
di: Yu Wang, et al.
Pubblicazione: (2025)
3D-RAD: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks
di: Gai, Xiaotang, et al.
Pubblicazione: (2025)
di: Gai, Xiaotang, et al.
Pubblicazione: (2025)
MedThink: Explaining Medical Visual Question Answering via Multimodal Decision-Making Rationale
di: Gai, Xiaotang, et al.
Pubblicazione: (2024)
di: Gai, Xiaotang, et al.
Pubblicazione: (2024)
Beam Weaver
di: Teles, Pedro, et al.
Pubblicazione: (2026)
di: Teles, Pedro, et al.
Pubblicazione: (2026)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
di: Wei, Cong, et al.
Pubblicazione: (2024)
di: Wei, Cong, et al.
Pubblicazione: (2024)
Wan-S2V: Audio-Driven Cinematic Video Generation
di: Gao, Xin, et al.
Pubblicazione: (2025)
di: Gao, Xin, et al.
Pubblicazione: (2025)
Wan-Animate: Unified Character Animation and Replacement with Holistic Replication
di: Cheng, Gang, et al.
Pubblicazione: (2025)
di: Cheng, Gang, et al.
Pubblicazione: (2025)
AnyText2: Visual Text Generation and Editing With Customizable Attributes
di: Tuo, Yuxiang, et al.
Pubblicazione: (2024)
di: Tuo, Yuxiang, et al.
Pubblicazione: (2024)
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
di: Luo, Xiangyang, et al.
Pubblicazione: (2025)
di: Luo, Xiangyang, et al.
Pubblicazione: (2025)
Moral licensing effect of work engagement: The role of psychological entitlement and relationship conflict with supervisors
di: Lianghua Zhang, et al.
Pubblicazione: (2024)
di: Lianghua Zhang, et al.
Pubblicazione: (2024)
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
di: Wang, Junjie, et al.
Pubblicazione: (2024)
di: Wang, Junjie, et al.
Pubblicazione: (2024)
FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
di: Fan, Ziyang, et al.
Pubblicazione: (2026)
di: Fan, Ziyang, et al.
Pubblicazione: (2026)
Exceptional extensions of local fields and the Carlitz--Wan conjecture
di: Ding, Zhiguo, et al.
Pubblicazione: (2025)
di: Ding, Zhiguo, et al.
Pubblicazione: (2025)
WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset
di: Qiu, Jiantao, et al.
Pubblicazione: (2024)
di: Qiu, Jiantao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
StyleBooth: Image Style Editing with Multimodal Instruction
di: Han, Zhen, et al.
Pubblicazione: (2024) -
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer
di: Han, Zhen, et al.
Pubblicazione: (2024) -
VACE: All-in-One Video Creation and Editing
di: Jiang, Zeyinzi, et al.
Pubblicazione: (2025) -
Locate, Assign, Refine: Taming Customized Promptable Image Inpainting
di: Pan, Yulin, et al.
Pubblicazione: (2024) -
ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing
di: Pan, Yulin, et al.
Pubblicazione: (2025)