DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jianzong, Tang, Chao, Wang, Jingbo, Zeng, Yanhong, Li, Xiangtai, Tong, Yunhai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DreamRelation: Bridging Customization and Relation Generation
von: Shi, Qingyu, et al.
Veröffentlicht: (2024)
von: Shi, Qingyu, et al.
Veröffentlicht: (2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
von: Chen, Yicheng, et al.
Veröffentlicht: (2024)
von: Chen, Yicheng, et al.
Veröffentlicht: (2024)
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
One-Step Distillation of Discrete Diffusion Image Generators via Fixed-Point Iteration
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
Explore In-Context Segmentation via Latent Diffusion Models
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
Manga Generation via Layout-controllable Diffusion
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
AMM-Diff: Adaptive Multi-Modality Diffusion Network for Missing Modality Imputation
von: Kebaili, Aghiles, et al.
Veröffentlicht: (2025)
von: Kebaili, Aghiles, et al.
Veröffentlicht: (2025)
Diffusion Models For Multi-Modal Generative Modeling
von: Chen, Changyou, et al.
Veröffentlicht: (2024)
von: Chen, Changyou, et al.
Veröffentlicht: (2024)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
von: Xu, Shilin, et al.
Veröffentlicht: (2023)
von: Xu, Shilin, et al.
Veröffentlicht: (2023)
Generative Classifier for Domain Generalization
von: Long, Shaocong, et al.
Veröffentlicht: (2025)
von: Long, Shaocong, et al.
Veröffentlicht: (2025)
Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models
von: Xing, Zhening, et al.
Veröffentlicht: (2024)
von: Xing, Zhening, et al.
Veröffentlicht: (2024)
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
von: Wang, Muyao, et al.
Veröffentlicht: (2026)
von: Wang, Muyao, et al.
Veröffentlicht: (2026)
MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
User-Friendly Customized Generation with Multi-Modal Prompts
von: Zhong, Linhao, et al.
Veröffentlicht: (2024)
von: Zhong, Linhao, et al.
Veröffentlicht: (2024)
VG4D: Vision-Language Model Goes 4D Video Recognition
von: Deng, Zhichao, et al.
Veröffentlicht: (2024)
von: Deng, Zhichao, et al.
Veröffentlicht: (2024)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
FOD-Diff: 3D Multi-Channel Patch Diffusion Model for Fiber Orientation Distribution
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Towards Open Vocabulary Learning: A Survey
von: Wu, Jianzong, et al.
Veröffentlicht: (2023)
von: Wu, Jianzong, et al.
Veröffentlicht: (2023)
Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
von: Zheng, Shuhong, et al.
Veröffentlicht: (2024)
von: Zheng, Shuhong, et al.
Veröffentlicht: (2024)
CyberV: Cybernetics for Test-time Scaling in Video Understanding
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
RecTok: Reconstruction Distillation along Rectified Flow
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
CRS-Diff: Controllable Remote Sensing Image Generation with Diffusion Model
von: Tang, Datao, et al.
Veröffentlicht: (2024)
von: Tang, Datao, et al.
Veröffentlicht: (2024)
Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain
von: Chao, Lianying, et al.
Veröffentlicht: (2026)
von: Chao, Lianying, et al.
Veröffentlicht: (2026)
CamFreeDiff: Camera-free Image to Panorama Generation with Diffusion Model
von: Yuan, Xiaoding, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaoding, et al.
Veröffentlicht: (2024)
DiffTF++: 3D-aware Diffusion Transformer for Large-Vocabulary 3D Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2024)
von: Cao, Ziang, et al.
Veröffentlicht: (2024)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2024)
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2024)
DiffX: Guide Your Layout to Cross-Modal Generative Modeling
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
von: Wang, Zitong, et al.
Veröffentlicht: (2025)
von: Wang, Zitong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DreamRelation: Bridging Customization and Relation Generation
von: Shi, Qingyu, et al.
Veröffentlicht: (2024) -
MotionBooth: Motion-Aware Customized Text-to-Video Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2024) -
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
von: Shi, Qingyu, et al.
Veröffentlicht: (2025) -
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
von: Chen, Yicheng, et al.
Veröffentlicht: (2024) -
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)