EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Shiyuan, Li, Ruihuang, Tao, Jiale, Shao, Shuai, Lu, Qinglin, Liao, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
von: Wang, Yukun, et al.
Veröffentlicht: (2026)
von: Wang, Yukun, et al.
Veröffentlicht: (2026)
VisionCreator: A Native Visual-Generation Agentic Model with Understanding, Thinking, Planning and Creation
von: Lai, Jinxiang, et al.
Veröffentlicht: (2026)
von: Lai, Jinxiang, et al.
Veröffentlicht: (2026)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
von: Yang, Shiyuan, et al.
Veröffentlicht: (2024)
von: Yang, Shiyuan, et al.
Veröffentlicht: (2024)
Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model
von: Gu, Zheng, et al.
Veröffentlicht: (2024)
von: Gu, Zheng, et al.
Veröffentlicht: (2024)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
von: Hu, Teng, et al.
Veröffentlicht: (2025)
von: Hu, Teng, et al.
Veröffentlicht: (2025)
VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model
von: Lai, Jinxiang, et al.
Veröffentlicht: (2026)
von: Lai, Jinxiang, et al.
Veröffentlicht: (2026)
InstantCharacter: Personalize Any Characters with a Scalable Diffusion Transformer Framework
von: Tao, Jiale, et al.
Veröffentlicht: (2025)
von: Tao, Jiale, et al.
Veröffentlicht: (2025)
FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded Sampling
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2024)
MuseumMaker: Continual Style Customization without Catastrophic Forgetting
von: Liu, Chenxi, et al.
Veröffentlicht: (2024)
von: Liu, Chenxi, et al.
Veröffentlicht: (2024)
Customized Visual Storytelling with Unified Multimodal LLMs
von: Li, Wei-Hua, et al.
Veröffentlicht: (2026)
von: Li, Wei-Hua, et al.
Veröffentlicht: (2026)
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
von: Yang, Jinrui, et al.
Veröffentlicht: (2026)
von: Yang, Jinrui, et al.
Veröffentlicht: (2026)
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation
von: Zhou, Zhengguang, et al.
Veröffentlicht: (2024)
von: Zhou, Zhengguang, et al.
Veröffentlicht: (2024)
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
von: Mao, Fangyuan, et al.
Veröffentlicht: (2025)
von: Mao, Fangyuan, et al.
Veröffentlicht: (2025)
DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
von: Zhang, Peiying, et al.
Veröffentlicht: (2025)
von: Zhang, Peiying, et al.
Veröffentlicht: (2025)
PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting
von: Wang, Linqing, et al.
Veröffentlicht: (2025)
von: Wang, Linqing, et al.
Veröffentlicht: (2025)
Style Customization of Text-to-Vector Generation with Image Diffusion Priors
von: Zhang, Peiying, et al.
Veröffentlicht: (2025)
von: Zhang, Peiying, et al.
Veröffentlicht: (2025)
DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation
von: He, Jing, et al.
Veröffentlicht: (2024)
von: He, Jing, et al.
Veröffentlicht: (2024)
Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models
von: Li, Ruibin, et al.
Veröffentlicht: (2024)
von: Li, Ruibin, et al.
Veröffentlicht: (2024)
USV: Unified Sparsification for Accelerating Video Diffusion Models
von: Wu, Xinjian, et al.
Veröffentlicht: (2025)
von: Wu, Xinjian, et al.
Veröffentlicht: (2025)
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
von: Li, Ruihuang, et al.
Veröffentlicht: (2025)
von: Li, Ruihuang, et al.
Veröffentlicht: (2025)
IC-Custom: Diverse Image Customization via In-Context Learning
von: Li, Yaowei, et al.
Veröffentlicht: (2025)
von: Li, Yaowei, et al.
Veröffentlicht: (2025)
ScatterFormer: Efficient Voxel Transformer with Scattered Linear Attention
von: He, Chenhang, et al.
Veröffentlicht: (2024)
von: He, Chenhang, et al.
Veröffentlicht: (2024)
Prohibited Items Segmentation via Occlusion-aware Bilayer Modeling
von: Ren, Yunhan, et al.
Veröffentlicht: (2025)
von: Ren, Yunhan, et al.
Veröffentlicht: (2025)
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
von: Tan, Shuai, et al.
Veröffentlicht: (2025)
von: Tan, Shuai, et al.
Veröffentlicht: (2025)
TMP: Temporal Motion Propagation for Online Video Super-Resolution
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2023)
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2023)
MTV-Inpaint: Multi-Task Long Video Inpainting
von: Yang, Shiyuan, et al.
Veröffentlicht: (2025)
von: Yang, Shiyuan, et al.
Veröffentlicht: (2025)
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
ScalingGaussian: Enhancing 3D Content Creation with Generative Gaussian Splatting
von: Chen, Shen, et al.
Veröffentlicht: (2024)
von: Chen, Shen, et al.
Veröffentlicht: (2024)
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
von: Tan, Chuangchuang, et al.
Veröffentlicht: (2025)
von: Tan, Chuangchuang, et al.
Veröffentlicht: (2025)
Generative Image Layer Decomposition with Visual Effects
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
von: Yang, Jinrui, et al.
Veröffentlicht: (2024)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion Models
von: Yang, Yang, et al.
Veröffentlicht: (2024)
von: Yang, Yang, et al.
Veröffentlicht: (2024)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing
von: Zeng, Ying, et al.
Veröffentlicht: (2026)
von: Zeng, Ying, et al.
Veröffentlicht: (2026)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
von: Wang, Yukun, et al.
Veröffentlicht: (2026) -
VisionCreator: A Native Visual-Generation Agentic Model with Understanding, Thinking, Planning and Creation
von: Lai, Jinxiang, et al.
Veröffentlicht: (2026) -
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026) -
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
von: Yang, Shiyuan, et al.
Veröffentlicht: (2024) -
Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model
von: Gu, Zheng, et al.
Veröffentlicht: (2024)