EVCtrl: Efficient Control Adapter for Visual Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Zixiang, Ma, Yue, Zhang, Yinhan, Mo, Shanhui, Liu, Dongrui, Zhang, Linfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation
por: Ma, Yue, et al.
Publicado: (2026)
por: Ma, Yue, et al.
Publicado: (2026)
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
por: Feng, Kunyu, et al.
Publicado: (2025)
por: Feng, Kunyu, et al.
Publicado: (2025)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
por: Zhang, Qizhe, et al.
Publicado: (2023)
por: Zhang, Qizhe, et al.
Publicado: (2023)
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
por: Cai, Peiliang, et al.
Publicado: (2026)
por: Cai, Peiliang, et al.
Publicado: (2026)
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
por: Jiang, Yue, et al.
Publicado: (2026)
por: Jiang, Yue, et al.
Publicado: (2026)
Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
por: Zhang, Yurong, et al.
Publicado: (2024)
por: Zhang, Yurong, et al.
Publicado: (2024)
Follow-Your-Color: Multi-Instance Sketch Colorization
por: Zhang, Yinhan, et al.
Publicado: (2025)
por: Zhang, Yinhan, et al.
Publicado: (2025)
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
por: Liu, Penghui, et al.
Publicado: (2025)
por: Liu, Penghui, et al.
Publicado: (2025)
Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning
por: Ma, Qianli, et al.
Publicado: (2024)
por: Ma, Qianli, et al.
Publicado: (2024)
SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping
por: Li, Jiajun, et al.
Publicado: (2025)
por: Li, Jiajun, et al.
Publicado: (2025)
Control and Realism: Best of Both Worlds in Layout-to-Image without Training
por: Li, Bonan, et al.
Publicado: (2025)
por: Li, Bonan, et al.
Publicado: (2025)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
por: Zhang, Yue, et al.
Publicado: (2024)
por: Zhang, Yue, et al.
Publicado: (2024)
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
por: Liu, Jiaming, et al.
Publicado: (2023)
por: Liu, Jiaming, et al.
Publicado: (2023)
Test-time Distribution Learning Adapter for Cross-modal Visual Reasoning
por: Zhang, Yi, et al.
Publicado: (2024)
por: Zhang, Yi, et al.
Publicado: (2024)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
por: Guo, Xun, et al.
Publicado: (2023)
por: Guo, Xun, et al.
Publicado: (2023)
EEdit: Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing
por: Yan, Zexuan, et al.
Publicado: (2025)
por: Yan, Zexuan, et al.
Publicado: (2025)
Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
por: Hirano, Shinnosuke, et al.
Publicado: (2025)
por: Hirano, Shinnosuke, et al.
Publicado: (2025)
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
por: Ma, Yue, et al.
Publicado: (2025)
por: Ma, Yue, et al.
Publicado: (2025)
Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
por: Long, Zeqian, et al.
Publicado: (2025)
por: Long, Zeqian, et al.
Publicado: (2025)
Memory Efficient Transformer Adapter for Dense Predictions
por: Zhang, Dong, et al.
Publicado: (2025)
por: Zhang, Dong, et al.
Publicado: (2025)
Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation
por: Li, Zizhuo, et al.
Publicado: (2026)
por: Li, Zizhuo, et al.
Publicado: (2026)
Pear: Pruning and Sharing Adapters in Visual Parameter-Efficient Fine-Tuning
por: Zhong, Yibo, et al.
Publicado: (2024)
por: Zhong, Yibo, et al.
Publicado: (2024)
FontAdapter: Instant Font Adaptation in Visual Text Generation
por: Koo, Myungkyu, et al.
Publicado: (2025)
por: Koo, Myungkyu, et al.
Publicado: (2025)
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
por: Gong, Yan, et al.
Publicado: (2025)
por: Gong, Yan, et al.
Publicado: (2025)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
por: Gao, Peng, et al.
Publicado: (2021)
por: Gao, Peng, et al.
Publicado: (2021)
Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control
por: Han, Yue, et al.
Publicado: (2024)
por: Han, Yue, et al.
Publicado: (2024)
Raw-JPEG Adapter: Efficient Raw Image Compression with JPEG
por: Afifi, Mahmoud, et al.
Publicado: (2025)
por: Afifi, Mahmoud, et al.
Publicado: (2025)
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
por: Liu, Ting, et al.
Publicado: (2024)
por: Liu, Ting, et al.
Publicado: (2024)
InstanceAnimator: Multi-Instance Sketch Video Colorization
por: Zhang, Yinhan, et al.
Publicado: (2026)
por: Zhang, Yinhan, et al.
Publicado: (2026)
Adapter-X: A Novel General Parameter-Efficient Fine-Tuning Framework for Vision
por: Li, Minglei, et al.
Publicado: (2024)
por: Li, Minglei, et al.
Publicado: (2024)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
por: Liu, Shanhui, et al.
Publicado: (2025)
por: Liu, Shanhui, et al.
Publicado: (2025)
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
por: Zhong, Weizhi, et al.
Publicado: (2025)
por: Zhong, Weizhi, et al.
Publicado: (2025)
Mixture of Physical Priors Adapter for Parameter-Efficient Fine-Tuning
por: Wang, Zhaozhi, et al.
Publicado: (2024)
por: Wang, Zhaozhi, et al.
Publicado: (2024)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
por: Cao, Meng, et al.
Publicado: (2024)
por: Cao, Meng, et al.
Publicado: (2024)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
por: Gong, Meiqi, et al.
Publicado: (2025)
por: Gong, Meiqi, et al.
Publicado: (2025)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
por: Chen, Junan, et al.
Publicado: (2025)
por: Chen, Junan, et al.
Publicado: (2025)
Hi-GaTA: Hierarchical Gated Temporal Aggregation Adapter for Surgical Video Report Generation
por: Sun, Kedi, et al.
Publicado: (2026)
por: Sun, Kedi, et al.
Publicado: (2026)
EdgeFM: Efficient Edge Inference for Vision-Language Models
por: Deng, Mengling, et al.
Publicado: (2026)
por: Deng, Mengling, et al.
Publicado: (2026)
Generalization Boosted Adapter for Open-Vocabulary Segmentation
por: Xu, Wenhao, et al.
Publicado: (2024)
por: Xu, Wenhao, et al.
Publicado: (2024)
DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding
por: Liu, Ting, et al.
Publicado: (2024)
por: Liu, Ting, et al.
Publicado: (2024)
Ejemplares similares
-
EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation
por: Ma, Yue, et al.
Publicado: (2026) -
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
por: Feng, Kunyu, et al.
Publicado: (2025) -
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
por: Zhang, Qizhe, et al.
Publicado: (2023) -
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
por: Cai, Peiliang, et al.
Publicado: (2026) -
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
por: Jiang, Yue, et al.
Publicado: (2026)