Learnable Sparsity for Vision Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yang, Jin, Er, Liang, Wenzhong, Dong, Yanfei, Khakzar, Ashkan, Torr, Philip, Stegmaier, Johannes, Kawaguchi, Kenji |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minimalist Concept Erasure in Generative Models
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
by: Jin, Er, et al.
Published: (2025)
by: Jin, Er, et al.
Published: (2025)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
by: Rezaei, Razieh, et al.
Published: (2024)
by: Rezaei, Razieh, et al.
Published: (2024)
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
by: Deb, Oishi, et al.
Published: (2025)
by: Deb, Oishi, et al.
Published: (2025)
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
by: Hemmat, Arshia, et al.
Published: (2024)
by: Hemmat, Arshia, et al.
Published: (2024)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Segment Anything for Cell Tracking
by: Chen, Zhu, et al.
Published: (2025)
by: Chen, Zhu, et al.
Published: (2025)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
by: Naghashyar, Lachin, et al.
Published: (2026)
by: Naghashyar, Lachin, et al.
Published: (2026)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction
by: Jin, Er, et al.
Published: (2025)
by: Jin, Er, et al.
Published: (2025)
SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
by: Gülhan, Fabian, et al.
Published: (2025)
by: Gülhan, Fabian, et al.
Published: (2025)
Unsupervised Learning for Feature Extraction and Temporal Alignment of 3D+t Point Clouds of Zebrafish Embryos
by: Chen, Zhu, et al.
Published: (2025)
by: Chen, Zhu, et al.
Published: (2025)
A Survey on Transferability of Adversarial Examples across Deep Neural Networks
by: Gu, Jindong, et al.
Published: (2023)
by: Gu, Jindong, et al.
Published: (2023)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
by: Luo, Haochen, et al.
Published: (2024)
by: Luo, Haochen, et al.
Published: (2024)
Annotated Biomedical Video Generation using Denoising Diffusion Probabilistic Models and Flow Fields
by: Yilmaz, Rüveyda, et al.
Published: (2024)
by: Yilmaz, Rüveyda, et al.
Published: (2024)
Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
by: Ma, Xiangkai, et al.
Published: (2025)
by: Ma, Xiangkai, et al.
Published: (2025)
Cascaded Diffusion Models for 2D and 3D Microscopy Image Synthesis to Enhance Cell Segmentation
by: Yilmaz, Rüveyda, et al.
Published: (2024)
by: Yilmaz, Rüveyda, et al.
Published: (2024)
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
by: Xu, Hengbo, et al.
Published: (2026)
by: Xu, Hengbo, et al.
Published: (2026)
Learning Camera Movement Control from Real-World Drone Videos
by: Hou, Yunzhong, et al.
Published: (2024)
by: Hou, Yunzhong, et al.
Published: (2024)
Which Model Generated This Image? A Model-Agnostic Approach for Origin Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
by: Xing, Songlong, et al.
Published: (2026)
by: Xing, Songlong, et al.
Published: (2026)
AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
by: Ganj, Ashkan, et al.
Published: (2025)
by: Ganj, Ashkan, et al.
Published: (2025)
A Pragmatic Note on Evaluating Generative Models with Fréchet Inception Distance for Retinal Image Synthesis
by: Wu, Yuli, et al.
Published: (2025)
by: Wu, Yuli, et al.
Published: (2025)
Replacement Learning: Training Vision Tasks with Fewer Learnable Parameters
by: Zhang, Yuming, et al.
Published: (2024)
by: Zhang, Yuming, et al.
Published: (2024)
Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Vision-Language Models Do Not Understand Negation
by: Alhamoud, Kumail, et al.
Published: (2025)
by: Alhamoud, Kumail, et al.
Published: (2025)
Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models
by: Han, Junlin, et al.
Published: (2024)
by: Han, Junlin, et al.
Published: (2024)
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
by: Li, Runjia, et al.
Published: (2025)
by: Li, Runjia, et al.
Published: (2025)
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
by: Yuan, Jianhao, et al.
Published: (2022)
by: Yuan, Jianhao, et al.
Published: (2022)
Semantic Score Distillation Sampling for Compositional Text-to-3D Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
CellStyle: Improved Zero-Shot Cell Segmentation via Style Transfer
by: Yilmaz, Rüveyda, et al.
Published: (2025)
by: Yilmaz, Rüveyda, et al.
Published: (2025)
Optimizing Retinal Prosthetic Stimuli with Conditional Invertible Neural Networks
by: Wu, Yuli, et al.
Published: (2024)
by: Wu, Yuli, et al.
Published: (2024)
PET-Adapter: Test-Time Domain Adaptation for Full and Limited-Angle PET Image Reconstruction
by: Yilmaz, Rüveyda, et al.
Published: (2026)
by: Yilmaz, Rüveyda, et al.
Published: (2026)
Towards Robust and Generalizable Gerchberg Saxton based Physics Inspired Neural Networks for Computer Generated Holography: A Sensitivity Analysis Framework
by: Amrutkar, Ankit, et al.
Published: (2025)
by: Amrutkar, Ankit, et al.
Published: (2025)
Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image Generation
by: Luo, Yihong, et al.
Published: (2025)
by: Luo, Yihong, et al.
Published: (2025)
Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation
by: Wang, Zihao, et al.
Published: (2026)
by: Wang, Zihao, et al.
Published: (2026)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
Similar Items
-
Minimalist Concept Erasure in Generative Models
by: Zhang, Yang, et al.
Published: (2025) -
Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
by: Jin, Er, et al.
Published: (2025) -
Learning Visual Prompts for Guiding the Attention of Vision Transformers
by: Rezaei, Razieh, et al.
Published: (2024) -
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
by: Deb, Oishi, et al.
Published: (2025) -
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
by: Hemmat, Arshia, et al.
Published: (2024)