Learnable Sparsity for Vision Generative Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yang, Jin, Er, Liang, Wenzhong, Dong, Yanfei, Khakzar, Ashkan, Torr, Philip, Stegmaier, Johannes, Kawaguchi, Kenji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Minimalist Concept Erasure in Generative Models
por: Zhang, Yang, et al.
Publicado: (2025)
por: Zhang, Yang, et al.
Publicado: (2025)
Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
por: Jin, Er, et al.
Publicado: (2025)
por: Jin, Er, et al.
Publicado: (2025)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
por: Rezaei, Razieh, et al.
Publicado: (2024)
por: Rezaei, Razieh, et al.
Publicado: (2024)
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
por: Deb, Oishi, et al.
Publicado: (2025)
por: Deb, Oishi, et al.
Publicado: (2025)
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
por: Hemmat, Arshia, et al.
Publicado: (2024)
por: Hemmat, Arshia, et al.
Publicado: (2024)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
por: Venhoff, Constantin, et al.
Publicado: (2025)
por: Venhoff, Constantin, et al.
Publicado: (2025)
Segment Anything for Cell Tracking
por: Chen, Zhu, et al.
Publicado: (2025)
por: Chen, Zhu, et al.
Publicado: (2025)
Latent Guard: a Safety Framework for Text-to-image Generation
por: Liu, Runtao, et al.
Publicado: (2024)
por: Liu, Runtao, et al.
Publicado: (2024)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
por: Naghashyar, Lachin, et al.
Publicado: (2026)
por: Naghashyar, Lachin, et al.
Publicado: (2026)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
por: Liu, Runtao, et al.
Publicado: (2024)
por: Liu, Runtao, et al.
Publicado: (2024)
LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction
por: Jin, Er, et al.
Publicado: (2025)
por: Jin, Er, et al.
Publicado: (2025)
SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
por: Gülhan, Fabian, et al.
Publicado: (2025)
por: Gülhan, Fabian, et al.
Publicado: (2025)
Unsupervised Learning for Feature Extraction and Temporal Alignment of 3D+t Point Clouds of Zebrafish Embryos
por: Chen, Zhu, et al.
Publicado: (2025)
por: Chen, Zhu, et al.
Publicado: (2025)
A Survey on Transferability of Adversarial Examples across Deep Neural Networks
por: Gu, Jindong, et al.
Publicado: (2023)
por: Gu, Jindong, et al.
Publicado: (2023)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
por: Luo, Haochen, et al.
Publicado: (2024)
por: Luo, Haochen, et al.
Publicado: (2024)
Annotated Biomedical Video Generation using Denoising Diffusion Probabilistic Models and Flow Fields
por: Yilmaz, Rüveyda, et al.
Publicado: (2024)
por: Yilmaz, Rüveyda, et al.
Publicado: (2024)
Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
por: Ma, Xiangkai, et al.
Publicado: (2025)
por: Ma, Xiangkai, et al.
Publicado: (2025)
Cascaded Diffusion Models for 2D and 3D Microscopy Image Synthesis to Enhance Cell Segmentation
por: Yilmaz, Rüveyda, et al.
Publicado: (2024)
por: Yilmaz, Rüveyda, et al.
Publicado: (2024)
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
por: Xu, Hengbo, et al.
Publicado: (2026)
por: Xu, Hengbo, et al.
Publicado: (2026)
Learning Camera Movement Control from Real-World Drone Videos
por: Hou, Yunzhong, et al.
Publicado: (2024)
por: Hou, Yunzhong, et al.
Publicado: (2024)
Which Model Generated This Image? A Model-Agnostic Approach for Origin Attribution
por: Liu, Fengyuan, et al.
Publicado: (2024)
por: Liu, Fengyuan, et al.
Publicado: (2024)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
por: Xing, Songlong, et al.
Publicado: (2026)
por: Xing, Songlong, et al.
Publicado: (2026)
AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
por: Ganj, Ashkan, et al.
Publicado: (2025)
por: Ganj, Ashkan, et al.
Publicado: (2025)
A Pragmatic Note on Evaluating Generative Models with Fréchet Inception Distance for Retinal Image Synthesis
por: Wu, Yuli, et al.
Publicado: (2025)
por: Wu, Yuli, et al.
Publicado: (2025)
Replacement Learning: Training Vision Tasks with Fewer Learnable Parameters
por: Zhang, Yuming, et al.
Publicado: (2024)
por: Zhang, Yuming, et al.
Publicado: (2024)
Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models
por: Zhang, Yang, et al.
Publicado: (2024)
por: Zhang, Yang, et al.
Publicado: (2024)
Vision-Language Models Do Not Understand Negation
por: Alhamoud, Kumail, et al.
Publicado: (2025)
por: Alhamoud, Kumail, et al.
Publicado: (2025)
Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior
por: Li, Xiang, et al.
Publicado: (2026)
por: Li, Xiang, et al.
Publicado: (2026)
VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models
por: Han, Junlin, et al.
Publicado: (2024)
por: Han, Junlin, et al.
Publicado: (2024)
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
por: Li, Runjia, et al.
Publicado: (2025)
por: Li, Runjia, et al.
Publicado: (2025)
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
por: Yuan, Jianhao, et al.
Publicado: (2022)
por: Yuan, Jianhao, et al.
Publicado: (2022)
Semantic Score Distillation Sampling for Compositional Text-to-3D Generation
por: Yang, Ling, et al.
Publicado: (2024)
por: Yang, Ling, et al.
Publicado: (2024)
CellStyle: Improved Zero-Shot Cell Segmentation via Style Transfer
por: Yilmaz, Rüveyda, et al.
Publicado: (2025)
por: Yilmaz, Rüveyda, et al.
Publicado: (2025)
Optimizing Retinal Prosthetic Stimuli with Conditional Invertible Neural Networks
por: Wu, Yuli, et al.
Publicado: (2024)
por: Wu, Yuli, et al.
Publicado: (2024)
PET-Adapter: Test-Time Domain Adaptation for Full and Limited-Angle PET Image Reconstruction
por: Yilmaz, Rüveyda, et al.
Publicado: (2026)
por: Yilmaz, Rüveyda, et al.
Publicado: (2026)
Towards Robust and Generalizable Gerchberg Saxton based Physics Inspired Neural Networks for Computer Generated Holography: A Sensitivity Analysis Framework
por: Amrutkar, Ankit, et al.
Publicado: (2025)
por: Amrutkar, Ankit, et al.
Publicado: (2025)
Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image Generation
por: Luo, Yihong, et al.
Publicado: (2025)
por: Luo, Yihong, et al.
Publicado: (2025)
Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation
por: Wang, Zihao, et al.
Publicado: (2026)
por: Wang, Zihao, et al.
Publicado: (2026)
Towards Interpreting Visual Information Processing in Vision-Language Models
por: Neo, Clement, et al.
Publicado: (2024)
por: Neo, Clement, et al.
Publicado: (2024)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
por: Zhu, Jiaying, et al.
Publicado: (2025)
por: Zhu, Jiaying, et al.
Publicado: (2025)
Ejemplares similares
-
Minimalist Concept Erasure in Generative Models
por: Zhang, Yang, et al.
Publicado: (2025) -
Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
por: Jin, Er, et al.
Publicado: (2025) -
Learning Visual Prompts for Guiding the Attention of Vision Transformers
por: Rezaei, Razieh, et al.
Publicado: (2024) -
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
por: Deb, Oishi, et al.
Publicado: (2025) -
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
por: Hemmat, Arshia, et al.
Publicado: (2024)