MULTIFLOW: Shifting Towards Task-Agnostic Vision-Language Pruning
Fuente:
arXiv
Guardado en:
| Autores principales: | Farina, Matteo, Mancini, Massimiliano, Cunegatti, Elia, Liu, Gaowen, Iacca, Giovanni, Ricci, Elisa |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
por: Farina, Matteo, et al.
Publicado: (2025)
por: Farina, Matteo, et al.
Publicado: (2025)
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
por: Farina, Matteo, et al.
Publicado: (2024)
por: Farina, Matteo, et al.
Publicado: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
por: Berasi, Davide, et al.
Publicado: (2025)
por: Berasi, Davide, et al.
Publicado: (2025)
Compositional Caching for Training-free Open-vocabulary Attribute Detection
por: Garosi, Marco, et al.
Publicado: (2025)
por: Garosi, Marco, et al.
Publicado: (2025)
Large Multimodal Models as General In-Context Classifiers
por: Garosi, Marco, et al.
Publicado: (2026)
por: Garosi, Marco, et al.
Publicado: (2026)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
por: Caldarella, Simone, et al.
Publicado: (2024)
por: Caldarella, Simone, et al.
Publicado: (2024)
Harnessing Large Language Models for Training-free Video Anomaly Detection
por: Zanella, Luca, et al.
Publicado: (2024)
por: Zanella, Luca, et al.
Publicado: (2024)
Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization
por: Monaco, Francesco Pio, et al.
Publicado: (2026)
por: Monaco, Francesco Pio, et al.
Publicado: (2026)
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
por: De Min, Thomas, et al.
Publicado: (2026)
por: De Min, Thomas, et al.
Publicado: (2026)
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
por: Sandri, Fabrizio, et al.
Publicado: (2025)
por: Sandri, Fabrizio, et al.
Publicado: (2025)
Test-time Vocabulary Adaptation for Language-driven Object Detection
por: Liu, Mingxuan, et al.
Publicado: (2025)
por: Liu, Mingxuan, et al.
Publicado: (2025)
Can Text-to-Video Generation help Video-Language Alignment?
por: Zanella, Luca, et al.
Publicado: (2025)
por: Zanella, Luca, et al.
Publicado: (2025)
Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
por: Guimard, Quentin, et al.
Publicado: (2025)
por: Guimard, Quentin, et al.
Publicado: (2025)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
por: Guimard, Quentin, et al.
Publicado: (2026)
por: Guimard, Quentin, et al.
Publicado: (2026)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
por: Das, Deepayan, et al.
Publicado: (2024)
por: Das, Deepayan, et al.
Publicado: (2024)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
por: Das, Deepayan, et al.
Publicado: (2025)
por: Das, Deepayan, et al.
Publicado: (2025)
Training-free Online Video Step Grounding
por: Zanella, Luca, et al.
Publicado: (2025)
por: Zanella, Luca, et al.
Publicado: (2025)
Unlearning Personal Data from a Single Image
por: De Min, Thomas, et al.
Publicado: (2024)
por: De Min, Thomas, et al.
Publicado: (2024)
Zeroth-Order Adaptive Neuron Alignment Based Pruning without Re-Training
por: Cunegatti, Elia, et al.
Publicado: (2024)
por: Cunegatti, Elia, et al.
Publicado: (2024)
Vision-by-Language for Training-Free Compositional Image Retrieval
por: Karthik, Shyamgopal, et al.
Publicado: (2023)
por: Karthik, Shyamgopal, et al.
Publicado: (2023)
Vocabulary-free Image Classification and Semantic Segmentation
por: Conti, Alessandro, et al.
Publicado: (2024)
por: Conti, Alessandro, et al.
Publicado: (2024)
Vocabulary-free Image Classification
por: Conti, Alessandro, et al.
Publicado: (2023)
por: Conti, Alessandro, et al.
Publicado: (2023)
On Large Multimodal Models as Open-World Image Classifiers
por: Conti, Alessandro, et al.
Publicado: (2025)
por: Conti, Alessandro, et al.
Publicado: (2025)
Safe Vision-Language Models via Unsafe Weights Manipulation
por: D'Incà, Moreno, et al.
Publicado: (2025)
por: D'Incà, Moreno, et al.
Publicado: (2025)
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
por: Huang, Yiran, et al.
Publicado: (2025)
por: Huang, Yiran, et al.
Publicado: (2025)
From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
por: Gentile, Francesco, et al.
Publicado: (2026)
por: Gentile, Francesco, et al.
Publicado: (2026)
Less is more: Summarizing Patch Tokens for efficient Multi-Label Class-Incremental Learning
por: De Min, Thomas, et al.
Publicado: (2024)
por: De Min, Thomas, et al.
Publicado: (2024)
VLTP: Vision-Language Guided Token Pruning for Task-Oriented Segmentation
por: Chen, Hanning, et al.
Publicado: (2024)
por: Chen, Hanning, et al.
Publicado: (2024)
Towards Joint Quantization and Token Pruning of Vision-Language Models
por: Li, Xinqing, et al.
Publicado: (2026)
por: Li, Xinqing, et al.
Publicado: (2026)
Understanding Sparse Neural Networks from their Topology via Multipartite Graph Representations
por: Cunegatti, Elia, et al.
Publicado: (2023)
por: Cunegatti, Elia, et al.
Publicado: (2023)
Comb, Prune, Distill: Towards Unified Pruning for Vision Model Compression
por: Schmitt, Jonas, et al.
Publicado: (2024)
por: Schmitt, Jonas, et al.
Publicado: (2024)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
por: Bonat, Laurence, et al.
Publicado: (2026)
por: Bonat, Laurence, et al.
Publicado: (2026)
Automatic benchmarking of large multimodal models via iterative experiment programming
por: Conti, Alessandro, et al.
Publicado: (2024)
por: Conti, Alessandro, et al.
Publicado: (2024)
Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models
por: Seutin, Corentin, et al.
Publicado: (2026)
por: Seutin, Corentin, et al.
Publicado: (2026)
Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models
por: Ma, Kexin, et al.
Publicado: (2026)
por: Ma, Kexin, et al.
Publicado: (2026)
Neuron-centric Hebbian Learning
por: Ferigo, Andrea, et al.
Publicado: (2024)
por: Ferigo, Andrea, et al.
Publicado: (2024)
RESTORE: Towards Feature Shift for Vision-Language Prompt Learning
por: Yang, Yuncheng, et al.
Publicado: (2024)
por: Yang, Yuncheng, et al.
Publicado: (2024)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
por: Liu, Jizhihui, et al.
Publicado: (2025)
por: Liu, Jizhihui, et al.
Publicado: (2025)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
por: D'Incà, Moreno, et al.
Publicado: (2024)
por: D'Incà, Moreno, et al.
Publicado: (2024)
Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge
por: Eliopoulos, Nick John, et al.
Publicado: (2024)
por: Eliopoulos, Nick John, et al.
Publicado: (2024)
Ejemplares similares
-
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
por: Farina, Matteo, et al.
Publicado: (2025) -
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
por: Farina, Matteo, et al.
Publicado: (2024) -
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
por: Berasi, Davide, et al.
Publicado: (2025) -
Compositional Caching for Training-free Open-vocabulary Attribute Detection
por: Garosi, Marco, et al.
Publicado: (2025) -
Large Multimodal Models as General In-Context Classifiers
por: Garosi, Marco, et al.
Publicado: (2026)