Salvato in:
| Autori principali: | Salman, Shaeke, Shams, Md Montasir Bin, Liu, Xiuwen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2401.15568 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
Malicious Path Manipulations via Exploitation of Representation Vulnerabilities of Vision-Language Navigation Systems
di: Islam, Chashi Mahiul, et al.
Pubblicazione: (2024)
di: Islam, Chashi Mahiul, et al.
Pubblicazione: (2024)
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
di: Shams, Montasir, et al.
Pubblicazione: (2025)
di: Shams, Montasir, et al.
Pubblicazione: (2025)
Intriguing Properties of Data Attribution on Diffusion Models
di: Zheng, Xiaosen, et al.
Pubblicazione: (2023)
di: Zheng, Xiaosen, et al.
Pubblicazione: (2023)
Intriguing properties of generative classifiers
di: Jaini, Priyank, et al.
Pubblicazione: (2023)
di: Jaini, Priyank, et al.
Pubblicazione: (2023)
Topological Alignment of Shared Vision-Language Embedding Space
di: You, Junwon, et al.
Pubblicazione: (2025)
di: You, Junwon, et al.
Pubblicazione: (2025)
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
di: Pantazopoulos, Georgios, et al.
Pubblicazione: (2024)
di: Pantazopoulos, Georgios, et al.
Pubblicazione: (2024)
Data-Driven Fairness Generalization for Deepfake Detection
di: Ezeakunne, Uzoamaka, et al.
Pubblicazione: (2024)
di: Ezeakunne, Uzoamaka, et al.
Pubblicazione: (2024)
Robust Asymmetric Heterogeneous Federated Learning with Corrupted Clients
di: Fang, Xiuwen, et al.
Pubblicazione: (2025)
di: Fang, Xiuwen, et al.
Pubblicazione: (2025)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
AI-Powered Deepfake Detection Using CNN and Vision Transformer Architectures
di: Urmi, Sifatullah Sheikh, et al.
Pubblicazione: (2026)
di: Urmi, Sifatullah Sheikh, et al.
Pubblicazione: (2026)
Improving Interpretation Faithfulness for Vision Transformers
di: Hu, Lijie, et al.
Pubblicazione: (2023)
di: Hu, Lijie, et al.
Pubblicazione: (2023)
ZAYAN: Disentangled Contrastive Transformer for Tabular Remote Sensing Data
di: Habib, Al Zadid Sultan Bin, et al.
Pubblicazione: (2026)
di: Habib, Al Zadid Sultan Bin, et al.
Pubblicazione: (2026)
Robust Multimodal Learning via Cross-Modal Proxy Tokens
di: Reza, Md Kaykobad, et al.
Pubblicazione: (2025)
di: Reza, Md Kaykobad, et al.
Pubblicazione: (2025)
DiffiT: Diffusion Vision Transformers for Image Generation
di: Hatamizadeh, Ali, et al.
Pubblicazione: (2023)
di: Hatamizadeh, Ali, et al.
Pubblicazione: (2023)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
di: Trager, Matthew, et al.
Pubblicazione: (2023)
di: Trager, Matthew, et al.
Pubblicazione: (2023)
Discovering Influential Neuron Path in Vision Transformers
di: Wang, Yifan, et al.
Pubblicazione: (2025)
di: Wang, Yifan, et al.
Pubblicazione: (2025)
ScaleKD: Strong Vision Transformers Could Be Excellent Teachers
di: Fan, Jiawei, et al.
Pubblicazione: (2024)
di: Fan, Jiawei, et al.
Pubblicazione: (2024)
Convolutional Neural Nets vs Vision Transformers: A SpaceNet Case Study with Balanced vs Imbalanced Regimes
di: Gothi, Akshar
Pubblicazione: (2025)
di: Gothi, Akshar
Pubblicazione: (2025)
Block-Recurrent Dynamics in Vision Transformers
di: Jacobs, Mozes, et al.
Pubblicazione: (2025)
di: Jacobs, Mozes, et al.
Pubblicazione: (2025)
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
di: Schlarmann, Christian, et al.
Pubblicazione: (2024)
di: Schlarmann, Christian, et al.
Pubblicazione: (2024)
On Background Bias of Post-Hoc Concept Embeddings in Computer Vision DNNs
di: Schwalbe, Gesina, et al.
Pubblicazione: (2025)
di: Schwalbe, Gesina, et al.
Pubblicazione: (2025)
ADAPT to Robustify Prompt Tuning Vision Transformers
di: Eskandar, Masih, et al.
Pubblicazione: (2024)
di: Eskandar, Masih, et al.
Pubblicazione: (2024)
Continual Adaptation of Vision Transformers for Federated Learning
di: Halbe, Shaunak, et al.
Pubblicazione: (2023)
di: Halbe, Shaunak, et al.
Pubblicazione: (2023)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
di: Kumar, Anantha Padmanaban Krishna
Pubblicazione: (2025)
di: Kumar, Anantha Padmanaban Krishna
Pubblicazione: (2025)
Accelerating Vision Transformers with Adaptive Patch Sizes
di: Choudhury, Rohan, et al.
Pubblicazione: (2025)
di: Choudhury, Rohan, et al.
Pubblicazione: (2025)
Class-Discriminative Attention Maps for Vision Transformers
di: Brocki, Lennart, et al.
Pubblicazione: (2023)
di: Brocki, Lennart, et al.
Pubblicazione: (2023)
Exploring Token Pruning in Vision State Space Models
di: Zhan, Zheng, et al.
Pubblicazione: (2024)
di: Zhan, Zheng, et al.
Pubblicazione: (2024)
AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
di: Munir, Mustafa, et al.
Pubblicazione: (2025)
di: Munir, Mustafa, et al.
Pubblicazione: (2025)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
di: Guimard, Quentin, et al.
Pubblicazione: (2026)
di: Guimard, Quentin, et al.
Pubblicazione: (2026)
SkipViT: Speeding Up Vision Transformers with a Token-Level Skip Connection
di: Ataiefard, Foozhan, et al.
Pubblicazione: (2024)
di: Ataiefard, Foozhan, et al.
Pubblicazione: (2024)
Vision-Based Localization and LLM-based Navigation for Indoor Environments
di: Rahimi, Keyan, et al.
Pubblicazione: (2025)
di: Rahimi, Keyan, et al.
Pubblicazione: (2025)
Oscillation-Reduced MXFP4 Training for Vision Transformers
di: Chen, Yuxiang, et al.
Pubblicazione: (2025)
di: Chen, Yuxiang, et al.
Pubblicazione: (2025)
Enhancing Vision Transformer Explainability Using Artificial Astrocytes
di: Echevarrieta-Catalan, Nicolas, et al.
Pubblicazione: (2025)
di: Echevarrieta-Catalan, Nicolas, et al.
Pubblicazione: (2025)
FasterViT: Fast Vision Transformers with Hierarchical Attention
di: Hatamizadeh, Ali, et al.
Pubblicazione: (2023)
di: Hatamizadeh, Ali, et al.
Pubblicazione: (2023)
Lightweight Model for Poultry Disease Detection from Fecal Images Using Multi-Color Space Feature Optimization and Machine Learning
di: Islam, A. K. M. Shoriful, et al.
Pubblicazione: (2025)
di: Islam, A. K. M. Shoriful, et al.
Pubblicazione: (2025)
Structure-Guided Adversarial Training of Diffusion Models
di: Yang, Ling, et al.
Pubblicazione: (2024)
di: Yang, Ling, et al.
Pubblicazione: (2024)
GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNs
di: Munir, Mustafa, et al.
Pubblicazione: (2024)
di: Munir, Mustafa, et al.
Pubblicazione: (2024)
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
di: Cao, Yuefan, et al.
Pubblicazione: (2025)
di: Cao, Yuefan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024) -
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024) -
Malicious Path Manipulations via Exploitation of Representation Vulnerabilities of Vision-Language Navigation Systems
di: Islam, Chashi Mahiul, et al.
Pubblicazione: (2024) -
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
di: Shams, Montasir, et al.
Pubblicazione: (2025) -
Intriguing Properties of Data Attribution on Diffusion Models
di: Zheng, Xiaosen, et al.
Pubblicazione: (2023)