Saved in:
| Main Authors: | Qiang, Yao, Li, Chengyin, Khanduri, Prashant, Zhu, Dongxiao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2309.08035 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
by: Sultan, Rafi Ibn, et al.
Published: (2026)
by: Sultan, Rafi Ibn, et al.
Published: (2026)
AutoProSAM: Automated Prompting SAM for 3D Multi-Organ Segmentation
by: Li, Chengyin, et al.
Published: (2023)
by: Li, Chengyin, et al.
Published: (2023)
GeoSAM: Fine-tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation
by: Sultan, Rafi Ibn, et al.
Published: (2023)
by: Sultan, Rafi Ibn, et al.
Published: (2023)
MulModSeg: Enhancing Unpaired Multi-Modal Medical Image Segmentation with Modality-Conditioned Text Embedding and Alternating Training
by: Li, Chengyin, et al.
Published: (2024)
by: Li, Chengyin, et al.
Published: (2024)
BiPVL-Seg: Bidirectional Progressive Vision-Language Fusion with Global-Local Alignment for Medical Image Segmentation
by: Sultan, Rafi Ibn, et al.
Published: (2025)
by: Sultan, Rafi Ibn, et al.
Published: (2025)
Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks
by: Zhao, Yingying, et al.
Published: (2026)
by: Zhao, Yingying, et al.
Published: (2026)
Robustness of Transformer-Based Fluence Map Prediction Under Clinically Realistic Perturbations
by: Mgboh, Ujunwa, et al.
Published: (2026)
by: Mgboh, Ujunwa, et al.
Published: (2026)
ComFe: An Interpretable Head for Vision Transformers
by: Mannix, Evelyn J., et al.
Published: (2024)
by: Mannix, Evelyn J., et al.
Published: (2024)
Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patches for Infrared Vision-Language Models
by: Hu, Chengyin, et al.
Published: (2026)
by: Hu, Chengyin, et al.
Published: (2026)
Interpretable Vision Transformers in Image Classification via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
Serial Low-rank Adaptation of Vision Transformer
by: Zhong, Houqiang, et al.
Published: (2025)
by: Zhong, Houqiang, et al.
Published: (2025)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
by: Böhle, Moritz, et al.
Published: (2023)
by: Böhle, Moritz, et al.
Published: (2023)
Interpretable Vision Transformers in Monocular Depth Estimation via SVDA
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
by: Ma, Chiyu, et al.
Published: (2024)
by: Ma, Chiyu, et al.
Published: (2024)
Hierarchical Vision Transformer with Prototypes for Interpretable Medical Image Classification
by: Gallée, Luisa, et al.
Published: (2025)
by: Gallée, Luisa, et al.
Published: (2025)
Context-Aware Token Selection and Packing for Enhanced Vision Transformer
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024)
by: Ganz, Roy, et al.
Published: (2024)
Neighbor-Aware Token Reduction via Hilbert Curve for Vision Transformers
by: Li, Yunge, et al.
Published: (2025)
by: Li, Yunge, et al.
Published: (2025)
When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models
by: Hu, Chengyin, et al.
Published: (2026)
by: Hu, Chengyin, et al.
Published: (2026)
Sparse Reasoning is Enough: Biological-Inspired Framework for Video Anomaly Detection with Large Pre-trained Models
by: Huang, He, et al.
Published: (2025)
by: Huang, He, et al.
Published: (2025)
Temporal-Spatial Object Relations Modeling for Vision-and-Language Navigation
by: Huang, Bowen, et al.
Published: (2024)
by: Huang, Bowen, et al.
Published: (2024)
FluenceFormer: Transformer-Driven Multi-Beam Fluence Map Regression for Radiotherapy Planning
by: Mgboh, Ujunwa, et al.
Published: (2025)
by: Mgboh, Ujunwa, et al.
Published: (2025)
Sparse but not Simpler: A Multi-Level Interpretability Analysis of Vision Transformers
by: Zhang, Siyu
Published: (2026)
by: Zhang, Siyu
Published: (2026)
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
by: Hong, Jung-Ho, et al.
Published: (2025)
by: Hong, Jung-Ho, et al.
Published: (2025)
[Re] Improving Interpretation Faithfulness for Vision Transformers
by: Kurek, Izabela, et al.
Published: (2025)
by: Kurek, Izabela, et al.
Published: (2025)
Multi-View Black-Box Physical Attacks on Infrared Pedestrian Detectors Using Adversarial Infrared Grid
by: Tiliwalidi, Kalibinuer, et al.
Published: (2024)
by: Tiliwalidi, Kalibinuer, et al.
Published: (2024)
Decision-Aware Attention Propagation for Vision Transformer Explainability
by: Jo, Sehyeong, et al.
Published: (2026)
by: Jo, Sehyeong, et al.
Published: (2026)
Fluence Map Prediction with Deep Learning: A Transformer-based Approach
by: Mgboh, Ujunwa, et al.
Published: (2025)
by: Mgboh, Ujunwa, et al.
Published: (2025)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
DeepHistoViT: An Interpretable Vision Transformer Framework for Histopathological Cancer Classification
by: Mosalpuri, Ravi, et al.
Published: (2026)
by: Mosalpuri, Ravi, et al.
Published: (2026)
GMAR: Gradient-Driven Multi-Head Attention Rollout for Vision Transformer Interpretability
by: Jo, Sehyeong, et al.
Published: (2025)
by: Jo, Sehyeong, et al.
Published: (2025)
Diffusion Once and Done: Degradation-Aware LoRA for Efficient All-in-One Image Restoration
by: Tang, Ni, et al.
Published: (2025)
by: Tang, Ni, et al.
Published: (2025)
PASTS: Progress-Aware Spatio-Temporal Transformer Speaker For Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2023)
by: Wang, Liuyi, et al.
Published: (2023)
HGFormer: Topology-Aware Vision Transformer with HyperGraph Learning
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Boundary-Aware Vision Transformer for Angiography Vascular Network Segmentation
by: Hezil, Nabil, et al.
Published: (2025)
by: Hezil, Nabil, et al.
Published: (2025)
Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
by: Liang, Wenhao, et al.
Published: (2025)
by: Liang, Wenhao, et al.
Published: (2025)
Thermal Topology Collapse: Universal Physical Patch Attacks on Infrared Vision Systems
by: Hu, Chengyin, et al.
Published: (2026)
by: Hu, Chengyin, et al.
Published: (2026)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Wavelet-Based Image Tokenizer for Vision Transformers
by: Zhu, Zhenhai, et al.
Published: (2024)
by: Zhu, Zhenhai, et al.
Published: (2024)
Similar Items
-
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023) -
WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
by: Sultan, Rafi Ibn, et al.
Published: (2026) -
AutoProSAM: Automated Prompting SAM for 3D Multi-Organ Segmentation
by: Li, Chengyin, et al.
Published: (2023) -
GeoSAM: Fine-tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation
by: Sultan, Rafi Ibn, et al.
Published: (2023) -
MulModSeg: Enhancing Unpaired Multi-Modal Medical Image Segmentation with Modality-Conditioned Text Embedding and Alternating Training
by: Li, Chengyin, et al.
Published: (2024)