Focus Your Attention: Towards Data-Intuitive Lightweight Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Gaurav, Suyash, Humayun, Muhammad Farhan, Heikkonen, Jukka, Chaudhary, Jatin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement
by: Gaurav, Suyash, et al.
Published: (2025)
by: Gaurav, Suyash, et al.
Published: (2025)
Pathway-based Progressive Inference (PaPI) for Energy-Efficient Continual Learning
by: Gaurav, Suyash, et al.
Published: (2025)
by: Gaurav, Suyash, et al.
Published: (2025)
From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
Elastic Attention Cores for Scalable Vision Transformers
by: Song, Alan Z., et al.
Published: (2026)
by: Song, Alan Z., et al.
Published: (2026)
Attention Transfer Is Not Universally Effective for Vision Transformers
by: Qin, Huaiyuan, et al.
Published: (2026)
by: Qin, Huaiyuan, et al.
Published: (2026)
Pre-training of Lightweight Vision Transformers on Small Datasets with Minimally Scaled Images
by: Tan, Jen Hong
Published: (2024)
by: Tan, Jen Hong
Published: (2024)
Opinion: Learning Intuitive Physics May Require More than Visual Data
by: Su, Ellen, et al.
Published: (2025)
by: Su, Ellen, et al.
Published: (2025)
MABViT -- Modified Attention Block Enhances Vision Transformers
by: Ramesh, Mahesh, et al.
Published: (2023)
by: Ramesh, Mahesh, et al.
Published: (2023)
Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
by: Shan, Jiquan, et al.
Published: (2025)
by: Shan, Jiquan, et al.
Published: (2025)
PCaM: A Progressive Focus Attention-Based Information Fusion Method for Improving Vision Transformer Domain Adaptation
by: Zang, Zelin, et al.
Published: (2025)
by: Zang, Zelin, et al.
Published: (2025)
What Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models
by: Humayun, Ahmed Imtiaz, et al.
Published: (2024)
by: Humayun, Ahmed Imtiaz, et al.
Published: (2024)
ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers
by: Hsu, Chih-Chung, et al.
Published: (2026)
by: Hsu, Chih-Chung, et al.
Published: (2026)
Discrete Cosine Transform Based Decorrelated Attention for Vision Transformers
by: Pan, Hongyi, et al.
Published: (2024)
by: Pan, Hongyi, et al.
Published: (2024)
Class-Discriminative Attention Maps for Vision Transformers
by: Brocki, Lennart, et al.
Published: (2023)
by: Brocki, Lennart, et al.
Published: (2023)
LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
by: Setyawan, Novendra, et al.
Published: (2024)
by: Setyawan, Novendra, et al.
Published: (2024)
Toward Unified Multimodal Representation Learning for Autonomous Driving
by: Tao, Ximeng, et al.
Published: (2026)
by: Tao, Ximeng, et al.
Published: (2026)
Ocular Disease Classification Using CNN with Deep Convolutional Generative Adversarial Network
by: Kunwar, Arun, et al.
Published: (2025)
by: Kunwar, Arun, et al.
Published: (2025)
Data-independent Module-aware Pruning for Hierarchical Vision Transformers
by: He, Yang, et al.
Published: (2024)
by: He, Yang, et al.
Published: (2024)
Make Your LVLM KV Cache More Lightweight
by: Chen, Xihao, et al.
Published: (2026)
by: Chen, Xihao, et al.
Published: (2026)
A Lightweight Transformer with Phase-Only Cross-Attention for Illumination-Invariant Biometric Authentication
by: Sharma, Arun K., et al.
Published: (2024)
by: Sharma, Arun K., et al.
Published: (2024)
Towards Improved Cervical Cancer Screening: Vision Transformer-Based Classification and Interpretability
by: Nguyen, Khoa Tuan, et al.
Published: (2025)
by: Nguyen, Khoa Tuan, et al.
Published: (2025)
Zero-Shot Generalization of Vision-Based RL Without Data Augmentation
by: Batra, Sumeet, et al.
Published: (2024)
by: Batra, Sumeet, et al.
Published: (2024)
ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition
by: Gupta, Rishabh, et al.
Published: (2026)
by: Gupta, Rishabh, et al.
Published: (2026)
Yo'LLaVA: Your Personalized Language and Vision Assistant
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
DAUNet: A Lightweight UNet Variant with Deformable Convolutions and Parameter-Free Attention for Medical Image Segmentation
by: Munir, Adnan, et al.
Published: (2025)
by: Munir, Adnan, et al.
Published: (2025)
Pointy - A Lightweight Transformer for Point Cloud Foundation Models
by: Szafer, Konrad, et al.
Published: (2026)
by: Szafer, Konrad, et al.
Published: (2026)
A Lightweight Large Vision-language Model for Multimodal Medical Images
by: Alsinglawi, Belal, et al.
Published: (2025)
by: Alsinglawi, Belal, et al.
Published: (2025)
S-E Pipeline: A Vision Transformer (ViT) based Resilient Classification Pipeline for Medical Imaging Against Adversarial Attacks
by: S, Neha A, et al.
Published: (2024)
by: S, Neha A, et al.
Published: (2024)
SplineCam: Exact Visualization and Characterization of Deep Network Geometry and Decision Boundaries
by: Humayun, Ahmed Imtiaz, et al.
Published: (2023)
by: Humayun, Ahmed Imtiaz, et al.
Published: (2023)
SegMate: Asymmetric Attention-Based Lightweight Architecture for Efficient Multi-Organ Segmentation
by: Bunea, Andrei-Alexandru, et al.
Published: (2026)
by: Bunea, Andrei-Alexandru, et al.
Published: (2026)
Visual Concept Networks: A Graph-Based Approach to Detecting Anomalous Data in Deep Neural Networks
by: Ganguly, Debargha, et al.
Published: (2024)
by: Ganguly, Debargha, et al.
Published: (2024)
Beyond ImageNet: Understanding Cross-Dataset Robustness of Lightweight Vision Models
by: Zhang, Weidong, et al.
Published: (2025)
by: Zhang, Weidong, et al.
Published: (2025)
GraphTARIF: Linear Graph Transformer with Augmented Rank and Improved Focus
by: Hu, Zhaolin, et al.
Published: (2025)
by: Hu, Zhaolin, et al.
Published: (2025)
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
Attention-Guided Dual-Stream Learning for Group Engagement Recognition: Fusing Transformer-Encoded Motion Dynamics with Scene Context via Adaptive Gating
by: Chowdhury, Saniah Kayenat, et al.
Published: (2026)
by: Chowdhury, Saniah Kayenat, et al.
Published: (2026)
Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
by: Srivastava, Divyansh, et al.
Published: (2025)
by: Srivastava, Divyansh, et al.
Published: (2025)
Similar Items
-
Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement
by: Gaurav, Suyash, et al.
Published: (2025) -
Pathway-based Progressive Inference (PaPI) for Energy-Efficient Continual Learning
by: Gaurav, Suyash, et al.
Published: (2025) -
From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)
by: Mishra, Suyash, et al.
Published: (2026) -
Elastic Attention Cores for Scalable Vision Transformers
by: Song, Alan Z., et al.
Published: (2026) -
Attention Transfer Is Not Universally Effective for Vision Transformers
by: Qin, Huaiyuan, et al.
Published: (2026)