Unified Local and Global Attention Interaction Modeling for Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Tan, Heldermon, Coy D., Toler-Franklin, Corey |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Event-based Solutions for Human-centered Applications: A Comprehensive Review
by: Adra, Mira, et al.
Published: (2025)
by: Adra, Mira, et al.
Published: (2025)
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
by: Bouzid, Hamza, et al.
Published: (2023)
by: Bouzid, Hamza, et al.
Published: (2023)
A Review of Pseudo-Labeling for Computer Vision
by: Kage, Patrick, et al.
Published: (2024)
by: Kage, Patrick, et al.
Published: (2024)
Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification
by: Ma, Shuxian, et al.
Published: (2025)
by: Ma, Shuxian, et al.
Published: (2025)
Efficient Neural Network Encoding for 3D Color Lookup Tables
by: Zehtab, Vahid, et al.
Published: (2024)
by: Zehtab, Vahid, et al.
Published: (2024)
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
by: Ma, Chong, et al.
Published: (2024)
by: Ma, Chong, et al.
Published: (2024)
Goal-conditioned reinforcement learning for ultrasound navigation guidance
by: Amadou, Abdoul Aziz, et al.
Published: (2024)
by: Amadou, Abdoul Aziz, et al.
Published: (2024)
Few-Class Arena: A Benchmark for Efficient Selection of Vision Models and Dataset Difficulty Measurement
by: Cao, Bryan Bo, et al.
Published: (2024)
by: Cao, Bryan Bo, et al.
Published: (2024)
A systematic review: Deep learning-based methods for pneumonia region detection
by: Xu, Xinmei
Published: (2024)
by: Xu, Xinmei
Published: (2024)
ViFiCon: Vision and Wireless Association Via Self-Supervised Contrastive Learning
by: Meegan, Nicholas, et al.
Published: (2022)
by: Meegan, Nicholas, et al.
Published: (2022)
Texture Discrimination via Hilbert Curve Path Based Information Quantifiers
by: Bariviera, Aurelio F., et al.
Published: (2024)
by: Bariviera, Aurelio F., et al.
Published: (2024)
Bridging SFT and DPO for Diffusion Model Alignment with Self-Sampling Preference Optimization
by: Zhang, Daoan, et al.
Published: (2024)
by: Zhang, Daoan, et al.
Published: (2024)
Nash Meets Wertheimer: Using Good Continuation in Jigsaw Puzzles
by: Khoroshiltseva, Marina, et al.
Published: (2024)
by: Khoroshiltseva, Marina, et al.
Published: (2024)
Pairwise Spatiotemporal Partial Trajectory Matching for Co-movement Analysis
by: Cardei, Maria, et al.
Published: (2024)
by: Cardei, Maria, et al.
Published: (2024)
StatsMerging: Statistics-Guided Model Merging via Task-Specific Teacher Distillation
by: Merugu, Ranjith, et al.
Published: (2025)
by: Merugu, Ranjith, et al.
Published: (2025)
FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare
by: Lekadir, Karim, et al.
Published: (2023)
by: Lekadir, Karim, et al.
Published: (2023)
Enhancing Explainable AI: A Hybrid Approach Combining GradCAM and LRP for CNN Interpretability
by: Dhore, Vaibhav, et al.
Published: (2024)
by: Dhore, Vaibhav, et al.
Published: (2024)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
An Overview of the Burer-Monteiro Method for Certifiable Robot Perception
by: Papalia, Alan, et al.
Published: (2024)
by: Papalia, Alan, et al.
Published: (2024)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
by: Maity, Subhajit, et al.
Published: (2025)
by: Maity, Subhajit, et al.
Published: (2025)
FLD+: Data-efficient Evaluation Metric for Generative Models
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Scalable Face Security Vision Foundation Model for Deepfake, Diffusion, and Spoofing Detection
by: Wang, Gaojian, et al.
Published: (2025)
by: Wang, Gaojian, et al.
Published: (2025)
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
by: Zhang, Shengkai, et al.
Published: (2024)
by: Zhang, Shengkai, et al.
Published: (2024)
Evaluating ML Robustness in GNSS Interference Classification, Characterization & Localization
by: Heublein, Lucas, et al.
Published: (2024)
by: Heublein, Lucas, et al.
Published: (2024)
Evaluation Metric for Quality Control and Generative Models in Histopathology Images
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Context-dependent Causality (the Non-Nonotonic Case)
by: Billfeld, Nir, et al.
Published: (2024)
by: Billfeld, Nir, et al.
Published: (2024)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
by: Martin, Michael R., et al.
Published: (2025)
by: Martin, Michael R., et al.
Published: (2025)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Normalizing Flow-Based Metric for Image Generation
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
WaveMix: A Resource-efficient Neural Network for Image Analysis
by: Jeevan, Pranav, et al.
Published: (2022)
by: Jeevan, Pranav, et al.
Published: (2022)
CerberusDet: Unified Multi-Dataset Object Detection
by: Tolstykh, Irina, et al.
Published: (2024)
by: Tolstykh, Irina, et al.
Published: (2024)
Optimal Blackjack Strategy Recommender: A Comprehensive Study on Computer Vision Integration for Enhanced Gameplay
by: Gupta, Krishnanshu, et al.
Published: (2024)
by: Gupta, Krishnanshu, et al.
Published: (2024)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
by: Perera, Amal S., et al.
Published: (2025)
by: Perera, Amal S., et al.
Published: (2025)
Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition
by: Carreira, Luiz G F, et al.
Published: (2026)
by: Carreira, Luiz G F, et al.
Published: (2026)
Low-Cost Tree Crown Dieback Estimation Using Deep Learning-Based Segmentation
by: Allen, M. J., et al.
Published: (2024)
by: Allen, M. J., et al.
Published: (2024)
Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis
by: Heyne, Catyana, et al.
Published: (2026)
by: Heyne, Catyana, et al.
Published: (2026)
Event Detection via Probability Density Function Regression
by: Peng, Clark, et al.
Published: (2024)
by: Peng, Clark, et al.
Published: (2024)
Fusion and Grouping Strategies in Deep Learning for Local Climate Zone Classification of Multimodal Remote Sensing Data
by: Thomas, Ancymol, et al.
Published: (2026)
by: Thomas, Ancymol, et al.
Published: (2026)
Transfer learning with generative models for object detection on limited datasets
by: Paiano, Matteo, et al.
Published: (2024)
by: Paiano, Matteo, et al.
Published: (2024)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
by: Diller, Christian, et al.
Published: (2023)
by: Diller, Christian, et al.
Published: (2023)
Similar Items
-
Event-based Solutions for Human-centered Applications: A Comprehensive Review
by: Adra, Mira, et al.
Published: (2025) -
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
by: Bouzid, Hamza, et al.
Published: (2023) -
A Review of Pseudo-Labeling for Computer Vision
by: Kage, Patrick, et al.
Published: (2024) -
Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification
by: Ma, Shuxian, et al.
Published: (2025) -
Efficient Neural Network Encoding for 3D Color Lookup Tables
by: Zehtab, Vahid, et al.
Published: (2024)