ROI-Aware Multiscale Cross-Attention Vision Transformer for Pest Image Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Ga-Eun, Son, Chang-Hwan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Locally Grouped and Scale-Guided Attention for Dense Pest Counting
by: Son, Chang-Hwan
Published: (2024)
by: Son, Chang-Hwan
Published: (2024)
Degradation-Agnostic Statistical Facial Feature Transformation for Blind Face Restoration in Adverse Weather Conditions
by: Son, Chang-Hwan, et al.
Published: (2025)
by: Son, Chang-Hwan, et al.
Published: (2025)
New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching
by: Son, Chang-Hwan, et al.
Published: (2021)
by: Son, Chang-Hwan, et al.
Published: (2021)
Decision-Aware Attention Propagation for Vision Transformer Explainability
by: Jo, Sehyeong, et al.
Published: (2026)
by: Jo, Sehyeong, et al.
Published: (2026)
Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision
by: Kim, Jinnyeong, et al.
Published: (2024)
by: Kim, Jinnyeong, et al.
Published: (2024)
MedROI: Codec-Agnostic Region of Interest-Centric Compression for Medical Images
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Structural Attention: Rethinking Transformer for Unpaired Medical Image Synthesis
by: Phan, Vu Minh Hieu, et al.
Published: (2024)
by: Phan, Vu Minh Hieu, et al.
Published: (2024)
Lightweight Spatiotemporal Highway Lane Detection via 3D-ResNet and PINet with ROI-Aware Attention
by: Raja, Sorna Shanmuga, et al.
Published: (2026)
by: Raja, Sorna Shanmuga, et al.
Published: (2026)
Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
by: Liang, Wenhao, et al.
Published: (2025)
by: Liang, Wenhao, et al.
Published: (2025)
DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
by: Wu, Yefeng, et al.
Published: (2025)
by: Wu, Yefeng, et al.
Published: (2025)
ROI-Packing: Efficient Region-Based Compression for Machine Vision
by: Eimon, Md Eimran Hossain, et al.
Published: (2025)
by: Eimon, Md Eimran Hossain, et al.
Published: (2025)
PestVL-Net: Enabling Multimodal Pest Learning via Fine-grained Vision-Language Interaction
by: Li, Xueheng, et al.
Published: (2026)
by: Li, Xueheng, et al.
Published: (2026)
Fieldscale: Locality-Aware Field-based Adaptive Rescaling for Thermal Infrared Image
by: Gil, Hyeonjae, et al.
Published: (2024)
by: Gil, Hyeonjae, et al.
Published: (2024)
RNA: Video Editing with ROI-based Neural Atlas
by: Lee, Jaekyeong, et al.
Published: (2024)
by: Lee, Jaekyeong, et al.
Published: (2024)
Intra-task Mutual Attention based Vision Transformer for Few-Shot Learning
by: Jiang, Weihao, et al.
Published: (2024)
by: Jiang, Weihao, et al.
Published: (2024)
Representative Attention For Vision Transformers
by: Li, Yuntong, et al.
Published: (2026)
by: Li, Yuntong, et al.
Published: (2026)
Vision Transformers with Hierarchical Attention
by: Liu, Yun, et al.
Published: (2021)
by: Liu, Yun, et al.
Published: (2021)
Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification
by: Gao, Xinle, et al.
Published: (2025)
by: Gao, Xinle, et al.
Published: (2025)
How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads
by: Baek, Ingeol, et al.
Published: (2025)
by: Baek, Ingeol, et al.
Published: (2025)
SAMReg: SAM-enabled Image Registration with ROI-based Correspondence
by: Huang, Shiqi, et al.
Published: (2024)
by: Huang, Shiqi, et al.
Published: (2024)
Estimating Extreme 3D Image Rotation with Transformer Cross-Attention
by: Dekel, Shay, et al.
Published: (2023)
by: Dekel, Shay, et al.
Published: (2023)
Crafting Query-Aware Selective Attention for Single Image Super-Resolution
by: Kim, Junyoung, et al.
Published: (2025)
by: Kim, Junyoung, et al.
Published: (2025)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
by: Song, Zijie, et al.
Published: (2023)
by: Song, Zijie, et al.
Published: (2023)
Structured Initialization for Attention in Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
Vision Transformers are Circulant Attention Learners
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Multi-manifold Attention for Vision Transformers
by: Konstantinidis, Dimitrios, et al.
Published: (2022)
by: Konstantinidis, Dimitrios, et al.
Published: (2022)
HAViT: Historical Attention Vision Transformer
by: Banik, Swarnendu, et al.
Published: (2026)
by: Banik, Swarnendu, et al.
Published: (2026)
DBAT: Dynamic Backward Attention Transformer for Material Segmentation with Cross-Resolution Patches
by: Heng, Yuwen, et al.
Published: (2023)
by: Heng, Yuwen, et al.
Published: (2023)
Local-Aware Global Attention Network for Person Re-Identification Based on Body and Hand Images
by: Baisa, Nathanael L.
Published: (2022)
by: Baisa, Nathanael L.
Published: (2022)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
MedFormer: Hierarchical Medical Vision Transformer with Content-Aware Dual Sparse Selection Attention
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
Interpretability-Aware Vision Transformer
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
by: Kwak, Min-Seop, et al.
Published: (2025)
by: Kwak, Min-Seop, et al.
Published: (2025)
CSTA: CNN-based Spatiotemporal Attention for Video Summarization
by: Son, Jaewon, et al.
Published: (2024)
by: Son, Jaewon, et al.
Published: (2024)
TERDNet: Transformer Encoder-Recurrent Decoder Network for Scene Change Detection
by: Yoon, Jiae, et al.
Published: (2026)
by: Yoon, Jiae, et al.
Published: (2026)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
A Saccade-inspired Approach to Image Classification using Vision Transformer Attention Maps
by: Dallain, Matthis, et al.
Published: (2026)
by: Dallain, Matthis, et al.
Published: (2026)
Multiscale Vision Transformers meet Bipartite Matching for efficient single-stage Action Localization
by: Ntinou, Ioanna, et al.
Published: (2023)
by: Ntinou, Ioanna, et al.
Published: (2023)
Improved Multiscale Structural Mapping with Supervertex Vision Transformer for the Detection of Alzheimer's Disease Neurodegeneration
by: Baek, Geonwoo, et al.
Published: (2026)
by: Baek, Geonwoo, et al.
Published: (2026)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Similar Items
-
Locally Grouped and Scale-Guided Attention for Dense Pest Counting
by: Son, Chang-Hwan
Published: (2024) -
Degradation-Agnostic Statistical Facial Feature Transformation for Blind Face Restoration in Adverse Weather Conditions
by: Son, Chang-Hwan, et al.
Published: (2025) -
New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching
by: Son, Chang-Hwan, et al.
Published: (2021) -
Decision-Aware Attention Propagation for Vision Transformer Explainability
by: Jo, Sehyeong, et al.
Published: (2026) -
Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision
by: Kim, Jinnyeong, et al.
Published: (2024)