Understanding attention-based encoder-decoder networks: a case study with chess scoresheet recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Hayashi, Sergio Y., Hirata, Nina S. T. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Limits of Spatial Imagery Reasoning in Frontier LLM Models
by: Hayashi, Sergio Y., et al.
Published: (2026)
by: Hayashi, Sergio Y., et al.
Published: (2026)
GCAM: Gaussian and causal-attention model of food fine-grained recognition
by: Zhuang, Guohang, et al.
Published: (2024)
by: Zhuang, Guohang, et al.
Published: (2024)
HCR-Net: A deep learning based script independent handwritten character recognition network
by: Chauhan, Vinod Kumar, et al.
Published: (2021)
by: Chauhan, Vinod Kumar, et al.
Published: (2021)
Sign language recognition from skeletal data using graph and recurrent neural networks
by: Mederos, B., et al.
Published: (2025)
by: Mederos, B., et al.
Published: (2025)
CDAN: Convolutional dense attention-guided network for low-light image enhancement
by: Shakibania, Hossein, et al.
Published: (2023)
by: Shakibania, Hossein, et al.
Published: (2023)
QUEST: A robust attention formulation using query-modulated spherical attention
by: Govindarajan, Hariprasath, et al.
Published: (2026)
by: Govindarajan, Hariprasath, et al.
Published: (2026)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
Pig behavior dataset and Spatial-temporal perception and enhancement networks based on the attention mechanism for pig behavior recognition
by: Qi, Fangzheng, et al.
Published: (2025)
by: Qi, Fangzheng, et al.
Published: (2025)
Developing emotion recognition for video conference software to support people with autism
by: Franzen, Marc, et al.
Published: (2021)
by: Franzen, Marc, et al.
Published: (2021)
Demographic-aware fine-grained visual recognition of pediatric wrist pathologies
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
Sensing technologies and machine learning methods for emotion recognition in autism: Systematic review
by: Banos, Oresti, et al.
Published: (2024)
by: Banos, Oresti, et al.
Published: (2024)
Moving object detection from multi-depth images with an attention-enhanced CNN
by: Shibukawa, Masato, et al.
Published: (2025)
by: Shibukawa, Masato, et al.
Published: (2025)
Vision Transformer attention alignment with human visual perception in aesthetic object evaluation
by: Carrasco, Miguel, et al.
Published: (2025)
by: Carrasco, Miguel, et al.
Published: (2025)
CALICO: Confident Active Learning with Integrated Calibration
by: Querol, Lorenzo S., et al.
Published: (2024)
by: Querol, Lorenzo S., et al.
Published: (2024)
Kinematic analysis of structural mechanics based on convolutional neural network
by: Zhang, Leye, et al.
Published: (2024)
by: Zhang, Leye, et al.
Published: (2024)
OPFormer: Object Pose Estimation leveraging foundation model with geometric encoding
by: Moroz, Artem, et al.
Published: (2025)
by: Moroz, Artem, et al.
Published: (2025)
Multi-scale Quaternion CNN and BiGRU with Cross Self-attention Feature Fusion for Fault Diagnosis of Bearing
by: Liu, Huanbai, et al.
Published: (2024)
by: Liu, Huanbai, et al.
Published: (2024)
MinkUNeXt-SI: Improving point cloud-based place recognition including spherical coordinates and LiDAR intensity
by: Vilella-Cantos, Judith, et al.
Published: (2025)
by: Vilella-Cantos, Judith, et al.
Published: (2025)
Classification for everyone : Building geography agnostic models for fairer recognition
by: Jindal, Akshat, et al.
Published: (2023)
by: Jindal, Akshat, et al.
Published: (2023)
A General Framework for Robust G-Invariance in G-Equivariant Networks
by: Sanborn, Sophia, et al.
Published: (2023)
by: Sanborn, Sophia, et al.
Published: (2023)
Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
by: Fan, Hehe, et al.
Published: (2025)
by: Fan, Hehe, et al.
Published: (2025)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents
by: Singh, Kunal, et al.
Published: (2025)
by: Singh, Kunal, et al.
Published: (2025)
An approach based on class activation maps for investigating the effects of data augmentation on neural networks for image classification
by: Dorneles, Lucas M., et al.
Published: (2025)
by: Dorneles, Lucas M., et al.
Published: (2025)
Intelligent recognition of GPR road hidden defect images based on feature fusion and attention mechanism
by: Lv, Haotian, et al.
Published: (2025)
by: Lv, Haotian, et al.
Published: (2025)
Patronus: Interpretable Diffusion Models with Prototypes
by: Weng, Nina, et al.
Published: (2025)
by: Weng, Nina, et al.
Published: (2025)
COOkeD: Ensemble-based OOD detection in the era of zero-shot CLIP
by: Humblot-Renaux, Galadrielle, et al.
Published: (2025)
by: Humblot-Renaux, Galadrielle, et al.
Published: (2025)
CLIP Can Understand Depth
by: Kim, Sohee, et al.
Published: (2024)
by: Kim, Sohee, et al.
Published: (2024)
Understanding Multi-View Transformers
by: Stary, Michal, et al.
Published: (2025)
by: Stary, Michal, et al.
Published: (2025)
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
by: Nguyen, Le Thien Phuc, et al.
Published: (2025)
by: Nguyen, Le Thien Phuc, et al.
Published: (2025)
Contrasting local and global modeling with machine learning and satellite data: A case study estimating tree canopy height in African savannas
by: Rolf, Esther, et al.
Published: (2024)
by: Rolf, Esther, et al.
Published: (2024)
On the Difficulty of Learning a Meta-network for Training Data Selection
by: Du, Zilin, et al.
Published: (2026)
by: Du, Zilin, et al.
Published: (2026)
MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for Mamba
by: Yoshimura, Masakazu, et al.
Published: (2024)
by: Yoshimura, Masakazu, et al.
Published: (2024)
Do Language Models Understand Time?
by: Ding, Xi, et al.
Published: (2024)
by: Ding, Xi, et al.
Published: (2024)
Understanding Visual Concepts Across Models
by: Trabucco, Brandon, et al.
Published: (2024)
by: Trabucco, Brandon, et al.
Published: (2024)
Understanding Implosion in Text-to-Image Generative Models
by: Ding, Wenxin, et al.
Published: (2024)
by: Ding, Wenxin, et al.
Published: (2024)
Understanding, Accelerating, and Improving MeanFlow Training
by: Kim, Jin-Young, et al.
Published: (2025)
by: Kim, Jin-Young, et al.
Published: (2025)
Adaptive Keyframe Sampling for Long Video Understanding
by: Tang, Xi, et al.
Published: (2025)
by: Tang, Xi, et al.
Published: (2025)
Dual Diffusion for Unified Image Generation and Understanding
by: Li, Zijie, et al.
Published: (2024)
by: Li, Zijie, et al.
Published: (2024)
Similar Items
-
Limits of Spatial Imagery Reasoning in Frontier LLM Models
by: Hayashi, Sergio Y., et al.
Published: (2026) -
GCAM: Gaussian and causal-attention model of food fine-grained recognition
by: Zhuang, Guohang, et al.
Published: (2024) -
HCR-Net: A deep learning based script independent handwritten character recognition network
by: Chauhan, Vinod Kumar, et al.
Published: (2021) -
Sign language recognition from skeletal data using graph and recurrent neural networks
by: Mederos, B., et al.
Published: (2025) -
CDAN: Convolutional dense attention-guided network for low-light image enhancement
by: Shakibania, Hossein, et al.
Published: (2023)