Explainable Image Recognition via Enhanced Slot-attention Based Classifier
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Bowen, Li, Liangzhi, Zhang, Jiahao, Nakashima, Yuta, Nagahara, Hajime |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
E-InMeMo: Enhanced Prompting for Visual In-Context Learning
by: Zhang, Jiahao, et al.
Published: (2025)
by: Zhang, Jiahao, et al.
Published: (2025)
PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
by: Zhang, Jiahao, et al.
Published: (2025)
by: Zhang, Jiahao, et al.
Published: (2025)
Enhancing Ambiguous Dynamic Facial Expression Recognition with Soft Label-based Data Augmentation
by: Kawamura, Ryosuke, et al.
Published: (2025)
by: Kawamura, Ryosuke, et al.
Published: (2025)
MIDAS: Mixing Ambiguous Data with Soft Labels for Dynamic Facial Expression Recognition
by: Kawamura, Ryosuke, et al.
Published: (2025)
by: Kawamura, Ryosuke, et al.
Published: (2025)
Point-Supervised Facial Expression Spotting with Gaussian-Based Instance-Adaptive Intensity Modeling
by: Deng, Yicheng, et al.
Published: (2025)
by: Deng, Yicheng, et al.
Published: (2025)
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training
by: Jiang, Zhouqiang, et al.
Published: (2024)
by: Jiang, Zhouqiang, et al.
Published: (2024)
Deep Polarization Cues for Single-shot Shape and Subsurface Scattering Estimation
by: Li, Chenhao, et al.
Published: (2024)
by: Li, Chenhao, et al.
Published: (2024)
SpotFormer: Multi-Scale Spatio-Temporal Transformer for Facial Expression Spotting
by: Deng, Yicheng, et al.
Published: (2024)
by: Deng, Yicheng, et al.
Published: (2024)
VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction
by: Zhang, Jiahao, et al.
Published: (2024)
by: Zhang, Jiahao, et al.
Published: (2024)
Coded-E2LF: Coded Aperture Light Field Imaging from Events
by: Tsuchida, Tomoya, et al.
Published: (2026)
by: Tsuchida, Tomoya, et al.
Published: (2026)
Multi-Scale Spatio-Temporal Graph Convolutional Network for Facial Expression Spotting
by: Deng, Yicheng, et al.
Published: (2024)
by: Deng, Yicheng, et al.
Published: (2024)
Stable Diffusion Exposed: Gender Bias from Prompt to Image
by: Wu, Yankun, et al.
Published: (2023)
by: Wu, Yankun, et al.
Published: (2023)
CALICO: Confident Active Learning with Integrated Calibration
by: Querol, Lorenzo S., et al.
Published: (2024)
by: Querol, Lorenzo S., et al.
Published: (2024)
NeISF++: Neural Incident Stokes Field for Polarized Inverse Rendering of Conductors and Dielectrics
by: Li, Chenhao, et al.
Published: (2024)
by: Li, Chenhao, et al.
Published: (2024)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
Single-Image Depth from Defocus with Coded Aperture and Diffusion Posterior Sampling
by: Kawachi, Hodaka, et al.
Published: (2025)
by: Kawachi, Hodaka, et al.
Published: (2025)
Explainable Part-Based Vehicle Classifier with Spatial Awareness
by: Caduff, Andreas, et al.
Published: (2026)
by: Caduff, Andreas, et al.
Published: (2026)
Exploring Visual Prompting: Robustness Inheritance and Beyond
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs
by: Pang, Zongshang, et al.
Published: (2025)
by: Pang, Zongshang, et al.
Published: (2025)
EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
by: Wei, Lu, et al.
Published: (2025)
by: Wei, Lu, et al.
Published: (2025)
Privacy in Image Datasets: A Case Study on Pregnancy Ultrasounds
by: Lohanimit, Rawisara, et al.
Published: (2026)
by: Lohanimit, Rawisara, et al.
Published: (2026)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022)
by: Li, Kunchang, et al.
Published: (2022)
Slot-VAE: Object-Centric Scene Generation with Slot Attention
by: Wang, Yanbo, et al.
Published: (2023)
by: Wang, Yanbo, et al.
Published: (2023)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Acquiring a Dynamic Light Field through a Single-Shot Coded Image
by: Mizuno, Ryoya, et al.
Published: (2022)
by: Mizuno, Ryoya, et al.
Published: (2022)
EMWaveNet: Physically Explainable Neural Network Based on Electromagnetic Propagation for SAR Target Recognition
by: Li, Zhuoxuan, et al.
Published: (2024)
by: Li, Zhuoxuan, et al.
Published: (2024)
OpenSlot: Mixed Open-Set Recognition with Object-Centric Learning
by: Yin, Xu, et al.
Published: (2024)
by: Yin, Xu, et al.
Published: (2024)
Predicting Video Slot Attention Queries from Random Slot-Feature Pairs
by: Zhao, Rongzhen, et al.
Published: (2025)
by: Zhao, Rongzhen, et al.
Published: (2025)
Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
by: Li, Jiahao, et al.
Published: (2025)
by: Li, Jiahao, et al.
Published: (2025)
Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
by: Lai, Yixuan, et al.
Published: (2026)
by: Lai, Yixuan, et al.
Published: (2026)
Time-Efficient Light-Field Acquisition Using Coded Aperture and Events
by: Habuchi, Shuji, et al.
Published: (2024)
by: Habuchi, Shuji, et al.
Published: (2024)
Continuous Sign Language Recognition Based on Motor attention mechanism and frame-level Self-distillation
by: Zhu, Qidan, et al.
Published: (2024)
by: Zhu, Qidan, et al.
Published: (2024)
Slot-VLM: SlowFast Slots for Video-Language Modeling
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
When Slots Compete: Slot Merging in Object-Centric Learning
by: Chatzisavvas, Christos, et al.
Published: (2026)
by: Chatzisavvas, Christos, et al.
Published: (2026)
From Global to Local: Social Bias Transfer in CLIP
by: Ramos, Ryan, et al.
Published: (2025)
by: Ramos, Ryan, et al.
Published: (2025)
Towards Counterfactual and Contrastive Explainability and Transparency of DCNN Image Classifiers
by: Tariq, Syed Ali, et al.
Published: (2025)
by: Tariq, Syed Ali, et al.
Published: (2025)
CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance Shifts
by: Dünkel, Olaf, et al.
Published: (2025)
by: Dünkel, Olaf, et al.
Published: (2025)
Enhancing Cross-Dataset Performance of Distracted Driving Detection With Score Softmax Classifier And Dynamic Gaussian Smoothing Supervision
by: Duan, Cong, et al.
Published: (2023)
by: Duan, Cong, et al.
Published: (2023)
Learning Global Object-Centric Representations via Disentangled Slot Attention
by: Chen, Tonglin, et al.
Published: (2024)
by: Chen, Tonglin, et al.
Published: (2024)
Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints
by: Xia, Jiahao, et al.
Published: (2025)
by: Xia, Jiahao, et al.
Published: (2025)
Similar Items
-
E-InMeMo: Enhanced Prompting for Visual In-Context Learning
by: Zhang, Jiahao, et al.
Published: (2025) -
PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
by: Zhang, Jiahao, et al.
Published: (2025) -
Enhancing Ambiguous Dynamic Facial Expression Recognition with Soft Label-based Data Augmentation
by: Kawamura, Ryosuke, et al.
Published: (2025) -
MIDAS: Mixing Ambiguous Data with Soft Labels for Dynamic Facial Expression Recognition
by: Kawamura, Ryosuke, et al.
Published: (2025) -
Point-Supervised Facial Expression Spotting with Gaussian-Based Instance-Adaptive Intensity Modeling
by: Deng, Yicheng, et al.
Published: (2025)