A Survey on Interpretability in Visual Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Qiyang, Gao, Chengzhi, Wang, Ruiping, Chen, Xilin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisKnow: Constructing Visual Knowledge Base for Object Understanding
by: Yao, Ziwei, et al.
Published: (2025)
by: Yao, Ziwei, et al.
Published: (2025)
Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models
by: Fu, Bin, et al.
Published: (2024)
by: Fu, Bin, et al.
Published: (2024)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
by: Bai, Ziyi, et al.
Published: (2024)
by: Bai, Ziyi, et al.
Published: (2024)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)
by: Xiao, Zhanqi, et al.
Published: (2026)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
GLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition
by: Wang, Tianyue, et al.
Published: (2025)
by: Wang, Tianyue, et al.
Published: (2025)
GEAR: GEometry-motion Alternating Refinement for Articulated Object Modeling with Gaussian Splatting
by: Li, Jialin, et al.
Published: (2026)
by: Li, Jialin, et al.
Published: (2026)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2025)
by: Xie, Senwei, et al.
Published: (2025)
GS-LTS: 3D Gaussian Splatting-Based Adaptive Modeling for Long-Term Service Robots
by: Fu, Bin, et al.
Published: (2025)
by: Fu, Bin, et al.
Published: (2025)
Spatial Information Bottleneck for Interpretable Visual Recognition
by: Shu, Kaixiang, et al.
Published: (2025)
by: Shu, Kaixiang, et al.
Published: (2025)
Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual Representation
by: Xiang, Wenzhao, et al.
Published: (2025)
by: Xiang, Wenzhao, et al.
Published: (2025)
Open Panoramic Segmentation
by: Zheng, Junwei, et al.
Published: (2024)
by: Zheng, Junwei, et al.
Published: (2024)
EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy
by: Ge, Xuanyu, et al.
Published: (2026)
by: Ge, Xuanyu, et al.
Published: (2026)
Scene-agnostic Pose Regression for Visual Localization
by: Zheng, Junwei, et al.
Published: (2025)
by: Zheng, Junwei, et al.
Published: (2025)
M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
A Survey on Visual Mamba
by: Zhang, Hanwei, et al.
Published: (2024)
by: Zhang, Hanwei, et al.
Published: (2024)
MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
Deep Learning for Cross-Domain Few-Shot Visual Recognition: A Survey
by: Xu, Huali, et al.
Published: (2023)
by: Xu, Huali, et al.
Published: (2023)
Deep Correlated Prompting for Visual Recognition with Missing Modalities
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
by: Li, Hang, et al.
Published: (2023)
by: Li, Hang, et al.
Published: (2023)
EventGait: Towards Robust Gait Recognition with Event Streams
by: Xu, Senyan, et al.
Published: (2026)
by: Xu, Senyan, et al.
Published: (2026)
Region Matters: Efficient and Reliable Region-Aware Visual Place Recognition
by: Chen, Shunpeng, et al.
Published: (2026)
by: Chen, Shunpeng, et al.
Published: (2026)
Image Recognition with Online Lightweight Vision Transformer: A Survey
by: Zhang, Zherui, et al.
Published: (2025)
by: Zhang, Zherui, et al.
Published: (2025)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
by: Liu, Ruiping, et al.
Published: (2026)
by: Liu, Ruiping, et al.
Published: (2026)
Visual Mamba: A Survey and New Outlooks
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
NCL++: Nested Collaborative Learning for Long-Tailed Visual Recognition
by: Tan, Zichang, et al.
Published: (2023)
by: Tan, Zichang, et al.
Published: (2023)
SeaFormer++: Squeeze-enhanced Axial Transformer for Mobile Visual Recognition
by: Wan, Qiang, et al.
Published: (2023)
by: Wan, Qiang, et al.
Published: (2023)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
by: Xiang, Wenzhao, et al.
Published: (2026)
by: Xiang, Wenzhao, et al.
Published: (2026)
SciceVPR: Stable Cross-Image Correlation Enhanced Model for Visual Place Recognition
by: Wan, Shanshan, et al.
Published: (2025)
by: Wan, Shanshan, et al.
Published: (2025)
Distributed Zero-Shot Learning for Visual Recognition
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection
by: Hong, Chengzhi, et al.
Published: (2026)
by: Hong, Chengzhi, et al.
Published: (2026)
un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
by: Li, Yinqi, et al.
Published: (2025)
by: Li, Yinqi, et al.
Published: (2025)
Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource
by: Yang, Lei, et al.
Published: (2025)
by: Yang, Lei, et al.
Published: (2025)
Survey on Hand Gesture Recognition from Visual Input
by: Linardakis, Manousos, et al.
Published: (2025)
by: Linardakis, Manousos, et al.
Published: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022)
by: Li, Kunchang, et al.
Published: (2022)
Beyond the Visible: A Survey on Cross-spectral Face Recognition
by: Anghelone, David, et al.
Published: (2022)
by: Anghelone, David, et al.
Published: (2022)
Towards Visual Grounding: A Survey
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
Interpretable Underwater Diver Gesture Recognition
by: Mangalvedhekar, Sudeep, et al.
Published: (2023)
by: Mangalvedhekar, Sudeep, et al.
Published: (2023)
TCFormer: Visual Recognition via Token Clustering Transformer
by: Zeng, Wang, et al.
Published: (2024)
by: Zeng, Wang, et al.
Published: (2024)
PVLR: Prompt-driven Visual-Linguistic Representation Learning for Multi-Label Image Recognition
by: Tan, Hao, et al.
Published: (2024)
by: Tan, Hao, et al.
Published: (2024)
Similar Items
-
VisKnow: Constructing Visual Knowledge Base for Object Understanding
by: Yao, Ziwei, et al.
Published: (2025) -
Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models
by: Fu, Bin, et al.
Published: (2024) -
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
by: Bai, Ziyi, et al.
Published: (2024) -
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026) -
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)