What's in a Name? Beyond Class Indices for Image Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Kai, Huang, Xiaohu, Li, Yandong, Vaze, Sagar, Li, Jie, Jia, Xuhui |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dissecting Out-of-Distribution Detection and Open-Set Recognition: A Critical Analysis of Methods and Benchmarks
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
HiLo: A Learning Framework for Generalized Category Discovery Robust to Domain Shifts
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
by: Huang, Xiaohu, et al.
Published: (2024)
by: Huang, Xiaohu, et al.
Published: (2024)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
by: Huang, Xiaohu, et al.
Published: (2024)
by: Huang, Xiaohu, et al.
Published: (2024)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
by: Li, Jinyuan, et al.
Published: (2024)
by: Li, Jinyuan, et al.
Published: (2024)
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language Explainer
by: Lei, Jiaming, et al.
Published: (2024)
by: Lei, Jiaming, et al.
Published: (2024)
CXR-AD: Component X-ray Image Dataset for Industrial Anomaly Detection
by: Bai, Haoyu, et al.
Published: (2025)
by: Bai, Haoyu, et al.
Published: (2025)
In-Context Translation: Towards Unifying Image Recognition, Processing, and Generation
by: Xue, Han, et al.
Published: (2024)
by: Xue, Han, et al.
Published: (2024)
Long-Tailed Anomaly Detection with Learnable Class Names
by: Ho, Chih-Hui, et al.
Published: (2024)
by: Ho, Chih-Hui, et al.
Published: (2024)
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
Absolute-Unified Multi-Class Anomaly Detection via Class-Agnostic Distribution Alignment
by: Guo, Jia, et al.
Published: (2024)
by: Guo, Jia, et al.
Published: (2024)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
by: Li, Jinyuan, et al.
Published: (2024)
by: Li, Jinyuan, et al.
Published: (2024)
A Proposal-Free Query-Guided Network for Grounded Multimodal Named Entity Recognition
by: Li, Hongbing, et al.
Published: (2026)
by: Li, Hongbing, et al.
Published: (2026)
NamedCurves: Learned Image Enhancement via Color Naming
by: Serrano-Lozano, David, et al.
Published: (2024)
by: Serrano-Lozano, David, et al.
Published: (2024)
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
by: Ma, Nanye, et al.
Published: (2025)
by: Ma, Nanye, et al.
Published: (2025)
Video Creation by Demonstration
by: Sun, Yihong, et al.
Published: (2024)
by: Sun, Yihong, et al.
Published: (2024)
Distribution-aware Interactive Attention Network and Large-scale Cloud Recognition Benchmark on FY-4A Satellite Image
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
Beyond Motion Cues and Structural Sparsity: Revisiting Small Moving Target Detection
by: Zhang, Guoyi, et al.
Published: (2025)
by: Zhang, Guoyi, et al.
Published: (2025)
A Forward and Backward Compatible Framework for Few-shot Class-incremental Pill Recognition
by: Zhang, Jinghua, et al.
Published: (2023)
by: Zhang, Jinghua, et al.
Published: (2023)
Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognition
by: Liu, Hui, et al.
Published: (2026)
by: Liu, Hui, et al.
Published: (2026)
Deep Tiny Network for Recognition-Oriented Face Image Quality Assessment
by: Peng, Baoyun, et al.
Published: (2021)
by: Peng, Baoyun, et al.
Published: (2021)
Video-based Sign Language Recognition without Temporal Segmentation
by: Huang, Jie, et al.
Published: (2018)
by: Huang, Jie, et al.
Published: (2018)
Beyond Image Super-Resolution for Image Recognition with Task-Driven Perceptual Loss
by: Kim, Jaeha, et al.
Published: (2024)
by: Kim, Jaeha, et al.
Published: (2024)
Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision
by: Jing, Jie, et al.
Published: (2024)
by: Jing, Jie, et al.
Published: (2024)
SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting
by: Wu, Hefeng, et al.
Published: (2023)
by: Wu, Hefeng, et al.
Published: (2023)
Data-Free Class-Incremental Gesture Recognition with Prototype-Guided Pseudo Feature Replay
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
The Comparison of Individual Cat Recognition Using Neural Networks
by: Li, Mingxuan, et al.
Published: (2024)
by: Li, Mingxuan, et al.
Published: (2024)
Diffusion Model Meets Non-Exemplar Class-Incremental Learning and Beyond
by: Zhang, Jichuan, et al.
Published: (2024)
by: Zhang, Jichuan, et al.
Published: (2024)
Information Bottleneck-based Causal Attention for Multi-label Medical Image Recognition
by: Cui, Xiaoxiao, et al.
Published: (2025)
by: Cui, Xiaoxiao, et al.
Published: (2025)
RegionDrag: Fast Region-Based Image Editing with Diffusion Models
by: Lu, Jingyi, et al.
Published: (2024)
by: Lu, Jingyi, et al.
Published: (2024)
Improving Interpretability of Deep Active Learning for Flood Inundation Mapping Through Class Ambiguity Indices Using Multi-spectral Satellite Imagery
by: Lee, Hyunho, et al.
Published: (2024)
by: Lee, Hyunho, et al.
Published: (2024)
GRITv2: Efficient and Light-weight Social Relation Recognition
by: Reddy, N K Sagar, et al.
Published: (2024)
by: Reddy, N K Sagar, et al.
Published: (2024)
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
by: Li, Jiaxuan, et al.
Published: (2023)
by: Li, Jiaxuan, et al.
Published: (2023)
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
by: Chan, Kelvin C. K., et al.
Published: (2024)
by: Chan, Kelvin C. K., et al.
Published: (2024)
Classes Are Not Equal: An Empirical Study on Image Recognition Fairness
by: Cui, Jiequan, et al.
Published: (2024)
by: Cui, Jiequan, et al.
Published: (2024)
Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer
by: Liu, Wenxuan, et al.
Published: (2025)
by: Liu, Wenxuan, et al.
Published: (2025)
What Is Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution
by: Ye, Xingsong, et al.
Published: (2026)
by: Ye, Xingsong, et al.
Published: (2026)
From Static to Dynamic: Adapting Landmark-Aware Image Models for Facial Expression Recognition in Videos
by: Chen, Yin, et al.
Published: (2023)
by: Chen, Yin, et al.
Published: (2023)
Similar Items
-
Dissecting Out-of-Distribution Detection and Open-Set Recognition: A Critical Analysis of Methods and Benchmarks
by: Wang, Hongjun, et al.
Published: (2024) -
HiLo: A Learning Framework for Generalized Category Discovery Robust to Domain Shifts
by: Wang, Hongjun, et al.
Published: (2024) -
SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt Tuning
by: Wang, Hongjun, et al.
Published: (2024) -
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
by: Huang, Xiaohu, et al.
Published: (2024) -
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
by: Huang, Xiaohu, et al.
Published: (2024)