[CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Akang, Deng, Xili, Hu, Zhanxuan, Zhao, Yi, Tai, Yonghang, Li, Huafeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Models
by: Hu, Zhanxuan, et al.
Published: (2025)
by: Hu, Zhanxuan, et al.
Published: (2025)
ConInfer: Context-Aware Inference for Training-Free Open-Vocabulary Remote Sensing Segmentation
by: Chen, Wenyang, et al.
Published: (2026)
by: Chen, Wenyang, et al.
Published: (2026)
A Hidden Stumbling Block in Generalized Category Discovery: Distracted Attention
by: Xu, Qiyu, et al.
Published: (2025)
by: Xu, Qiyu, et al.
Published: (2025)
Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification
by: Zhou, Linhan, et al.
Published: (2025)
by: Zhou, Linhan, et al.
Published: (2025)
Ensembling Diffusion Models via Adaptive Feature Aggregation
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
by: Marouani, Alexis, et al.
Published: (2026)
by: Marouani, Alexis, et al.
Published: (2026)
Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
by: Nguyen, Y Hop, et al.
Published: (2025)
by: Nguyen, Y Hop, et al.
Published: (2025)
TiCLS : Tightly Coupled Language Text Spotter
by: Jang, Leeje, et al.
Published: (2026)
by: Jang, Leeje, et al.
Published: (2026)
GaitASMS: Gait Recognition by Adaptive Structured Spatial Representation and Multi-Scale Temporal Aggregation
by: Sun, Yan, et al.
Published: (2023)
by: Sun, Yan, et al.
Published: (2023)
Multi-Expert Adaptive Selection: Task-Balancing for All-in-One Image Restoration
by: Yu, Xiaoyan, et al.
Published: (2024)
by: Yu, Xiaoyan, et al.
Published: (2024)
Text-Region Matching for Multi-Label Image Recognition with Missing Labels
by: Ma, Leilei, et al.
Published: (2024)
by: Ma, Leilei, et al.
Published: (2024)
Incomplete Multi-Label Image Recognition by Co-learning Semantic-Aware Features and Label Recovery
by: He, Zhi-Fen, et al.
Published: (2025)
by: He, Zhi-Fen, et al.
Published: (2025)
Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies
by: Gorostegui, Juan Ignacio Bustos, et al.
Published: (2026)
by: Gorostegui, Juan Ignacio Bustos, et al.
Published: (2026)
An Optimized PatchMatch for Multi-scale and Multi-feature Label Fusion
by: Giraud, Rémi, et al.
Published: (2019)
by: Giraud, Rémi, et al.
Published: (2019)
DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
by: Liu, Haijing, et al.
Published: (2025)
by: Liu, Haijing, et al.
Published: (2025)
Good Enough: Is it Worth Improving your Label Quality?
by: Jaus, Alexander, et al.
Published: (2025)
by: Jaus, Alexander, et al.
Published: (2025)
Pedestrian Attribute Recognition as Label-balanced Multi-label Learning
by: Zhou, Yibo, et al.
Published: (2024)
by: Zhou, Yibo, et al.
Published: (2024)
Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models
by: Kong, Dehong, et al.
Published: (2024)
by: Kong, Dehong, et al.
Published: (2024)
Unsqueeze [CLS] Bottleneck to Learn Rich Representations
by: Su, Qing, et al.
Published: (2024)
by: Su, Qing, et al.
Published: (2024)
Unlocking [CLS] Features for Continual Post-Training
by: Yildirim, Murat Onur, et al.
Published: (2025)
by: Yildirim, Murat Onur, et al.
Published: (2025)
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
MAPLE: Multi-Path Adaptive Propagation with Level-Aware Embeddings for Hierarchical Multi-Label Image Classification
by: Koloski, Boshko, et al.
Published: (2026)
by: Koloski, Boshko, et al.
Published: (2026)
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition
by: Liu, Haijing, et al.
Published: (2024)
by: Liu, Haijing, et al.
Published: (2024)
Counterfactual Reasoning for Multi-Label Image Classification via Patching-Based Training
by: Xie, Ming-Kun, et al.
Published: (2024)
by: Xie, Ming-Kun, et al.
Published: (2024)
Adaptive Dynamic Dehazing via Instruction-Driven and Task-Feedback Closed-Loop Optimization for Diverse Downstream Task Adaptation
by: Zhang, Yafei, et al.
Published: (2026)
by: Zhang, Yafei, et al.
Published: (2026)
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
by: Miller, Kevin, et al.
Published: (2025)
by: Miller, Kevin, et al.
Published: (2025)
Synthetic Data Alone is Enough? Rethinking Data Scarcity in Pediatric Rare Disease Recognition
by: Feng, Ganlin, et al.
Published: (2026)
by: Feng, Ganlin, et al.
Published: (2026)
Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusion
by: Yang, Zengyi, et al.
Published: (2026)
by: Yang, Zengyi, et al.
Published: (2026)
EPIR: An Efficient Patch Tokenization, Integration and Representation Framework for Micro-expression Recognition
by: Wang, Junbo, et al.
Published: (2026)
by: Wang, Junbo, et al.
Published: (2026)
Learning Semantic-Aware Threshold for Multi-Label Image Recognition with Partial Labels
by: Ruan, Haoxian, et al.
Published: (2025)
by: Ruan, Haoxian, et al.
Published: (2025)
Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
by: Rasekh, Ali, et al.
Published: (2025)
by: Rasekh, Ali, et al.
Published: (2025)
AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception
by: Li, Xilai, et al.
Published: (2025)
by: Li, Xilai, et al.
Published: (2025)
Hiding Your Signals: A Security Analysis of PPG-based Biometric Authentication
by: Li, Lin, et al.
Published: (2022)
by: Li, Lin, et al.
Published: (2022)
CASP: Few-Shot Class-Incremental Learning with CLS Token Attention Steering Prompts
by: Huang, Shuai, et al.
Published: (2026)
by: Huang, Shuai, et al.
Published: (2026)
Query-Based Adaptive Aggregation for Multi-Dataset Joint Training Toward Universal Visual Place Recognition
by: Xiao, Jiuhong, et al.
Published: (2025)
by: Xiao, Jiuhong, et al.
Published: (2025)
PEVA-Net: Prompt-Enhanced View Aggregation Network for Zero/Few-Shot Multi-View 3D Shape Recognition
by: Lin, Dongyun, et al.
Published: (2024)
by: Lin, Dongyun, et al.
Published: (2024)
When the Small-Loss Trick is Not Enough: Multi-Label Image Classification with Noisy Labels Applied to CCTV Sewer Inspections
by: Chelouche, Keryan, et al.
Published: (2024)
by: Chelouche, Keryan, et al.
Published: (2024)
SSPA: Split-and-Synthesize Prompting with Gated Alignments for Multi-Label Image Recognition
by: Tan, Hao, et al.
Published: (2024)
by: Tan, Hao, et al.
Published: (2024)
How Much Data are Enough? Investigating Dataset Requirements for Patch-Based Brain MRI Segmentation Tasks
by: Wang, Dongang, et al.
Published: (2024)
by: Wang, Dongang, et al.
Published: (2024)
Similar Items
-
SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Models
by: Hu, Zhanxuan, et al.
Published: (2025) -
ConInfer: Context-Aware Inference for Training-Free Open-Vocabulary Remote Sensing Segmentation
by: Chen, Wenyang, et al.
Published: (2026) -
A Hidden Stumbling Block in Generalized Category Discovery: Distracted Attention
by: Xu, Qiyu, et al.
Published: (2025) -
Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification
by: Zhou, Linhan, et al.
Published: (2025) -
Ensembling Diffusion Models via Adaptive Feature Aggregation
by: Wang, Cong, et al.
Published: (2024)