When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Zirui, Tan, Haosheng, Pu, Yuhan, Deng, Zhijie, Shen, Zhouan, Hu, Keyu, Wei, Jiaheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Label Smoothing Improves Gradient Ascent in LLM Unlearning
by: Pang, Zirui, et al.
Published: (2025)
by: Pang, Zirui, et al.
Published: (2025)
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
LM-mixup: Text Data Augmentation via Language Model based Mixup
by: Deng, Zhijie, et al.
Published: (2025)
by: Deng, Zhijie, et al.
Published: (2025)
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
by: Deng, Zhijie, et al.
Published: (2025)
by: Deng, Zhijie, et al.
Published: (2025)
When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs
by: Dai, Aobotao, et al.
Published: (2025)
by: Dai, Aobotao, et al.
Published: (2025)
When Reasoning Meets Its Laws
by: Zhang, Junyu, et al.
Published: (2025)
by: Zhang, Junyu, et al.
Published: (2025)
When Segmentation Meets Hyperspectral Image: New Paradigm for Hyperspectral Image Classification
by: Zhou, Weilian, et al.
Published: (2025)
by: Zhou, Weilian, et al.
Published: (2025)
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
RW-Net: Enhancing Few-Shot Point Cloud Classification with a Wavelet Transform Projection-based Network
by: Zhang, Haosheng, et al.
Published: (2025)
by: Zhang, Haosheng, et al.
Published: (2025)
Rethinking VLMs and LLMs for Image Classification
by: Cooper, Avi, et al.
Published: (2024)
by: Cooper, Avi, et al.
Published: (2024)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
by: An, Keyu, et al.
Published: (2024)
by: An, Keyu, et al.
Published: (2024)
CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator
by: Pu, Yuhan, et al.
Published: (2026)
by: Pu, Yuhan, et al.
Published: (2026)
Text-Region Matching for Multi-Label Image Recognition with Missing Labels
by: Ma, Leilei, et al.
Published: (2024)
by: Ma, Leilei, et al.
Published: (2024)
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
by: Zhuo, Le, et al.
Published: (2025)
by: Zhuo, Le, et al.
Published: (2025)
Bayesian Exploration of Pre-trained Models for Low-shot Image Classification
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
When MoE Meets Blockchain: A Trustworthy Distributed Framework of Large Models
by: Zhu, Weihao, et al.
Published: (2025)
by: Zhu, Weihao, et al.
Published: (2025)
CLAReSNet: When Convolution Meets Latent Attention for Hyperspectral Image Classification
by: Bandyopadhyay, Asmit, et al.
Published: (2025)
by: Bandyopadhyay, Asmit, et al.
Published: (2025)
Reliable Representation Learning for Incomplete Multi-View Missing Multi-Label Classification
by: Liu, Chengliang, et al.
Published: (2023)
by: Liu, Chengliang, et al.
Published: (2023)
Hierarchical Multi-Label Classification with Missing Information for Benthic Habitat Imagery
by: Xu, Isaac, et al.
Published: (2024)
by: Xu, Isaac, et al.
Published: (2024)
When Large Vision-Language Models Meet Person Re-Identification
by: Wang, Qizao, et al.
Published: (2024)
by: Wang, Qizao, et al.
Published: (2024)
On VLMs for Diverse Tasks in Multimodal Meme Classification
by: Gavit, Deepesh, et al.
Published: (2025)
by: Gavit, Deepesh, et al.
Published: (2025)
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
by: Tian, Yuan, et al.
Published: (2026)
by: Tian, Yuan, et al.
Published: (2026)
MBMamba: When Memory Buffer Meets Mamba for Structure-Aware Image Deblurring
by: Gao, Hu, et al.
Published: (2025)
by: Gao, Hu, et al.
Published: (2025)
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2025)
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2025)
When the Small-Loss Trick is Not Enough: Multi-Label Image Classification with Noisy Labels Applied to CCTV Sewer Inspections
by: Chelouche, Keryan, et al.
Published: (2024)
by: Chelouche, Keryan, et al.
Published: (2024)
ProtoGuard-guided PROPEL: Class-Aware Prototype Enhancement and Progressive Labeling for Incremental 3D Point Cloud Segmentation
by: Li, Haosheng, et al.
Published: (2025)
by: Li, Haosheng, et al.
Published: (2025)
When Text Embedding Meets Large Language Model: A Comprehensive Survey
by: Nie, Zhijie, et al.
Published: (2024)
by: Nie, Zhijie, et al.
Published: (2024)
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
by: Gao, Hongcheng, et al.
Published: (2024)
by: Gao, Hongcheng, et al.
Published: (2024)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
Revisiting Early-Learning Regularization When Federated Learning Meets Noisy Labels
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
When Invariant Representation Learning Meets Label Shift: Insufficiency and Theoretical Insights
by: Luo, You-Wei, et al.
Published: (2024)
by: Luo, You-Wei, et al.
Published: (2024)
Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
by: Luo, Yifu, et al.
Published: (2025)
by: Luo, Yifu, et al.
Published: (2025)
When Test-Time Adaptation Meets Self-Supervised Models
by: Han, Jisu, et al.
Published: (2025)
by: Han, Jisu, et al.
Published: (2025)
When Person Re-Identification Meets Event Camera: A Benchmark Dataset and An Attribute-guided Re-Identification Framework
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
QuizRank: Picking Images by Quizzing VLMs
by: Ji, Tenghao, et al.
Published: (2025)
by: Ji, Tenghao, et al.
Published: (2025)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
Label Dependencies-aware Set Prediction Networks for Multi-label Text Classification
by: Xinkai, Du, et al.
Published: (2023)
by: Xinkai, Du, et al.
Published: (2023)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
by: Huang, Hailang, et al.
Published: (2024)
by: Huang, Hailang, et al.
Published: (2024)
Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
by: Xu, Mingjie, et al.
Published: (2025)
by: Xu, Mingjie, et al.
Published: (2025)
Similar Items
-
Label Smoothing Improves Gradient Ascent in LLM Unlearning
by: Pang, Zirui, et al.
Published: (2025) -
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
by: Zheng, Hao, et al.
Published: (2025) -
LM-mixup: Text Data Augmentation via Language Model based Mixup
by: Deng, Zhijie, et al.
Published: (2025) -
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
by: Deng, Zhijie, et al.
Published: (2025) -
When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs
by: Dai, Aobotao, et al.
Published: (2025)