Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Jie, Su, Yiyang, Liu, Xiaoming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpretable Perception and Reasoning for Audiovisual Geolocation
by: Su, Yiyang, et al.
Published: (2026)
by: Su, Yiyang, et al.
Published: (2026)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
by: Lee, Naeun, et al.
Published: (2026)
by: Lee, Naeun, et al.
Published: (2026)
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
by: Qiang, Chenhui, et al.
Published: (2025)
by: Qiang, Chenhui, et al.
Published: (2025)
FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition
by: Zhu, Jie, et al.
Published: (2025)
by: Zhu, Jie, et al.
Published: (2025)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
by: Wang, Haicheng, et al.
Published: (2026)
by: Wang, Haicheng, et al.
Published: (2026)
When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
by: Guo, Longteng, et al.
Published: (2026)
by: Guo, Longteng, et al.
Published: (2026)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
Visual Jigsaw Post-Training Improves MLLMs
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
by: Yin, Zhihan, et al.
Published: (2026)
by: Yin, Zhihan, et al.
Published: (2026)
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
Language-guided Hierarchical Fine-grained Image Forgery Detection and Localization
by: Guo, Xiao, et al.
Published: (2024)
by: Guo, Xiao, et al.
Published: (2024)
HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
by: Su, Yiyang, et al.
Published: (2025)
by: Su, Yiyang, et al.
Published: (2025)
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
by: Xiao, Rui, et al.
Published: (2026)
by: Xiao, Rui, et al.
Published: (2026)
LocalScore: Local Density-Aware Similarity Scoring for Biometrics
by: Su, Yiyang, et al.
Published: (2026)
by: Su, Yiyang, et al.
Published: (2026)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
by: Lin, Haoqiang, et al.
Published: (2025)
by: Lin, Haoqiang, et al.
Published: (2025)
DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
by: Kumar, Raja, et al.
Published: (2025)
by: Kumar, Raja, et al.
Published: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
by: Zhang, Bob, et al.
Published: (2025)
by: Zhang, Bob, et al.
Published: (2025)
UB-FineNet: Urban Building Fine-grained Classification Network for Open-access Satellite Images
by: He, Zhiyi, et al.
Published: (2024)
by: He, Zhiyi, et al.
Published: (2024)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
by: Chang, Boyu, et al.
Published: (2026)
by: Chang, Boyu, et al.
Published: (2026)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
by: Lai, Yingxin, et al.
Published: (2026)
by: Lai, Yingxin, et al.
Published: (2026)
SARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition
by: Yang, Jingxiao, et al.
Published: (2026)
by: Yang, Jingxiao, et al.
Published: (2026)
KeyPoint Relative Position Encoding for Face Recognition
by: Kim, Minchul, et al.
Published: (2024)
by: Kim, Minchul, et al.
Published: (2024)
SapiensID: Foundation for Human Recognition
by: Kim, Minchul, et al.
Published: (2025)
by: Kim, Minchul, et al.
Published: (2025)
Open-Set Biometrics: Beyond Good Closed-Set Models
by: Su, Yiyang, et al.
Published: (2024)
by: Su, Yiyang, et al.
Published: (2024)
SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
by: Peng, Yibo, et al.
Published: (2026)
by: Peng, Yibo, et al.
Published: (2026)
Motion Generation from Fine-grained Textual Descriptions
by: Li, Kunhang, et al.
Published: (2024)
by: Li, Kunhang, et al.
Published: (2024)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
by: Hao, Yunzhuo, et al.
Published: (2025)
by: Hao, Yunzhuo, et al.
Published: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
Democratizing Fine-grained Visual Recognition with Large Language Models
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
by: Yu, Tianyu, et al.
Published: (2023)
by: Yu, Tianyu, et al.
Published: (2023)
Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios
by: Li, Xiaomin, et al.
Published: (2026)
by: Li, Xiaomin, et al.
Published: (2026)
Similar Items
-
Interpretable Perception and Reasoning for Audiovisual Geolocation
by: Su, Yiyang, et al.
Published: (2026) -
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025) -
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025) -
Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
by: Lee, Naeun, et al.
Published: (2026) -
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
by: Qiang, Chenhui, et al.
Published: (2025)