Benchmarking Multimodal Models for Fine-Grained Image Analysis: A Comparative Study Across Diverse Visual Features
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Evstafev, Evgenii |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Fine-Grained Attention and Geometric Correspondence Model for Musculoskeletal Risk Classification in Athletes Using Multimodal Visual and Skeletal Features
von: Rahman, Md. Abdur, et al.
Veröffentlicht: (2025)
von: Rahman, Md. Abdur, et al.
Veröffentlicht: (2025)
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
von: Wang, Bingli, et al.
Veröffentlicht: (2026)
von: Wang, Bingli, et al.
Veröffentlicht: (2026)
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2025)
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2025)
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
von: Xiao, Yicheng, et al.
Veröffentlicht: (2026)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2026)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
XR-VLM: Cross-Relationship Modeling with Multi-part Prompts and Visual Features for Fine-Grained Recognition
von: Wang, Chuanming, et al.
Veröffentlicht: (2025)
von: Wang, Chuanming, et al.
Veröffentlicht: (2025)
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
von: Shan, Haozhe, et al.
Veröffentlicht: (2026)
von: Shan, Haozhe, et al.
Veröffentlicht: (2026)
MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
von: Cao, Yue, et al.
Veröffentlicht: (2024)
von: Cao, Yue, et al.
Veröffentlicht: (2024)
Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks
von: Jin, Jing, et al.
Veröffentlicht: (2026)
von: Jin, Jing, et al.
Veröffentlicht: (2026)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
von: Liu, Runzhou, et al.
Veröffentlicht: (2026)
von: Liu, Runzhou, et al.
Veröffentlicht: (2026)
Concept Drift and Long-Tailed Distribution in Fine-Grained Visual Categorization: Benchmark and Method
von: Ye, Shuo, et al.
Veröffentlicht: (2023)
von: Ye, Shuo, et al.
Veröffentlicht: (2023)
FG$^2$: Fine-Grained Cross-View Localization by Fine-Grained Feature Matching
von: Xia, Zimin, et al.
Veröffentlicht: (2025)
von: Xia, Zimin, et al.
Veröffentlicht: (2025)
Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification
von: Jin, Yinghao, et al.
Veröffentlicht: (2025)
von: Jin, Yinghao, et al.
Veröffentlicht: (2025)
OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models
von: Weng, Tengjin, et al.
Veröffentlicht: (2026)
von: Weng, Tengjin, et al.
Veröffentlicht: (2026)
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
von: Zhang, Jesen, et al.
Veröffentlicht: (2025)
von: Zhang, Jesen, et al.
Veröffentlicht: (2025)
Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
von: Gong, Haozhen, et al.
Veröffentlicht: (2025)
von: Gong, Haozhen, et al.
Veröffentlicht: (2025)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
On Learning Discriminative Features from Synthesized Data for Self-Supervised Fine-Grained Visual Recognition
von: Wang, Zihu, et al.
Veröffentlicht: (2024)
von: Wang, Zihu, et al.
Veröffentlicht: (2024)
Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
SGIA: Enhancing Fine-Grained Visual Classification with Sequence Generative Image Augmentation
von: Liao, Qiyu, et al.
Veröffentlicht: (2024)
von: Liao, Qiyu, et al.
Veröffentlicht: (2024)
MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
von: Xue, Junxiao, et al.
Veröffentlicht: (2026)
von: Xue, Junxiao, et al.
Veröffentlicht: (2026)
Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
von: Pang, Cong, et al.
Veröffentlicht: (2025)
von: Pang, Cong, et al.
Veröffentlicht: (2025)
Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection
von: Xiao, Yao, et al.
Veröffentlicht: (2026)
von: Xiao, Yao, et al.
Veröffentlicht: (2026)
Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification
von: Qiu, Shulei, et al.
Veröffentlicht: (2024)
von: Qiu, Shulei, et al.
Veröffentlicht: (2024)
LoDisc: Learning Global-Local Discriminative Features for Self-Supervised Fine-Grained Visual Recognition
von: Shi, Jialu, et al.
Veröffentlicht: (2024)
von: Shi, Jialu, et al.
Veröffentlicht: (2024)
YourSkatingCoach: A Figure Skating Video Benchmark for Fine-Grained Element Analysis
von: Chen, Wei-Yi, et al.
Veröffentlicht: (2024)
von: Chen, Wei-Yi, et al.
Veröffentlicht: (2024)
Visually-Guided Controllable Medical Image Generation via Fine-Grained Semantic Disentanglement
von: Huang, Xin, et al.
Veröffentlicht: (2026)
von: Huang, Xin, et al.
Veröffentlicht: (2026)
Fine-Grained Zero-Shot Composed Image Retrieval with Complementary Visual-Semantic Integration
von: Ye, Yongcong, et al.
Veröffentlicht: (2026)
von: Ye, Yongcong, et al.
Veröffentlicht: (2026)
Optimizing Domain-Specific Image Retrieval: A Benchmark of FAISS and Annoy with Fine-Tuned Features
von: Rahman, MD Shaikh, et al.
Veröffentlicht: (2024)
von: Rahman, MD Shaikh, et al.
Veröffentlicht: (2024)
ShiftedBronzes: Benchmarking and Analysis of Domain Fine-Grained Classification in Open-World Settings
von: Zhou, Rixin, et al.
Veröffentlicht: (2024)
von: Zhou, Rixin, et al.
Veröffentlicht: (2024)
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
von: Zhu, Fengbin, et al.
Veröffentlicht: (2024)
von: Zhu, Fengbin, et al.
Veröffentlicht: (2024)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control
von: Chen, Shi, et al.
Veröffentlicht: (2026)
von: Chen, Shi, et al.
Veröffentlicht: (2026)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
DCFG: Diverse Cross-Channel Fine-Grained Feature Learning and Progressive Fusion Siamese Tracker for Thermal Infrared Target Tracking
von: Xiong, Ruoyan, et al.
Veröffentlicht: (2025)
von: Xiong, Ruoyan, et al.
Veröffentlicht: (2025)
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
von: Liu, Tianyi, et al.
Veröffentlicht: (2026)
von: Liu, Tianyi, et al.
Veröffentlicht: (2026)
Diffusion Representations for Fine-Grained Image Classification: A Marine Plankton Case Study
von: Juscafresa, A. Nieto, et al.
Veröffentlicht: (2026)
von: Juscafresa, A. Nieto, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Fine-Grained Attention and Geometric Correspondence Model for Musculoskeletal Risk Classification in Athletes Using Multimodal Visual and Skeletal Features
von: Rahman, Md. Abdur, et al.
Veröffentlicht: (2025) -
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
von: Wang, Bingli, et al.
Veröffentlicht: (2026) -
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2025) -
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
von: Xiao, Yicheng, et al.
Veröffentlicht: (2026) -
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)