African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Geigle, Gregor, Timofte, Radu, Glavaš, Goran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
von: Geigle, Gregor, et al.
Veröffentlicht: (2025)
von: Geigle, Gregor, et al.
Veröffentlicht: (2025)
InstructIR: High-Quality Image Restoration Following Human Instructions
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models
von: Wei, Canshi
Veröffentlicht: (2024)
von: Wei, Canshi
Veröffentlicht: (2024)
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024)
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
von: Cui, Chenhang, et al.
Veröffentlicht: (2024)
von: Cui, Chenhang, et al.
Veröffentlicht: (2024)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
von: Oh, Hongseok, et al.
Veröffentlicht: (2025)
von: Oh, Hongseok, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Image Classification of Marine Mammals
von: Qi, Yijiashun, et al.
Veröffentlicht: (2024)
von: Qi, Yijiashun, et al.
Veröffentlicht: (2024)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
von: He, Xingwei, et al.
Veröffentlicht: (2024)
von: He, Xingwei, et al.
Veröffentlicht: (2024)
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
von: Kim, Dain, et al.
Veröffentlicht: (2026)
von: Kim, Dain, et al.
Veröffentlicht: (2026)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
Beyond Accuracy Optimization: Computer Vision Losses for Large Language Model Fine-Tuning
von: Cambrin, Daniele Rege, et al.
Veröffentlicht: (2024)
von: Cambrin, Daniele Rege, et al.
Veröffentlicht: (2024)
From Code to Prediction: Fine-Tuning LLMs for Neural Network Performance Classification in NNGPT
von: Hanouneh, Mahmoud, et al.
Veröffentlicht: (2026)
von: Hanouneh, Mahmoud, et al.
Veröffentlicht: (2026)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
von: Chen, Junzhe, et al.
Veröffentlicht: (2024)
von: Chen, Junzhe, et al.
Veröffentlicht: (2024)
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
von: Jesani, Krunal, et al.
Veröffentlicht: (2025)
von: Jesani, Krunal, et al.
Veröffentlicht: (2025)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
von: Ma, Ruiqi, et al.
Veröffentlicht: (2025)
von: Ma, Ruiqi, et al.
Veröffentlicht: (2025)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
von: Ray, Sourjyadip, et al.
Veröffentlicht: (2024)
von: Ray, Sourjyadip, et al.
Veröffentlicht: (2024)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
von: Liao, Yuan-Hong, et al.
Veröffentlicht: (2024)
von: Liao, Yuan-Hong, et al.
Veröffentlicht: (2024)
Benchmarking Deflection and Hallucination in Large Vision-Language Models
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2026)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2026)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
von: Gao, Jianjun, et al.
Veröffentlicht: (2024)
von: Gao, Jianjun, et al.
Veröffentlicht: (2024)
Practical Manipulation Model for Robust Deepfake Detection
von: Hopf, Benedikt, et al.
Veröffentlicht: (2025)
von: Hopf, Benedikt, et al.
Veröffentlicht: (2025)
Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models
von: Uzun, Tolgay Atinc, et al.
Veröffentlicht: (2026)
von: Uzun, Tolgay Atinc, et al.
Veröffentlicht: (2026)
Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization
von: Song, Yuhang, et al.
Veröffentlicht: (2024)
von: Song, Yuhang, et al.
Veröffentlicht: (2024)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
von: Xiao, Wenyi, et al.
Veröffentlicht: (2024)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2024)
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
von: Lovenia, Holy, et al.
Veröffentlicht: (2023)
von: Lovenia, Holy, et al.
Veröffentlicht: (2023)
Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis
von: Mittal, Yash, et al.
Veröffentlicht: (2025)
von: Mittal, Yash, et al.
Veröffentlicht: (2025)
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification
von: Qu, Renyi, et al.
Veröffentlicht: (2024)
von: Qu, Renyi, et al.
Veröffentlicht: (2024)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
von: Lawrence, Logan, et al.
Veröffentlicht: (2025)
von: Lawrence, Logan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
von: Geigle, Gregor, et al.
Veröffentlicht: (2024) -
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
von: Geigle, Gregor, et al.
Veröffentlicht: (2023) -
mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs
von: Geigle, Gregor, et al.
Veröffentlicht: (2023) -
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
von: Geigle, Gregor, et al.
Veröffentlicht: (2025) -
InstructIR: High-Quality Image Restoration Following Human Instructions
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)