Evaluating Attribute Comprehension in Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Haiwen, Yang, Zixi, Liu, Yuanzhi, Wang, Xinran, He, Zheqi, Liang, Kongming, Ma, Zhanyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024)
by: Wang, Xinran, et al.
Published: (2024)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
Detailed Object Description with Controllable Dimensions
by: Wang, Xinran, et al.
Published: (2024)
by: Wang, Xinran, et al.
Published: (2024)
Benchmarking Segmentation Models with Mask-Preserved Attribute Editing
by: Yin, Zijin, et al.
Published: (2024)
by: Yin, Zijin, et al.
Published: (2024)
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
by: Yin, Zijin, et al.
Published: (2026)
by: Yin, Zijin, et al.
Published: (2026)
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
by: Yang, Zixi, et al.
Published: (2025)
by: Yang, Zixi, et al.
Published: (2025)
Efficient Face Super-Resolution via Wavelet-based Feature Enhancement Network
by: Li, Wenjie, et al.
Published: (2024)
by: Li, Wenjie, et al.
Published: (2024)
Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
by: Li, Baoteng, et al.
Published: (2026)
by: Li, Baoteng, et al.
Published: (2026)
ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
by: Gao, Jiayi, et al.
Published: (2025)
by: Gao, Jiayi, et al.
Published: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
by: Wang, Xinran, et al.
Published: (2026)
by: Wang, Xinran, et al.
Published: (2026)
Polyp-E: Benchmarking the Robustness of Deep Segmentation Models via Polyp Editing
by: Wei, Runpu, et al.
Published: (2024)
by: Wei, Runpu, et al.
Published: (2024)
PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Image Segmentation
by: Yan, Zhonghao, et al.
Published: (2025)
by: Yan, Zhonghao, et al.
Published: (2025)
MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation
by: Yang, Ling, et al.
Published: (2023)
by: Yang, Ling, et al.
Published: (2023)
REVAL: A Comprehension Evaluation on Reliability and Values of Large Vision-Language Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Comprehensive Attribution: Inherently Explainable Vision Model with Feature Detector
by: Zhang, Xianren, et al.
Published: (2024)
by: Zhang, Xianren, et al.
Published: (2024)
Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images
by: Yang, Yuxuan, et al.
Published: (2026)
by: Yang, Yuxuan, et al.
Published: (2026)
Multi-modal Attribute Prompting for Vision-Language Models
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Toward Generalizable Forgery Detection and Reasoning
by: Gao, Yueying, et al.
Published: (2025)
by: Gao, Yueying, et al.
Published: (2025)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
by: Wu, Xuyang, et al.
Published: (2024)
by: Wu, Xuyang, et al.
Published: (2024)
Resampling Benchmark for Efficient Comprehensive Evaluation of Large Vision-Language Models
by: Suzuki, Teppei, et al.
Published: (2025)
by: Suzuki, Teppei, et al.
Published: (2025)
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
by: Ying, Kaining, et al.
Published: (2024)
by: Ying, Kaining, et al.
Published: (2024)
A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
by: Wu, Tianhe, et al.
Published: (2024)
by: Wu, Tianhe, et al.
Published: (2024)
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
by: Wei, Runpu, et al.
Published: (2025)
by: Wei, Runpu, et al.
Published: (2025)
PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models
by: Li, Yuliang, et al.
Published: (2026)
by: Li, Yuliang, et al.
Published: (2026)
IncreFA: Breaking the Static Wall of Generative Model Attribution
by: Qin, Haotian, et al.
Published: (2026)
by: Qin, Haotian, et al.
Published: (2026)
Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers
by: Zhang, Shuo, et al.
Published: (2026)
by: Zhang, Shuo, et al.
Published: (2026)
Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
by: Chen, Qizhou, et al.
Published: (2024)
by: Chen, Qizhou, et al.
Published: (2024)
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
by: Yan, Zhonghao, et al.
Published: (2025)
by: Yan, Zhonghao, et al.
Published: (2025)
Training A Small Emotional Vision Language Model for Visual Art Comprehension
by: Zhang, Jing, et al.
Published: (2024)
by: Zhang, Jing, et al.
Published: (2024)
Unveiling Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
Generative Visual Chain-of-Thought for Image Editing
by: Yin, Zijin, et al.
Published: (2026)
by: Yin, Zijin, et al.
Published: (2026)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
by: Fu, Chaoyou, et al.
Published: (2023)
by: Fu, Chaoyou, et al.
Published: (2023)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
by: Lu, Fan, et al.
Published: (2024)
by: Lu, Fan, et al.
Published: (2024)
ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2023)
by: Tian, Xinyu, et al.
Published: (2023)
SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models
by: Yue, Tongtian, et al.
Published: (2024)
by: Yue, Tongtian, et al.
Published: (2024)
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
by: Wang, Zekun, et al.
Published: (2025)
by: Wang, Zekun, et al.
Published: (2025)
Similar Items
-
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024) -
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025) -
Detailed Object Description with Controllable Dimensions
by: Wang, Xinran, et al.
Published: (2024) -
Benchmarking Segmentation Models with Mask-Preserved Attribute Editing
by: Yin, Zijin, et al.
Published: (2024) -
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
by: Yin, Zijin, et al.
Published: (2026)