Can Vision-Language Models Replace Human Annotators: A Case Study with CelebA Dataset
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Haoming, Zhong, Feifei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attention IoU: Examining Biases in CelebA using Attention Maps
von: Serianni, Aaron, et al.
Veröffentlicht: (2025)
von: Serianni, Aaron, et al.
Veröffentlicht: (2025)
Beyond Performance Disparities: A Three-Level Audit of Representational Harm in CelebA
von: Park, Sieun, et al.
Veröffentlicht: (2026)
von: Park, Sieun, et al.
Veröffentlicht: (2026)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2026)
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2026)
Fine-tuning Pre-trained Vision-Language Models in a Human-Annotation-Free Manner
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
Pre-Trained Vision-Language Models as Partial Annotators
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2024)
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2024)
OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
von: Zhang, Yulong
Veröffentlicht: (2025)
von: Zhang, Yulong
Veröffentlicht: (2025)
Can Vision-Language Models Understand Construction Workers? An Exploratory Study
von: Bui, Hieu, et al.
Veröffentlicht: (2026)
von: Bui, Hieu, et al.
Veröffentlicht: (2026)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
von: Nandy, Abhilash, et al.
Veröffentlicht: (2024)
von: Nandy, Abhilash, et al.
Veröffentlicht: (2024)
Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
von: Blankemeier, Louis, et al.
Veröffentlicht: (2024)
von: Blankemeier, Louis, et al.
Veröffentlicht: (2024)
Longitudinal Vestibular Schwannoma Dataset with Consensus-based Human-in-the-loop Annotations
von: Wijethilake, Navodini, et al.
Veröffentlicht: (2025)
von: Wijethilake, Navodini, et al.
Veröffentlicht: (2025)
MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?
von: Li, Guanzhen, et al.
Veröffentlicht: (2024)
von: Li, Guanzhen, et al.
Veröffentlicht: (2024)
VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
von: Wu, Kui, et al.
Veröffentlicht: (2025)
von: Wu, Kui, et al.
Veröffentlicht: (2025)
Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis
von: Lu, Sheng, et al.
Veröffentlicht: (2026)
von: Lu, Sheng, et al.
Veröffentlicht: (2026)
Efficient and Comprehensive Feature Extraction in Large Vision-Language Model for Pathology Analysis
von: Zhang, Shengxuming, et al.
Veröffentlicht: (2024)
von: Zhang, Shengxuming, et al.
Veröffentlicht: (2024)
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
von: Fan, Sicheng, et al.
Veröffentlicht: (2026)
von: Fan, Sicheng, et al.
Veröffentlicht: (2026)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
von: Roy, Parthib, et al.
Veröffentlicht: (2024)
von: Roy, Parthib, et al.
Veröffentlicht: (2024)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
Sanitizing Manufacturing Dataset Labels Using Vision-Language Models
von: Mahjourian, Nazanin, et al.
Veröffentlicht: (2025)
von: Mahjourian, Nazanin, et al.
Veröffentlicht: (2025)
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
von: Zhou, Yuhao, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2026)
ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models
von: Zhong, Jing, et al.
Veröffentlicht: (2025)
von: Zhong, Jing, et al.
Veröffentlicht: (2025)
UrbanSense:A Framework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models
von: Yin, Jun, et al.
Veröffentlicht: (2025)
von: Yin, Jun, et al.
Veröffentlicht: (2025)
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
von: Zhang, Yongting, et al.
Veröffentlicht: (2024)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
von: Bao, Han, et al.
Veröffentlicht: (2024)
von: Bao, Han, et al.
Veröffentlicht: (2024)
Can Vision Language Models Understand Mimed Actions?
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
Privacy-Preserving Computer Vision for Industry: Three Case Studies in Human-Centric Manufacturing
von: De Coninck, Sander, et al.
Veröffentlicht: (2025)
von: De Coninck, Sander, et al.
Veröffentlicht: (2025)
Can Machines Imitate Humans? Integrative Turing-like tests for Language and Vision Demonstrate a Narrowing Gap
von: Zhang, Mengmi, et al.
Veröffentlicht: (2022)
von: Zhang, Mengmi, et al.
Veröffentlicht: (2022)
Avoid Wasted Annotation Costs in Open-set Active Learning with Pre-trained Vision-Language Model
von: Heo, Jaehyuk, et al.
Veröffentlicht: (2024)
von: Heo, Jaehyuk, et al.
Veröffentlicht: (2024)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
von: Liu, Jinlong, et al.
Veröffentlicht: (2026)
von: Liu, Jinlong, et al.
Veröffentlicht: (2026)
Conformal Predictions for Human Action Recognition with Vision-Language Models
von: Tim, Bary, et al.
Veröffentlicht: (2025)
von: Tim, Bary, et al.
Veröffentlicht: (2025)
ClimateIQA: A New Dataset and Benchmark to Advance Vision-Language Models in Meteorology Anomalies Analysis
von: Chen, Jian, et al.
Veröffentlicht: (2024)
von: Chen, Jian, et al.
Veröffentlicht: (2024)
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT
von: Du, Chengyi, et al.
Veröffentlicht: (2026)
von: Du, Chengyi, et al.
Veröffentlicht: (2026)
GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
von: Zhang, Kuan, et al.
Veröffentlicht: (2026)
von: Zhang, Kuan, et al.
Veröffentlicht: (2026)
Landsat30-AU: A Vision-Language Dataset for Australian Landsat Imagery
von: Ma, Sai, et al.
Veröffentlicht: (2025)
von: Ma, Sai, et al.
Veröffentlicht: (2025)
Can Vision-Language Models Solve Visual Math Equations?
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
ReasonEdit: Editing Vision-Language Models using Human Reasoning
von: Qiu, Jiaxing, et al.
Veröffentlicht: (2026)
von: Qiu, Jiaxing, et al.
Veröffentlicht: (2026)
Why Do Vision Language Models Struggle To Recognize Human Emotions?
von: Agarwal, Madhav, et al.
Veröffentlicht: (2026)
von: Agarwal, Madhav, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Attention IoU: Examining Biases in CelebA using Attention Maps
von: Serianni, Aaron, et al.
Veröffentlicht: (2025) -
Beyond Performance Disparities: A Three-Level Audit of Representational Harm in CelebA
von: Park, Sieun, et al.
Veröffentlicht: (2026) -
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2026) -
Fine-tuning Pre-trained Vision-Language Models in a Human-Annotation-Free Manner
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026) -
Pre-Trained Vision-Language Models as Partial Annotators
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2024)