Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Toibazar, Daulet, Wang, Kesen, Mohamed, Sherif, Al-Badawi, Abdulaziz, Alfulayt, Abdulrahman, Moreno, Pedro J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
von: Al-Homoud, Haneen, et al.
Veröffentlicht: (2025)
von: Al-Homoud, Haneen, et al.
Veröffentlicht: (2025)
A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation
von: Wang, Kesen, et al.
Veröffentlicht: (2025)
von: Wang, Kesen, et al.
Veröffentlicht: (2025)
Multi-Agent Interactive Question Generation Framework for Long Document Understanding
von: Wang, Kesen, et al.
Veröffentlicht: (2025)
von: Wang, Kesen, et al.
Veröffentlicht: (2025)
FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models
von: Sahili, Zahraa Al, et al.
Veröffentlicht: (2025)
von: Sahili, Zahraa Al, et al.
Veröffentlicht: (2025)
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
von: Narayanan, Aravind, et al.
Veröffentlicht: (2025)
von: Narayanan, Aravind, et al.
Veröffentlicht: (2025)
Clapper: Compact Learning and Video Representation in VLMs
von: Kong, Lingyu, et al.
Veröffentlicht: (2025)
von: Kong, Lingyu, et al.
Veröffentlicht: (2025)
Should VLMs be Pre-trained with Image Data?
von: Keh, Sedrick, et al.
Veröffentlicht: (2025)
von: Keh, Sedrick, et al.
Veröffentlicht: (2025)
Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers
von: Eltahir, Mohamed, et al.
Veröffentlicht: (2025)
von: Eltahir, Mohamed, et al.
Veröffentlicht: (2025)
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
von: Luo, Zhi, et al.
Veröffentlicht: (2025)
von: Luo, Zhi, et al.
Veröffentlicht: (2025)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
von: Hayes, Kevin David, et al.
Veröffentlicht: (2025)
von: Hayes, Kevin David, et al.
Veröffentlicht: (2025)
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
von: Poggi, Nicolas, et al.
Veröffentlicht: (2025)
von: Poggi, Nicolas, et al.
Veröffentlicht: (2025)
Building Trust in Virtual Immunohistochemistry: Automated Assessment of Image Quality
von: Kataria, Tushar, et al.
Veröffentlicht: (2025)
von: Kataria, Tushar, et al.
Veröffentlicht: (2025)
Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions
von: Saha, Oindrila, et al.
Veröffentlicht: (2024)
von: Saha, Oindrila, et al.
Veröffentlicht: (2024)
Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought
von: Zhao, Kesen, et al.
Veröffentlicht: (2026)
von: Zhao, Kesen, et al.
Veröffentlicht: (2026)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
Image Recognition with Vision and Language Embeddings of VLMs
von: Volkov, Illia, et al.
Veröffentlicht: (2025)
von: Volkov, Illia, et al.
Veröffentlicht: (2025)
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
von: Wang, Weizhi, et al.
Veröffentlicht: (2024)
von: Wang, Weizhi, et al.
Veröffentlicht: (2024)
RSDiff: Remote Sensing Image Generation from Text Using Diffusion Model
von: Sebaq, Ahmad, et al.
Veröffentlicht: (2023)
von: Sebaq, Ahmad, et al.
Veröffentlicht: (2023)
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
T2T-VICL: Unlocking the Boundaries of Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
von: Xia, Shao-Jun, et al.
Veröffentlicht: (2025)
von: Xia, Shao-Jun, et al.
Veröffentlicht: (2025)
CopyJudge: Automated Copyright Infringement Identification and Mitigation in Text-to-Image Diffusion Models
von: Liu, Shunchang, et al.
Veröffentlicht: (2025)
von: Liu, Shunchang, et al.
Veröffentlicht: (2025)
Data Factory with Minimal Human Effort Using VLMs
von: Ye, Jiaojiao, et al.
Veröffentlicht: (2025)
von: Ye, Jiaojiao, et al.
Veröffentlicht: (2025)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models
von: Wang, Zhendong, et al.
Veröffentlicht: (2025)
von: Wang, Zhendong, et al.
Veröffentlicht: (2025)
Learning to Customize Text-to-Image Diffusion In Diverse Context
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
CoRe: Context-Regularized Text Embedding Learning for Text-to-Image Personalization
von: Wu, Feize, et al.
Veröffentlicht: (2024)
von: Wu, Feize, et al.
Veröffentlicht: (2024)
TF-TI2I: Training-Free Text-and-Image-to-Image Generation via Multi-Modal Implicit-Context Learning in Text-to-Image Models
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2025)
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2025)
Beyond Static Perception: Integrating Temporal Context into VLMs for Cloth Folding
von: Barbany, Oriol, et al.
Veröffentlicht: (2025)
von: Barbany, Oriol, et al.
Veröffentlicht: (2025)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
Benchmarking Compact VLMs for Clip-Level Surveillance Anomaly Detection Under Weak Supervision
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
TCIG: Two-Stage Controlled Image Generation with Quality Enhancement through Diffusion
von: Mohamed, Salaheldin
Veröffentlicht: (2024)
von: Mohamed, Salaheldin
Veröffentlicht: (2024)
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
von: Jadhav, Avadhoot, et al.
Veröffentlicht: (2025)
von: Jadhav, Avadhoot, et al.
Veröffentlicht: (2025)
Quality and Quantity: Unveiling a Million High-Quality Images for Text-to-Image Synthesis in Fashion Design
von: Yu, Jia, et al.
Veröffentlicht: (2023)
von: Yu, Jia, et al.
Veröffentlicht: (2023)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
von: Zhang, Yabo, et al.
Veröffentlicht: (2024)
von: Zhang, Yabo, et al.
Veröffentlicht: (2024)
DragNeXt: Rethinking Drag-Based Image Editing
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
von: Malakouti, Sina, et al.
Veröffentlicht: (2024)
von: Malakouti, Sina, et al.
Veröffentlicht: (2024)
Lean Unet: A Compact Model for Image Segmentation
von: Hassler, Ture, et al.
Veröffentlicht: (2025)
von: Hassler, Ture, et al.
Veröffentlicht: (2025)
Robustness and Transferability of Pix2Geomodel for Bidirectional Facies Property Translation in a Complex Reservoir
von: Al-Fakih, Abdulrahman, et al.
Veröffentlicht: (2026)
von: Al-Fakih, Abdulrahman, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
von: Al-Homoud, Haneen, et al.
Veröffentlicht: (2025) -
A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation
von: Wang, Kesen, et al.
Veröffentlicht: (2025) -
Multi-Agent Interactive Question Generation Framework for Long Document Understanding
von: Wang, Kesen, et al.
Veröffentlicht: (2025) -
FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models
von: Sahili, Zahraa Al, et al.
Veröffentlicht: (2025) -
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
von: Narayanan, Aravind, et al.
Veröffentlicht: (2025)