A Conformal Risk Control Framework for Granular Word Assessment and Uncertainty Calibration of CLIPScore Quality Estimates
Fuente:
arXiv
Saved in:
| Main Authors: | Gomes, Gonçalo, Martins, Bruno, Zerva, Chrysoula |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
by: Gomes, Gonçalo, et al.
Published: (2025)
by: Gomes, Gonçalo, et al.
Published: (2025)
Non-Exchangeable Conformal Language Generation with Nearest Neighbors
by: Ulmer, Dennis, et al.
Published: (2024)
by: Ulmer, Dennis, et al.
Published: (2024)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
by: Padhi, Trilok, et al.
Published: (2025)
by: Padhi, Trilok, et al.
Published: (2025)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
by: Xie, Yupeng, et al.
Published: (2025)
by: Xie, Yupeng, et al.
Published: (2025)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
by: Leng, Jixuan, et al.
Published: (2025)
by: Leng, Jixuan, et al.
Published: (2025)
Decoupling Perception and Calibration: Label-Efficient Image Quality Assessment Framework
by: Li, Xinyue, et al.
Published: (2026)
by: Li, Xinyue, et al.
Published: (2026)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
by: Lan, Zhibin, et al.
Published: (2024)
by: Lan, Zhibin, et al.
Published: (2024)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
by: Freitas, Miguel Monte e, et al.
Published: (2026)
by: Freitas, Miguel Monte e, et al.
Published: (2026)
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
by: Xiao, Wenyi, et al.
Published: (2026)
by: Xiao, Wenyi, et al.
Published: (2026)
RingGesture: A Ring-Based Mid-Air Gesture Typing System Powered by a Deep-Learning Word Prediction Framework
by: Shen, Junxiao, et al.
Published: (2024)
by: Shen, Junxiao, et al.
Published: (2024)
Beyond Words: Multimodal LLM Knows When to Speak
by: Liao, Zikai, et al.
Published: (2025)
by: Liao, Zikai, et al.
Published: (2025)
Designing Practical Models for Isolated Word Visual Speech Recognition
by: Panagos, Iason Ioannis, et al.
Published: (2025)
by: Panagos, Iason Ioannis, et al.
Published: (2025)
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
by: Zhou, Ziwei, et al.
Published: (2026)
by: Zhou, Ziwei, et al.
Published: (2026)
DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios
by: Artham, Sainithin, et al.
Published: (2026)
by: Artham, Sainithin, et al.
Published: (2026)
Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models
by: Chen, Zijun, et al.
Published: (2024)
by: Chen, Zijun, et al.
Published: (2024)
Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
by: Kriz, Anita, et al.
Published: (2025)
by: Kriz, Anita, et al.
Published: (2025)
Using Images to Find Context-Independent Word Representations in Vector Space
by: Kumar, Harsh
Published: (2024)
by: Kumar, Harsh
Published: (2024)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
by: Guo, Ziyu, et al.
Published: (2026)
by: Guo, Ziyu, et al.
Published: (2026)
Efficient Architectures for High Resolution Vision-Language Models
by: Carvalho, Miguel, et al.
Published: (2025)
by: Carvalho, Miguel, et al.
Published: (2025)
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
by: Broomfield, Julius, et al.
Published: (2025)
by: Broomfield, Julius, et al.
Published: (2025)
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
by: Huang, Jinsheng, et al.
Published: (2024)
by: Huang, Jinsheng, et al.
Published: (2024)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
by: He, Zoe Wanying, et al.
Published: (2025)
by: He, Zoe Wanying, et al.
Published: (2025)
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
by: Jung, Jaeyoon, et al.
Published: (2026)
by: Jung, Jaeyoon, et al.
Published: (2026)
Accurate and Well-Calibrated ICD Code Assignment Through Attention Over Diverse Label Embeddings
by: Gomes, Gonçalo, et al.
Published: (2024)
by: Gomes, Gonçalo, et al.
Published: (2024)
A Theoretical and Practical Framework for Evaluating Uncertainty Calibration in Object Detection
by: Conde, Pedro, et al.
Published: (2023)
by: Conde, Pedro, et al.
Published: (2023)
SpecPL: Disentangling Spectral Granularity for Prompt Learning
by: Zhou, Jingtao, et al.
Published: (2026)
by: Zhou, Jingtao, et al.
Published: (2026)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
by: Kim, Mingyeong, et al.
Published: (2026)
by: Kim, Mingyeong, et al.
Published: (2026)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
by: Carvalho, Miguel, et al.
Published: (2025)
by: Carvalho, Miguel, et al.
Published: (2025)
SLVideo: A Sign Language Video Moment Retrieval Framework
by: Martins, Gonçalo Vinagre, et al.
Published: (2024)
by: Martins, Gonçalo Vinagre, et al.
Published: (2024)
CSLRConformer: A Data-Centric Conformer Approach for Continuous Arabic Sign Language Recognition on the Isharah Datase
by: Elden, Fatimah Mohamed Emad
Published: (2025)
by: Elden, Fatimah Mohamed Emad
Published: (2025)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
by: Yuan, Fan, et al.
Published: (2025)
by: Yuan, Fan, et al.
Published: (2025)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
by: Tang, Fei, et al.
Published: (2026)
by: Tang, Fei, et al.
Published: (2026)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
by: Park, Seongheon, et al.
Published: (2026)
by: Park, Seongheon, et al.
Published: (2026)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
by: Oh, Hongseok, et al.
Published: (2025)
by: Oh, Hongseok, et al.
Published: (2025)
How Confident are Video Models? Empowering Video Models to Express their Uncertainty
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
by: Seo, Hoigi, et al.
Published: (2025)
by: Seo, Hoigi, et al.
Published: (2025)
SAUGE: Taming SAM for Uncertainty-Aligned Multi-Granularity Edge Detection
by: Liufu, Xing, et al.
Published: (2024)
by: Liufu, Xing, et al.
Published: (2024)
Scientific Reasoning: Assessment of Multimodal Generative LLMs
by: Dreyer, Florian, et al.
Published: (2025)
by: Dreyer, Florian, et al.
Published: (2025)
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
by: Lab, Shanghai AI, et al.
Published: (2025)
by: Lab, Shanghai AI, et al.
Published: (2025)
Similar Items
-
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
by: Gomes, Gonçalo, et al.
Published: (2025) -
Non-Exchangeable Conformal Language Generation with Nearest Neighbors
by: Ulmer, Dennis, et al.
Published: (2024) -
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
by: Padhi, Trilok, et al.
Published: (2025) -
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
by: Xie, Yupeng, et al.
Published: (2025) -
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
by: Leng, Jixuan, et al.
Published: (2025)