ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
Fuente:
arXiv
Salvato in:
| Autori principali: | Shrivastava, Ayush, Gangani, Kirtan, Jain, Laksh, Goel, Mayank, Batra, Nipun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scalable Methods for Brick Kiln Detection and Compliance Monitoring from Satellite Imagery: A Deployment Case Study in India
di: Mondal, Rishabh, et al.
Pubblicazione: (2024)
di: Mondal, Rishabh, et al.
Pubblicazione: (2024)
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
di: Simonyan, Aleksandr, et al.
Pubblicazione: (2026)
di: Simonyan, Aleksandr, et al.
Pubblicazione: (2026)
Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery
di: Chen, Hao, et al.
Pubblicazione: (2026)
di: Chen, Hao, et al.
Pubblicazione: (2026)
Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations
di: Singh, Jaisidh, et al.
Pubblicazione: (2024)
di: Singh, Jaisidh, et al.
Pubblicazione: (2024)
Eye in the Sky: Detection and Compliance Monitoring of Brick Kilns using Satellite Imagery
di: Mondal, Rishabh, et al.
Pubblicazione: (2024)
di: Mondal, Rishabh, et al.
Pubblicazione: (2024)
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
di: Ma, Qiwei, et al.
Pubblicazione: (2025)
di: Ma, Qiwei, et al.
Pubblicazione: (2025)
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
di: Roberts, Josselin Somerville, et al.
Pubblicazione: (2024)
di: Roberts, Josselin Somerville, et al.
Pubblicazione: (2024)
Benchmarking Vision Language Models for Cultural Understanding
di: Nayak, Shravan, et al.
Pubblicazione: (2024)
di: Nayak, Shravan, et al.
Pubblicazione: (2024)
MultiMedEval: A Benchmark and a Toolkit for Evaluating Medical Vision-Language Models
di: Royer, Corentin, et al.
Pubblicazione: (2024)
di: Royer, Corentin, et al.
Pubblicazione: (2024)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
di: Kamath, Amita, et al.
Pubblicazione: (2025)
di: Kamath, Amita, et al.
Pubblicazione: (2025)
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
di: Zhang, Yiman, et al.
Pubblicazione: (2025)
di: Zhang, Yiman, et al.
Pubblicazione: (2025)
Deep Reinforcement Learning for Urban Air Quality Management: Multi-Objective Optimization of Pollution Mitigation Booth Placement in Metropolitan Environments
di: Rajesh, Kirtan, et al.
Pubblicazione: (2025)
di: Rajesh, Kirtan, et al.
Pubblicazione: (2025)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
di: Yan, Bei, et al.
Pubblicazione: (2024)
di: Yan, Bei, et al.
Pubblicazione: (2024)
Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models
di: Wei, Jie, et al.
Pubblicazione: (2025)
di: Wei, Jie, et al.
Pubblicazione: (2025)
Unifying 2D and 3D Vision-Language Understanding
di: Jain, Ayush, et al.
Pubblicazione: (2025)
di: Jain, Ayush, et al.
Pubblicazione: (2025)
TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
di: Roy, Susim, et al.
Pubblicazione: (2025)
di: Roy, Susim, et al.
Pubblicazione: (2025)
Self-Supervised Spatial Correspondence Across Modalities
di: Shrivastava, Ayush, et al.
Pubblicazione: (2025)
di: Shrivastava, Ayush, et al.
Pubblicazione: (2025)
Self-Supervised Any-Point Tracking by Contrastive Random Walks
di: Shrivastava, Ayush, et al.
Pubblicazione: (2024)
di: Shrivastava, Ayush, et al.
Pubblicazione: (2024)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
di: Wang, Sibo, et al.
Pubblicazione: (2024)
di: Wang, Sibo, et al.
Pubblicazione: (2024)
When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset
di: Das, Sarmistha, et al.
Pubblicazione: (2025)
di: Das, Sarmistha, et al.
Pubblicazione: (2025)
Landsat30-AU: A Vision-Language Dataset for Australian Landsat Imagery
di: Ma, Sai, et al.
Pubblicazione: (2025)
di: Ma, Sai, et al.
Pubblicazione: (2025)
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
di: Shah, Arya, et al.
Pubblicazione: (2026)
di: Shah, Arya, et al.
Pubblicazione: (2026)
Prompt Triage: Structured Optimization Enhances Vision-Language Model Performance on Medical Imaging Benchmarks
di: Singhvi, Arnav, et al.
Pubblicazione: (2025)
di: Singhvi, Arnav, et al.
Pubblicazione: (2025)
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models
di: Roger, Alexis, et al.
Pubblicazione: (2025)
di: Roger, Alexis, et al.
Pubblicazione: (2025)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
di: Unmesh, Asim, et al.
Pubblicazione: (2026)
di: Unmesh, Asim, et al.
Pubblicazione: (2026)
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
di: Rawte, Vipula, et al.
Pubblicazione: (2024)
di: Rawte, Vipula, et al.
Pubblicazione: (2024)
VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
di: Kang, Hyeonsu, et al.
Pubblicazione: (2025)
di: Kang, Hyeonsu, et al.
Pubblicazione: (2025)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
di: Xu, Juangui, et al.
Pubblicazione: (2025)
di: Xu, Juangui, et al.
Pubblicazione: (2025)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
di: Yan, Bei, et al.
Pubblicazione: (2024)
di: Yan, Bei, et al.
Pubblicazione: (2024)
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
di: Chen, Qian, et al.
Pubblicazione: (2026)
di: Chen, Qian, et al.
Pubblicazione: (2026)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
di: Wang, Yi Ru, et al.
Pubblicazione: (2025)
di: Wang, Yi Ru, et al.
Pubblicazione: (2025)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
di: Lee, Youngwan, et al.
Pubblicazione: (2025)
di: Lee, Youngwan, et al.
Pubblicazione: (2025)
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
di: Wu, Xue, et al.
Pubblicazione: (2026)
di: Wu, Xue, et al.
Pubblicazione: (2026)
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
di: Perez, Alejandra, et al.
Pubblicazione: (2026)
di: Perez, Alejandra, et al.
Pubblicazione: (2026)
EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation
di: Han, Shuhao, et al.
Pubblicazione: (2024)
di: Han, Shuhao, et al.
Pubblicazione: (2024)
DiffuSyn Bench: Evaluating Vision-Language Models on Real-World Complexities with Diffusion-Generated Synthetic Benchmarks
di: Zhou, Haokun, et al.
Pubblicazione: (2024)
di: Zhou, Haokun, et al.
Pubblicazione: (2024)
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation
di: Xun, Yue, et al.
Pubblicazione: (2026)
di: Xun, Yue, et al.
Pubblicazione: (2026)
Discovering Failure Modes in Vision-Language Models using RL
di: Jain, Kanishk, et al.
Pubblicazione: (2026)
di: Jain, Kanishk, et al.
Pubblicazione: (2026)
Uncertainty-Aware Evaluation for Vision-Language Models
di: Kostumov, Vasily, et al.
Pubblicazione: (2024)
di: Kostumov, Vasily, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Scalable Methods for Brick Kiln Detection and Compliance Monitoring from Satellite Imagery: A Deployment Case Study in India
di: Mondal, Rishabh, et al.
Pubblicazione: (2024) -
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
di: Simonyan, Aleksandr, et al.
Pubblicazione: (2026) -
Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery
di: Chen, Hao, et al.
Pubblicazione: (2026) -
Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations
di: Singh, Jaisidh, et al.
Pubblicazione: (2024) -
Eye in the Sky: Detection and Compliance Monitoring of Brick Kilns using Satellite Imagery
di: Mondal, Rishabh, et al.
Pubblicazione: (2024)