Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rykov, Elisei, Petrushina, Kseniia, Titova, Kseniia, Razzhigaev, Anton, Panchenko, Alexander, Konovalov, Vasily |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
SmurfCat at SemEval-2024 Task 6: Leveraging Synthetic Data for Hallucination Detection
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
Multimodal Evaluation of Russian-language Architectures
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
S3: A Simple Strong Sample-effective Multimodal Dialog System
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
Common Sense Reasoning for Deepfake Detection
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
Beyond Detection: Rethinking Education in the Age of AI-writing
von: Marina, Maria, et al.
Veröffentlicht: (2026)
von: Marina, Maria, et al.
Veröffentlicht: (2026)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2026)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2026)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models
von: Conwell, Colin, et al.
Veröffentlicht: (2024)
von: Conwell, Colin, et al.
Veröffentlicht: (2024)
Real-World Transferable Adversarial Attack on Face-Recognition Systems
von: Kaznacheev, Andrey, et al.
Veröffentlicht: (2025)
von: Kaznacheev, Andrey, et al.
Veröffentlicht: (2025)
Sparse and Transferable Universal Singular Vectors Attack
von: Kuvshinova, Kseniia, et al.
Veröffentlicht: (2024)
von: Kuvshinova, Kseniia, et al.
Veröffentlicht: (2024)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
When to Think and When to Look: Uncertainty-Guided Lookback
von: Bi, Jing, et al.
Veröffentlicht: (2025)
von: Bi, Jing, et al.
Veröffentlicht: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
Towards Automatic Evaluation for Image Transcreation
von: Khanuja, Simran, et al.
Veröffentlicht: (2024)
von: Khanuja, Simran, et al.
Veröffentlicht: (2024)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
von: Liang, Yunhao, et al.
Veröffentlicht: (2026)
von: Liang, Yunhao, et al.
Veröffentlicht: (2026)
New Capability to Look Up an ASL Sign from a Video Example
von: Neidle, Carol, et al.
Veröffentlicht: (2024)
von: Neidle, Carol, et al.
Veröffentlicht: (2024)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
von: Jian, Pu, et al.
Veröffentlicht: (2025)
von: Jian, Pu, et al.
Veröffentlicht: (2025)
Where do Large Vision-Language Models Look at when Answering Questions?
von: Xing, Xiaoying, et al.
Veröffentlicht: (2025)
von: Xing, Xiaoying, et al.
Veröffentlicht: (2025)
HSEmotion Team at ABAW-10 Competition: Facial Expression Recognition, Valence-Arousal Estimation, Action Unit Detection and Fine-Grained Violence Classification
von: Savchenko, Andrey V., et al.
Veröffentlicht: (2026)
von: Savchenko, Andrey V., et al.
Veröffentlicht: (2026)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025)
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
A Unified Agentic Framework for Evaluating Conditional Image Generation
von: Wang, Jifang, et al.
Veröffentlicht: (2025)
von: Wang, Jifang, et al.
Veröffentlicht: (2025)
LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
von: Wan, Zhongwei, et al.
Veröffentlicht: (2024)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2024)
Mars Spectrometry 2: Gas Chromatography -- Second place solution
von: Konovalov, Dmitry A.
Veröffentlicht: (2024)
von: Konovalov, Dmitry A.
Veröffentlicht: (2024)
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
von: Long, Fuchen, et al.
Veröffentlicht: (2024)
von: Long, Fuchen, et al.
Veröffentlicht: (2024)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
von: Zinkovich, Viktoriia, et al.
Veröffentlicht: (2025)
von: Zinkovich, Viktoriia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts
von: Rykov, Elisei, et al.
Veröffentlicht: (2025) -
SmurfCat at SemEval-2024 Task 6: Leveraging Synthetic Data for Hallucination Detection
von: Rykov, Elisei, et al.
Veröffentlicht: (2024) -
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
von: Rykov, Elisei, et al.
Veröffentlicht: (2025) -
Multimodal Evaluation of Russian-language Architectures
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025) -
S3: A Simple Strong Sample-effective Multimodal Dialog System
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)