Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment
Fuente:
arXiv
Guardado en:
| Autores principales: | Driggers-Ellis, Christopher, Tibrewal, Nachiketh, Bogulla, Rohit, Khanna, Harsh, Youm, Sangpil, Grant, Christan, Dorr, Bonnie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Effect of Data Partitioning Strategy on Model Generalizability: A Case Study of Morphological Segmentation
por: Liu, Zoey, et al.
Publicado: (2024)
por: Liu, Zoey, et al.
Publicado: (2024)
Revisiting Semantic Role Labeling: Efficient Structured Inference with Dependency-Informed Analysis
por: Youm, Sangpil, et al.
Publicado: (2026)
por: Youm, Sangpil, et al.
Publicado: (2026)
AMREx: AMR for Explainable Fact Verification
por: Jayaweera, Chathuri, et al.
Publicado: (2024)
por: Jayaweera, Chathuri, et al.
Publicado: (2024)
DAHRS: Divergence-Aware Hallucination-Remediated SRL Projection
por: Youm, Sangpil, et al.
Publicado: (2024)
por: Youm, Sangpil, et al.
Publicado: (2024)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
por: Cherif, Ahmed
Publicado: (2026)
por: Cherif, Ahmed
Publicado: (2026)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
por: Rathva, Harsh, et al.
Publicado: (2025)
por: Rathva, Harsh, et al.
Publicado: (2025)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
por: Simhi, Adi, et al.
Publicado: (2024)
por: Simhi, Adi, et al.
Publicado: (2024)
BHRAM-IL: A Benchmark for Hallucination Recognition and Assessment in Multiple Indian Languages
por: Terdalkar, Hrishikesh, et al.
Publicado: (2025)
por: Terdalkar, Hrishikesh, et al.
Publicado: (2025)
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
por: Salla, Rohit Kumar, et al.
Publicado: (2025)
por: Salla, Rohit Kumar, et al.
Publicado: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
por: Rudman, William, et al.
Publicado: (2026)
por: Rudman, William, et al.
Publicado: (2026)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
por: Lee, Kyuho, et al.
Publicado: (2025)
por: Lee, Kyuho, et al.
Publicado: (2025)
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
por: Pan, Xinghan
Publicado: (2025)
por: Pan, Xinghan
Publicado: (2025)
A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph Generation
por: Sun, Shuzhou, et al.
Publicado: (2025)
por: Sun, Shuzhou, et al.
Publicado: (2025)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
por: Radosky, Lukas, et al.
Publicado: (2026)
por: Radosky, Lukas, et al.
Publicado: (2026)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
por: Li, Xiaoyang, et al.
Publicado: (2025)
por: Li, Xiaoyang, et al.
Publicado: (2025)
Improving Consistency in Large Language Models through Chain of Guidance
por: Raj, Harsh, et al.
Publicado: (2025)
por: Raj, Harsh, et al.
Publicado: (2025)
M3LEO: A Multi-Modal, Multi-Label Earth Observation Dataset Integrating Interferometric SAR and Multispectral Data
por: Allen, Matthew J, et al.
Publicado: (2024)
por: Allen, Matthew J, et al.
Publicado: (2024)
Invariant Representation via Decoupling Style and Spurious Features from Images
por: Li, Ruimeng, et al.
Publicado: (2023)
por: Li, Ruimeng, et al.
Publicado: (2023)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
por: Agarwal, Mahak, et al.
Publicado: (2025)
por: Agarwal, Mahak, et al.
Publicado: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
por: Liu, Zhongxin, et al.
Publicado: (2025)
por: Liu, Zhongxin, et al.
Publicado: (2025)
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
por: Pu, Bowei, et al.
Publicado: (2025)
por: Pu, Bowei, et al.
Publicado: (2025)
Global-Local Similarity for Efficient Fine-Grained Image Recognition with Vision Transformers
por: Rios, Edwin Arkel, et al.
Publicado: (2024)
por: Rios, Edwin Arkel, et al.
Publicado: (2024)
TEncDM: Understanding the Properties of the Diffusion Model in the Space of Language Model Encodings
por: Shabalin, Alexander, et al.
Publicado: (2024)
por: Shabalin, Alexander, et al.
Publicado: (2024)
When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets
por: Dai, Zeshi, et al.
Publicado: (2025)
por: Dai, Zeshi, et al.
Publicado: (2025)
Tversky Neural Networks: Psychologically Plausible Deep Learning with Differentiable Tversky Similarity
por: Doumbouya, Moussa Koulako Bala, et al.
Publicado: (2025)
por: Doumbouya, Moussa Koulako Bala, et al.
Publicado: (2025)
$\textit{sweet}$- An Open Source Modular Platform for Contactless Hand Vascular Biometric Experiments
por: Geissbühler, David, et al.
Publicado: (2024)
por: Geissbühler, David, et al.
Publicado: (2024)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
por: Zhu, Qian, et al.
Publicado: (2026)
por: Zhu, Qian, et al.
Publicado: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
por: Wu, Shuai, et al.
Publicado: (2026)
por: Wu, Shuai, et al.
Publicado: (2026)
Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
por: Reddy, Sandeep, et al.
Publicado: (2025)
por: Reddy, Sandeep, et al.
Publicado: (2025)
Adaptive Activation Cancellation for Hallucination Mitigation in Large Language Models
por: Yocam, Eric, et al.
Publicado: (2026)
por: Yocam, Eric, et al.
Publicado: (2026)
Evolutionary Algorithms Approach For Search Based On Semantic Document Similarity
por: Muniyappa, Chandrashekar, et al.
Publicado: (2025)
por: Muniyappa, Chandrashekar, et al.
Publicado: (2025)
Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
por: Ackermann, Richard, et al.
Publicado: (2025)
por: Ackermann, Richard, et al.
Publicado: (2025)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
por: Alberts, Lize, et al.
Publicado: (2024)
por: Alberts, Lize, et al.
Publicado: (2024)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
por: Shen, Meng, et al.
Publicado: (2026)
por: Shen, Meng, et al.
Publicado: (2026)
On the Limitations of Vision-Language Models in Understanding Image Transforms
por: Anis, Ahmad Mustafa, et al.
Publicado: (2025)
por: Anis, Ahmad Mustafa, et al.
Publicado: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
por: Jiang, Xintong, et al.
Publicado: (2025)
por: Jiang, Xintong, et al.
Publicado: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
por: Du, Bangde, et al.
Publicado: (2025)
por: Du, Bangde, et al.
Publicado: (2025)
Who Gets the Kidney? Human-AI Alignment, Indecision, and Moral Values
por: Dickerson, John P., et al.
Publicado: (2025)
por: Dickerson, John P., et al.
Publicado: (2025)
Ejemplares similares
-
The Effect of Data Partitioning Strategy on Model Generalizability: A Case Study of Morphological Segmentation
por: Liu, Zoey, et al.
Publicado: (2024) -
Revisiting Semantic Role Labeling: Efficient Structured Inference with Dependency-Informed Analysis
por: Youm, Sangpil, et al.
Publicado: (2026) -
AMREx: AMR for Explainable Fact Verification
por: Jayaweera, Chathuri, et al.
Publicado: (2024) -
DAHRS: Divergence-Aware Hallucination-Remediated SRL Projection
por: Youm, Sangpil, et al.
Publicado: (2024) -
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
por: Cherif, Ahmed
Publicado: (2026)