German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Mascarell, Laura, Chalumattu, Ribin, Rios, Annette |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedHal: An Evaluation Dataset for Medical Hallucination Detection
by: Mehenni, Gaya, et al.
Published: (2025)
by: Mehenni, Gaya, et al.
Published: (2025)
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
by: Benkirane, Kenza, et al.
Published: (2024)
by: Benkirane, Kenza, et al.
Published: (2024)
Seeing Through the Fog: A Cost-Effectiveness Analysis of Hallucination Detection Systems
by: Thomas, Alexander, et al.
Published: (2024)
by: Thomas, Alexander, et al.
Published: (2024)
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
by: Li, Shanghao, et al.
Published: (2025)
by: Li, Shanghao, et al.
Published: (2025)
AustroTox: A Dataset for Target-Based Austrian German Offensive Language Detection
by: Pachinger, Pia, et al.
Published: (2024)
by: Pachinger, Pia, et al.
Published: (2024)
Ontology-Constrained Generation of Domain-Specific Clinical Summaries
by: Mehenni, Gaya, et al.
Published: (2024)
by: Mehenni, Gaya, et al.
Published: (2024)
Listen to the Layers: Mitigating Hallucinations with Inter-Layer Disagreement
by: Subbalakshmi, Koduvayur, et al.
Published: (2026)
by: Subbalakshmi, Koduvayur, et al.
Published: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?
by: Argese, Alex, et al.
Published: (2026)
by: Argese, Alex, et al.
Published: (2026)
TRACE: Trajectory Correction from Cross-layer Evidence for Hallucination Reduction
by: Ranade, Tej Sanibh
Published: (2026)
by: Ranade, Tej Sanibh
Published: (2026)
Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks
by: Buszydlik, Aleksander, et al.
Published: (2023)
by: Buszydlik, Aleksander, et al.
Published: (2023)
Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
by: Pan, Muyu, et al.
Published: (2025)
by: Pan, Muyu, et al.
Published: (2025)
Historical Ink: Semantic Shift Detection for 19th Century Spanish
by: Montes, Tony, et al.
Published: (2024)
by: Montes, Tony, et al.
Published: (2024)
Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models
by: Gilhuly, Colleen, et al.
Published: (2025)
by: Gilhuly, Colleen, et al.
Published: (2025)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
by: Quevedo, Ernesto, et al.
Published: (2024)
by: Quevedo, Ernesto, et al.
Published: (2024)
Reducing Hallucinations in Summarization via Reinforcement Learning with Entity Hallucination Index
by: Katwe, Praveenkumar, et al.
Published: (2025)
by: Katwe, Praveenkumar, et al.
Published: (2025)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
by: Yuan, Weikang, et al.
Published: (2025)
by: Yuan, Weikang, et al.
Published: (2025)
UrduLLaMA 1.0: Dataset Curation, Preprocessing, and Evaluation in Low-Resource Settings
by: Fiaz, Layba, et al.
Published: (2025)
by: Fiaz, Layba, et al.
Published: (2025)
A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text
by: Sagae, Alicia, et al.
Published: (2025)
by: Sagae, Alicia, et al.
Published: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA
by: Simione II, Robert L
Published: (2024)
by: Simione II, Robert L
Published: (2024)
Detecting and Steering LLMs' Empathy in Action
by: Cadile, Juan P.
Published: (2025)
by: Cadile, Juan P.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Detecting AI-Generated Texts in Cross-Domains
by: Zhou, You, et al.
Published: (2024)
by: Zhou, You, et al.
Published: (2024)
Identifying Bias in Machine-generated Text Detection
by: Stowe, Kevin, et al.
Published: (2025)
by: Stowe, Kevin, et al.
Published: (2025)
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
by: Stowe, Kevin, et al.
Published: (2026)
by: Stowe, Kevin, et al.
Published: (2026)
Detecting Data Contamination in LLMs via In-Context Learning
by: Zawalski, Michał, et al.
Published: (2025)
by: Zawalski, Michał, et al.
Published: (2025)
Blessing or curse? A survey on the Impact of Generative AI on Fake News
by: Loth, Alexander, et al.
Published: (2024)
by: Loth, Alexander, et al.
Published: (2024)
CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification
by: Ye, Severin, et al.
Published: (2026)
by: Ye, Severin, et al.
Published: (2026)
The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
by: Larsen, Erik
Published: (2025)
by: Larsen, Erik
Published: (2025)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
by: Michail, Andrianos, et al.
Published: (2024)
by: Michail, Andrianos, et al.
Published: (2024)
On the Effectiveness of LLM-Specific Fine-Tuning for Detecting AI-Generated Text
by: Gromadzki, Michał, et al.
Published: (2026)
by: Gromadzki, Michał, et al.
Published: (2026)
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Paying Attention to Deflections: Mining Pragmatic Nuances for Whataboutism Detection in Online Discourse
by: Phi, Khiem, et al.
Published: (2024)
by: Phi, Khiem, et al.
Published: (2024)
AEyeDE: An Attention-Based Attribution Framework for AI-Generated Text Detection
by: Nourbakhsh, Aria, et al.
Published: (2026)
by: Nourbakhsh, Aria, et al.
Published: (2026)
Low-Resource Court Judgment Summarization for Common Law Systems
by: Liu, Shuaiqi, et al.
Published: (2024)
by: Liu, Shuaiqi, et al.
Published: (2024)
Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative Likelihood
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Exploring News Summarization and Enrichment in a Highly Resource-Scarce Indian Language: A Case Study of Mizo
by: Bala, Abhinaba, et al.
Published: (2024)
by: Bala, Abhinaba, et al.
Published: (2024)
Enhancing Sentiment Classification and Irony Detection in Large Language Models through Advanced Prompt Engineering Techniques
by: Schmitt, Marvin, et al.
Published: (2026)
by: Schmitt, Marvin, et al.
Published: (2026)
Similar Items
-
MedHal: An Evaluation Dataset for Medical Hallucination Detection
by: Mehenni, Gaya, et al.
Published: (2025) -
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
by: Benkirane, Kenza, et al.
Published: (2024) -
Seeing Through the Fog: A Cost-Effectiveness Analysis of Hallucination Detection Systems
by: Thomas, Alexander, et al.
Published: (2024) -
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
by: Li, Shanghao, et al.
Published: (2025) -
AustroTox: A Dataset for Target-Based Austrian German Offensive Language Detection
by: Pachinger, Pia, et al.
Published: (2024)