Guardado en:
| Autores principales: | Mehenni, Gaya, Lamarche, Fabrice, Rios-Ibacache, Odette, Kildea, John, Zouaq, Amal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2504.08596 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ontology-Constrained Generation of Domain-Specific Clinical Summaries
por: Mehenni, Gaya, et al.
Publicado: (2024)
por: Mehenni, Gaya, et al.
Publicado: (2024)
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
por: Mascarell, Laura, et al.
Publicado: (2024)
por: Mascarell, Laura, et al.
Publicado: (2024)
MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
por: V., Diego Fajardo, et al.
Publicado: (2025)
por: V., Diego Fajardo, et al.
Publicado: (2025)
MedAide: Leveraging Large Language Models for On-Premise Medical Assistance on Edge Devices
por: Basit, Abdul, et al.
Publicado: (2024)
por: Basit, Abdul, et al.
Publicado: (2024)
Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?
por: Argese, Alex, et al.
Publicado: (2026)
por: Argese, Alex, et al.
Publicado: (2026)
Seeing Through the Fog: A Cost-Effectiveness Analysis of Hallucination Detection Systems
por: Thomas, Alexander, et al.
Publicado: (2024)
por: Thomas, Alexander, et al.
Publicado: (2024)
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
por: Li, Shanghao, et al.
Publicado: (2025)
por: Li, Shanghao, et al.
Publicado: (2025)
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
por: Benkirane, Kenza, et al.
Publicado: (2024)
por: Benkirane, Kenza, et al.
Publicado: (2024)
CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification
por: Ye, Severin, et al.
Publicado: (2026)
por: Ye, Severin, et al.
Publicado: (2026)
How to Evaluate Medical AI
por: Kopanichuk, Ilia, et al.
Publicado: (2025)
por: Kopanichuk, Ilia, et al.
Publicado: (2025)
Listen to the Layers: Mitigating Hallucinations with Inter-Layer Disagreement
por: Subbalakshmi, Koduvayur, et al.
Publicado: (2026)
por: Subbalakshmi, Koduvayur, et al.
Publicado: (2026)
TRACE: Trajectory Correction from Cross-layer Evidence for Hallucination Reduction
por: Ranade, Tej Sanibh
Publicado: (2026)
por: Ranade, Tej Sanibh
Publicado: (2026)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
por: Yuan, Weikang, et al.
Publicado: (2025)
por: Yuan, Weikang, et al.
Publicado: (2025)
Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks
por: Buszydlik, Aleksander, et al.
Publicado: (2023)
por: Buszydlik, Aleksander, et al.
Publicado: (2023)
UrduLLaMA 1.0: Dataset Curation, Preprocessing, and Evaluation in Low-Resource Settings
por: Fiaz, Layba, et al.
Publicado: (2025)
por: Fiaz, Layba, et al.
Publicado: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
por: Naeem, Numaan, et al.
Publicado: (2025)
por: Naeem, Numaan, et al.
Publicado: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
por: Stowe, Kevin, et al.
Publicado: (2026)
por: Stowe, Kevin, et al.
Publicado: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
por: Hawkins, John
Publicado: (2025)
por: Hawkins, John
Publicado: (2025)
Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
por: Pan, Muyu, et al.
Publicado: (2025)
por: Pan, Muyu, et al.
Publicado: (2025)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
por: Michail, Andrianos, et al.
Publicado: (2024)
por: Michail, Andrianos, et al.
Publicado: (2024)
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
por: Quevedo, Ernesto, et al.
Publicado: (2024)
por: Quevedo, Ernesto, et al.
Publicado: (2024)
Reducing Hallucinations in Summarization via Reinforcement Learning with Entity Hallucination Index
por: Katwe, Praveenkumar, et al.
Publicado: (2025)
por: Katwe, Praveenkumar, et al.
Publicado: (2025)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
por: Palit, Sayon, et al.
Publicado: (2025)
por: Palit, Sayon, et al.
Publicado: (2025)
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection
por: Hanif, Ikhlasul Akmal, et al.
Publicado: (2025)
por: Hanif, Ikhlasul Akmal, et al.
Publicado: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
por: Ovcharov, Volodymyr
Publicado: (2026)
por: Ovcharov, Volodymyr
Publicado: (2026)
Serialisation Strategy Matters: How FHIR Data Format Affects LLM Medication Reconciliation
por: Pator, Sanjoy
Publicado: (2026)
por: Pator, Sanjoy
Publicado: (2026)
A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text
por: Sagae, Alicia, et al.
Publicado: (2025)
por: Sagae, Alicia, et al.
Publicado: (2025)
No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA
por: Simione II, Robert L
Publicado: (2024)
por: Simione II, Robert L
Publicado: (2024)
Accurate and Energy Efficient: Local Retrieval-Augmented Generation Models Outperform Commercial Large Language Models in Medical Tasks
por: Vrettos, Konstantinos, et al.
Publicado: (2025)
por: Vrettos, Konstantinos, et al.
Publicado: (2025)
Knowledge Graphs, Large Language Models, and Hallucinations: An NLP Perspective
por: Lavrinovics, Ernests, et al.
Publicado: (2024)
por: Lavrinovics, Ernests, et al.
Publicado: (2024)
Low-Resource Court Judgment Summarization for Common Law Systems
por: Liu, Shuaiqi, et al.
Publicado: (2024)
por: Liu, Shuaiqi, et al.
Publicado: (2024)
Detecting and Steering LLMs' Empathy in Action
por: Cadile, Juan P.
Publicado: (2025)
por: Cadile, Juan P.
Publicado: (2025)
AI-Powered Detection of Inappropriate Language in Medical School Curricula
por: Salavati, Chiman, et al.
Publicado: (2025)
por: Salavati, Chiman, et al.
Publicado: (2025)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
por: Orgad, Hadas, et al.
Publicado: (2024)
por: Orgad, Hadas, et al.
Publicado: (2024)
Identifying Bias in Machine-generated Text Detection
por: Stowe, Kevin, et al.
Publicado: (2025)
por: Stowe, Kevin, et al.
Publicado: (2025)
Detecting AI-Generated Texts in Cross-Domains
por: Zhou, You, et al.
Publicado: (2024)
por: Zhou, You, et al.
Publicado: (2024)
Ejemplares similares
-
Ontology-Constrained Generation of Domain-Specific Clinical Summaries
por: Mehenni, Gaya, et al.
Publicado: (2024) -
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
por: Mascarell, Laura, et al.
Publicado: (2024) -
MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
por: V., Diego Fajardo, et al.
Publicado: (2025) -
MedAide: Leveraging Large Language Models for On-Premise Medical Assistance on Edge Devices
por: Basit, Abdul, et al.
Publicado: (2024) -
Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?
por: Argese, Alex, et al.
Publicado: (2026)