HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
Fuente:
arXiv
Guardado en:
| Autor principal: | Cherif, Ahmed |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
por: Zhu, Qian, et al.
Publicado: (2026)
por: Zhu, Qian, et al.
Publicado: (2026)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
por: Zhang, Luyan, et al.
Publicado: (2025)
por: Zhang, Luyan, et al.
Publicado: (2025)
Towards Conditioning Clinical Text Generation for User Control
por: Koraş, Osman Alperen, et al.
Publicado: (2025)
por: Koraş, Osman Alperen, et al.
Publicado: (2025)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
por: Han, Lifeng, et al.
Publicado: (2016)
por: Han, Lifeng, et al.
Publicado: (2016)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
por: Dayarathne, Ranul, et al.
Publicado: (2025)
por: Dayarathne, Ranul, et al.
Publicado: (2025)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
por: Hashemi, Helia, et al.
Publicado: (2024)
por: Hashemi, Helia, et al.
Publicado: (2024)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
por: Nguyen, Huyen, et al.
Publicado: (2026)
por: Nguyen, Huyen, et al.
Publicado: (2026)
Challenges and Opportunities of NLP for HR Applications: A Discussion Paper
por: Leidner, Jochen L., et al.
Publicado: (2024)
por: Leidner, Jochen L., et al.
Publicado: (2024)
ReTreVal: Reasoning Tree with Validation -- A Hybrid Framework for Enhanced LLM Multi-Step Reasoning
por: HS, Abhishek, et al.
Publicado: (2026)
por: HS, Abhishek, et al.
Publicado: (2026)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
por: Hasan, Mohammed Rakibul
Publicado: (2026)
por: Hasan, Mohammed Rakibul
Publicado: (2026)
AskSport: Web Application for Sports Question-Answering
por: Onofre, Enzo B, et al.
Publicado: (2025)
por: Onofre, Enzo B, et al.
Publicado: (2025)
How to Evaluate Medical AI
por: Kopanichuk, Ilia, et al.
Publicado: (2025)
por: Kopanichuk, Ilia, et al.
Publicado: (2025)
Improving the Capabilities of Large Language Model Based Marketing Analytics Copilots With Semantic Search And Fine-Tuning
por: Gao, Yilin, et al.
Publicado: (2024)
por: Gao, Yilin, et al.
Publicado: (2024)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
por: Rodriguez, David, et al.
Publicado: (2025)
por: Rodriguez, David, et al.
Publicado: (2025)
Comparative Analysis of AI Agent Architectures for Entity Relationship Classification
por: Berijanian, Maryam, et al.
Publicado: (2025)
por: Berijanian, Maryam, et al.
Publicado: (2025)
AI-Powered Detection of Inappropriate Language in Medical School Curricula
por: Salavati, Chiman, et al.
Publicado: (2025)
por: Salavati, Chiman, et al.
Publicado: (2025)
Mitigating Structural Noise in Low-Resource S2TT: An Optimized Cascaded Nepali-English Pipeline with Punctuation Restoration
por: Chongbang, Tangsang, et al.
Publicado: (2026)
por: Chongbang, Tangsang, et al.
Publicado: (2026)
LLMs Simulate Big Five Personality Traits: Further Evidence
por: Sorokovikova, Aleksandra, et al.
Publicado: (2024)
por: Sorokovikova, Aleksandra, et al.
Publicado: (2024)
Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding
por: Rollman, John C., et al.
Publicado: (2025)
por: Rollman, John C., et al.
Publicado: (2025)
BLT: Can Large Language Models Handle Basic Legal Text?
por: Blair-Stanek, Andrew, et al.
Publicado: (2023)
por: Blair-Stanek, Andrew, et al.
Publicado: (2023)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
por: Simhi, Adi, et al.
Publicado: (2024)
por: Simhi, Adi, et al.
Publicado: (2024)
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
por: Sauter, Andreas, et al.
Publicado: (2026)
por: Sauter, Andreas, et al.
Publicado: (2026)
Leveraging Large Language Models to Extract and Translate Medical Information in Doctors' Notes for Health Records and Diagnostic Billing Codes
por: Hartnett, Peter, et al.
Publicado: (2026)
por: Hartnett, Peter, et al.
Publicado: (2026)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
por: Liu, Zhongxin, et al.
Publicado: (2025)
por: Liu, Zhongxin, et al.
Publicado: (2025)
Social and Ethical Risks Posed by General-Purpose LLMs for Settling Newcomers in Canada
por: Nejadgholi, Isar, et al.
Publicado: (2024)
por: Nejadgholi, Isar, et al.
Publicado: (2024)
The Need for Guardrails with Large Language Models in Medical Safety-Critical Settings: An Artificial Intelligence Application in the Pharmacovigilance Ecosystem
por: Hakim, Joe B, et al.
Publicado: (2024)
por: Hakim, Joe B, et al.
Publicado: (2024)
A Llama walks into the 'Bar': Efficient Supervised Fine-Tuning for Legal Reasoning in the Multi-state Bar Exam
por: Fernandes, Rean, et al.
Publicado: (2025)
por: Fernandes, Rean, et al.
Publicado: (2025)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
por: Sun, Mingrui, et al.
Publicado: (2026)
por: Sun, Mingrui, et al.
Publicado: (2026)
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
por: Mullens, Drake, et al.
Publicado: (2026)
por: Mullens, Drake, et al.
Publicado: (2026)
Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data
por: de Campos, Andresa Rodrigues, et al.
Publicado: (2026)
por: de Campos, Andresa Rodrigues, et al.
Publicado: (2026)
Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts
por: Simons, Arno
Publicado: (2026)
por: Simons, Arno
Publicado: (2026)
Exploring the Structure of AI-Induced Language Change in Scientific English
por: Galpin, Riley, et al.
Publicado: (2025)
por: Galpin, Riley, et al.
Publicado: (2025)
VERITAS-NLI : Validation and Extraction of Reliable Information Through Automated Scraping and Natural Language Inference
por: Shah, Arjun, et al.
Publicado: (2024)
por: Shah, Arjun, et al.
Publicado: (2024)
Teaching a Language Model to Speak the Language of Tools
por: Emanuilov, Simeon
Publicado: (2025)
por: Emanuilov, Simeon
Publicado: (2025)
OEMA: Ontology-Enhanced Multi-Agent Collaboration Framework for Zero-Shot Clinical Named Entity Recognition
por: Tao, Xinli, et al.
Publicado: (2025)
por: Tao, Xinli, et al.
Publicado: (2025)
Scaling Laws for State Dynamics in Large Language Models
por: Li, Jacob X, et al.
Publicado: (2025)
por: Li, Jacob X, et al.
Publicado: (2025)
Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design
por: Charles, Richard M., et al.
Publicado: (2025)
por: Charles, Richard M., et al.
Publicado: (2025)
Project Riley: Multimodal Multi-Agent LLM Collaboration with Emotional Reasoning and Voting
por: Ortigoso, Ana Rita, et al.
Publicado: (2025)
por: Ortigoso, Ana Rita, et al.
Publicado: (2025)
From Instructor to Collaborator: What a 90-Participant Study Reveals about Human-Agent Collaboration in a Mobile Serious Game
por: Korre, Danai
Publicado: (2026)
por: Korre, Danai
Publicado: (2026)
AutoTRIZ: Automating Engineering Innovation with TRIZ and Large Language Models
por: Jiang, Shuo, et al.
Publicado: (2024)
por: Jiang, Shuo, et al.
Publicado: (2024)
Ejemplares similares
-
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
por: Zhu, Qian, et al.
Publicado: (2026) -
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
por: Zhang, Luyan, et al.
Publicado: (2025) -
Towards Conditioning Clinical Text Generation for User Control
por: Koraş, Osman Alperen, et al.
Publicado: (2025) -
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
por: Han, Lifeng, et al.
Publicado: (2016) -
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
por: Dayarathne, Ranul, et al.
Publicado: (2025)