Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends
Fuente:
arXiv
Saved in:
| Main Authors: | Ramprasad, Sanjana, Ferracane, Elisa, Lipton, Zachary C. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues
by: Binici, Kuluhan, et al.
Published: (2024)
by: Binici, Kuluhan, et al.
Published: (2024)
LLM-Select: Feature Selection with Large Language Models
by: Jeong, Daniel P., et al.
Published: (2024)
by: Jeong, Daniel P., et al.
Published: (2024)
Correction with Backtracking Reduces Hallucination in Summarization
by: Liu, Zhenzhen, et al.
Published: (2023)
by: Liu, Zhenzhen, et al.
Published: (2023)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization
by: Ryu, Sangwon, et al.
Published: (2026)
by: Ryu, Sangwon, et al.
Published: (2026)
Hallucination Detection-Guided Preference Optimization for Clinical Summarization
by: Seethakantha, Shamanth Kuthpadi, et al.
Published: (2026)
by: Seethakantha, Shamanth Kuthpadi, et al.
Published: (2026)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
by: Chae, Kyubyung, et al.
Published: (2024)
by: Chae, Kyubyung, et al.
Published: (2024)
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
by: Liu, Yinhong, et al.
Published: (2025)
by: Liu, Yinhong, et al.
Published: (2025)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
by: Jin, Keyan, et al.
Published: (2025)
by: Jin, Keyan, et al.
Published: (2025)
Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
by: Chrysostomou, George, et al.
Published: (2023)
by: Chrysostomou, George, et al.
Published: (2023)
CADS: A Systematic Literature Review on the Challenges of Abstractive Dialogue Summarization
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
by: Bao, Forrest Sheng, et al.
Published: (2024)
by: Bao, Forrest Sheng, et al.
Published: (2024)
Lessons from the Field: An Adaptable Lifecycle Approach to Applied Dialogue Summarization
by: Chawla, Kushal, et al.
Published: (2026)
by: Chawla, Kushal, et al.
Published: (2026)
Enhancing Consistency of Werewolf AI through Dialogue Summarization and Persona Information
by: Tanaka, Yoshiki, et al.
Published: (2026)
by: Tanaka, Yoshiki, et al.
Published: (2026)
Systematic Exploration of Dialogue Summarization Approaches for Reproducibility, Comparative Assessment, and Methodological Innovations for Advancing Natural Language Processing in Abstractive Summarization
by: Gogireddy, Yugandhar Reddy, et al.
Published: (2024)
by: Gogireddy, Yugandhar Reddy, et al.
Published: (2024)
Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection
by: He, Jianfeng, et al.
Published: (2024)
by: He, Jianfeng, et al.
Published: (2024)
Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models
by: Wang, Qingyue, et al.
Published: (2023)
by: Wang, Qingyue, et al.
Published: (2023)
QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
by: Ghebriout, Mohamed Imed Eddine, et al.
Published: (2025)
by: Ghebriout, Mohamed Imed Eddine, et al.
Published: (2025)
Utilizing GPT to Enhance Text Summarization: A Strategy to Minimize Hallucinations
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
Sequence-Level Certainty Reduces Hallucination In Knowledge-Grounded Dialogue Generation
by: Wan, Yixin, et al.
Published: (2023)
by: Wan, Yixin, et al.
Published: (2023)
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
by: Zhu, Ruizhe, et al.
Published: (2025)
by: Zhu, Ruizhe, et al.
Published: (2025)
Hallucination Diversity-Aware Active Learning for Text Summarization
by: Xia, Yu, et al.
Published: (2024)
by: Xia, Yu, et al.
Published: (2024)
Personalized Language Modeling from Personalized Human Feedback
by: Li, Xinyu, et al.
Published: (2024)
by: Li, Xinyu, et al.
Published: (2024)
Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization
by: Mei, Xiaoyong, et al.
Published: (2026)
by: Mei, Xiaoyong, et al.
Published: (2026)
Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
by: Zhong, Yang, et al.
Published: (2025)
by: Zhong, Yang, et al.
Published: (2025)
A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue Generation
by: Yu, Jifan, et al.
Published: (2024)
by: Yu, Jifan, et al.
Published: (2024)
Analyzing the Performance of Large Language Models on Code Summarization
by: Haldar, Rajarshi, et al.
Published: (2024)
by: Haldar, Rajarshi, et al.
Published: (2024)
Tell Me Why: Designing an Explainable LLM-based Dialogue System for Student Problem Behavior Diagnosis
by: Fan, Zhilin, et al.
Published: (2026)
by: Fan, Zhilin, et al.
Published: (2026)
Learning to Summarize from LLM-generated Feedback
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
by: Chen, Kedi, et al.
Published: (2024)
by: Chen, Kedi, et al.
Published: (2024)
Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
by: Jeong, Daniel P., et al.
Published: (2024)
by: Jeong, Daniel P., et al.
Published: (2024)
STRUM-LLM: Attributed and Structured Contrastive Summarization
by: Gunel, Beliz, et al.
Published: (2024)
by: Gunel, Beliz, et al.
Published: (2024)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
by: Yuan, Dong, et al.
Published: (2024)
by: Yuan, Dong, et al.
Published: (2024)
Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations
by: Chen, Yi-Pei, et al.
Published: (2024)
by: Chen, Yi-Pei, et al.
Published: (2024)
Similar Items
-
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
by: Ramprasad, Sanjana, et al.
Published: (2024) -
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024) -
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024) -
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025) -
MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues
by: Binici, Kuluhan, et al.
Published: (2024)