From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization
Fuente:
arXiv
Salvato in:
| Autori principali: | Belem, Catarina G., Pezeshkpour, Pouya, Iso, Hayate, Maekawa, Seiji, Bhutani, Nikita, Hruschka, Estevam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education
di: Iso, Hayate, et al.
Pubblicazione: (2025)
di: Iso, Hayate, et al.
Pubblicazione: (2025)
Less is More for Long Document Summary Evaluation by LLMs
di: Wu, Yunshu, et al.
Pubblicazione: (2023)
di: Wu, Yunshu, et al.
Pubblicazione: (2023)
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
di: Maekawa, Seiji, et al.
Pubblicazione: (2025)
di: Maekawa, Seiji, et al.
Pubblicazione: (2025)
Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual Data
di: Maekawa, Seiji, et al.
Pubblicazione: (2024)
di: Maekawa, Seiji, et al.
Pubblicazione: (2024)
Multi-Conditional Ranking with Large Language Models
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2024)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2024)
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025)
Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
di: Maekawa, Seiji, et al.
Pubblicazione: (2025)
di: Maekawa, Seiji, et al.
Pubblicazione: (2025)
Align then Train: Efficient Retrieval Adapter Learning
di: Maekawa, Seiji, et al.
Pubblicazione: (2026)
di: Maekawa, Seiji, et al.
Pubblicazione: (2026)
Insight-RAG: Enhancing LLMs with Insight-Driven Augmentation
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025)
From Task Solving to Robust Real-World Adaptation in LLM Agents
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2026)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2026)
Retrieval Helps or Hurts? A Deeper Dive into the Efficacy of Retrieval Augmentation to Language Models
di: Maekawa, Seiji, et al.
Pubblicazione: (2024)
di: Maekawa, Seiji, et al.
Pubblicazione: (2024)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2026)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2026)
Reasoning Capacity in Multi-Agent Systems: Limitations, Challenges and Human-Centered Solutions
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2024)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2024)
From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
di: Bayat, Farima Fatahi, et al.
Pubblicazione: (2025)
di: Bayat, Farima Fatahi, et al.
Pubblicazione: (2025)
Natural Language Processing for Human Resources: A Survey
di: Otani, Naoki, et al.
Pubblicazione: (2024)
di: Otani, Naoki, et al.
Pubblicazione: (2024)
XATU: A Fine-grained Instruction-based Benchmark for Explainable Text Updates
di: Zhang, Haopeng, et al.
Pubblicazione: (2023)
di: Zhang, Haopeng, et al.
Pubblicazione: (2023)
OmniTQA: A Cost-Aware System for Hybrid Query Processing over Semi-Structured Data
di: Shahbazi, Nima, et al.
Pubblicazione: (2026)
di: Shahbazi, Nima, et al.
Pubblicazione: (2026)
Same Content, Different Representations: A Controlled Study for Table QA
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025)
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-$k$
di: Taguchi, Chihiro, et al.
Pubblicazione: (2025)
di: Taguchi, Chihiro, et al.
Pubblicazione: (2025)
Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
di: Otani, Naoki, et al.
Pubblicazione: (2026)
di: Otani, Naoki, et al.
Pubblicazione: (2026)
AutoTemplate: A Simple Recipe for Lexically Constrained Text Generation
di: Iso, Hayate
Pubblicazione: (2022)
di: Iso, Hayate
Pubblicazione: (2022)
Noisy Pairing and Partial Supervision for Stylized Opinion Summarization
di: Iso, Hayate, et al.
Pubblicazione: (2022)
di: Iso, Hayate, et al.
Pubblicazione: (2022)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
di: Niwa, Ayana, et al.
Pubblicazione: (2024)
di: Niwa, Ayana, et al.
Pubblicazione: (2024)
Verification-Aware Planning for Multi-Agent Systems
di: Xu, Tianyang, et al.
Pubblicazione: (2025)
di: Xu, Tianyang, et al.
Pubblicazione: (2025)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
di: Davoodi, Arash Gholami, et al.
Pubblicazione: (2024)
di: Davoodi, Arash Gholami, et al.
Pubblicazione: (2024)
A Blueprint Architecture of Compound AI Systems for Enterprise
di: Kandogan, Eser, et al.
Pubblicazione: (2024)
di: Kandogan, Eser, et al.
Pubblicazione: (2024)
A Dynamic Self-Evolving Extraction System
di: Amin-Naseri, Moin, et al.
Pubblicazione: (2026)
di: Amin-Naseri, Moin, et al.
Pubblicazione: (2026)
Towards Probabilistic Question Answering Over Tabular Data
di: Shen, Chen, et al.
Pubblicazione: (2025)
di: Shen, Chen, et al.
Pubblicazione: (2025)
RECAP: REwriting Conversations for Intent Understanding in Agentic Planning
di: Mitra, Kushan, et al.
Pubblicazione: (2025)
di: Mitra, Kushan, et al.
Pubblicazione: (2025)
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
di: Li, Chuyuan, et al.
Pubblicazione: (2025)
di: Li, Chuyuan, et al.
Pubblicazione: (2025)
FactLens: Benchmarking Fine-Grained Fact Verification
di: Mitra, Kushan, et al.
Pubblicazione: (2024)
di: Mitra, Kushan, et al.
Pubblicazione: (2024)
How to Steer Your Multi-Agent System: Human-LLM Collaborative Planning
di: He, Zeyu, et al.
Pubblicazione: (2026)
di: He, Zeyu, et al.
Pubblicazione: (2026)
Orchestrating Agents and Data for Enterprise: A Blueprint Architecture for Compound AI
di: Kandogan, Eser, et al.
Pubblicazione: (2025)
di: Kandogan, Eser, et al.
Pubblicazione: (2025)
Characterizing Large Language Models as Rationalizers of Knowledge-intensive Tasks
di: Mishra, Aditi, et al.
Pubblicazione: (2023)
di: Mishra, Aditi, et al.
Pubblicazione: (2023)
Do Multi-Document Summarization Models Synthesize?
di: DeYoung, Jay, et al.
Pubblicazione: (2023)
di: DeYoung, Jay, et al.
Pubblicazione: (2023)
Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models
di: Davoodi, Arash Gholami, et al.
Pubblicazione: (2026)
di: Davoodi, Arash Gholami, et al.
Pubblicazione: (2026)
Blue Data Intelligence Layer: Streaming Data and Agents for Multi-source Multi-modal Data-Centric Applications
di: Aminnaseri, Moin, et al.
Pubblicazione: (2026)
di: Aminnaseri, Moin, et al.
Pubblicazione: (2026)
GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews
di: Darrin, Maxime, et al.
Pubblicazione: (2024)
di: Darrin, Maxime, et al.
Pubblicazione: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
di: Tang, Liyan, et al.
Pubblicazione: (2024)
di: Tang, Liyan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education
di: Iso, Hayate, et al.
Pubblicazione: (2025) -
Less is More for Long Document Summary Evaluation by LLMs
di: Wu, Yunshu, et al.
Pubblicazione: (2023) -
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
di: Maekawa, Seiji, et al.
Pubblicazione: (2025) -
Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual Data
di: Maekawa, Seiji, et al.
Pubblicazione: (2024) -
Multi-Conditional Ranking with Large Language Models
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2024)