Less is More for Long Document Summary Evaluation by LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yunshu, Iso, Hayate, Pezeshkpour, Pouya, Bhutani, Nikita, Hruschka, Estevam |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education
by: Iso, Hayate, et al.
Published: (2025)
by: Iso, Hayate, et al.
Published: (2025)
From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization
by: Belem, Catarina G., et al.
Published: (2024)
by: Belem, Catarina G., et al.
Published: (2024)
Insight-RAG: Enhancing LLMs with Insight-Driven Augmentation
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
From Task Solving to Robust Real-World Adaptation in LLM Agents
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
Multi-Conditional Ranking with Large Language Models
by: Pezeshkpour, Pouya, et al.
Published: (2024)
by: Pezeshkpour, Pouya, et al.
Published: (2024)
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
Reasoning Capacity in Multi-Agent Systems: Limitations, Challenges and Human-Centered Solutions
by: Pezeshkpour, Pouya, et al.
Published: (2024)
by: Pezeshkpour, Pouya, et al.
Published: (2024)
Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual Data
by: Maekawa, Seiji, et al.
Published: (2024)
by: Maekawa, Seiji, et al.
Published: (2024)
From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
by: Bayat, Farima Fatahi, et al.
Published: (2025)
by: Bayat, Farima Fatahi, et al.
Published: (2025)
Natural Language Processing for Human Resources: A Survey
by: Otani, Naoki, et al.
Published: (2024)
by: Otani, Naoki, et al.
Published: (2024)
Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
Align then Train: Efficient Retrieval Adapter Learning
by: Maekawa, Seiji, et al.
Published: (2026)
by: Maekawa, Seiji, et al.
Published: (2026)
XATU: A Fine-grained Instruction-based Benchmark for Explainable Text Updates
by: Zhang, Haopeng, et al.
Published: (2023)
by: Zhang, Haopeng, et al.
Published: (2023)
Retrieval Helps or Hurts? A Deeper Dive into the Efficacy of Retrieval Augmentation to Language Models
by: Maekawa, Seiji, et al.
Published: (2024)
by: Maekawa, Seiji, et al.
Published: (2024)
Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
by: Otani, Naoki, et al.
Published: (2026)
by: Otani, Naoki, et al.
Published: (2026)
AutoTemplate: A Simple Recipe for Lexically Constrained Text Generation
by: Iso, Hayate
Published: (2022)
by: Iso, Hayate
Published: (2022)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
by: Niwa, Ayana, et al.
Published: (2024)
by: Niwa, Ayana, et al.
Published: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
by: Davoodi, Arash Gholami, et al.
Published: (2024)
by: Davoodi, Arash Gholami, et al.
Published: (2024)
Noisy Pairing and Partial Supervision for Stylized Opinion Summarization
by: Iso, Hayate, et al.
Published: (2022)
by: Iso, Hayate, et al.
Published: (2022)
A Blueprint Architecture of Compound AI Systems for Enterprise
by: Kandogan, Eser, et al.
Published: (2024)
by: Kandogan, Eser, et al.
Published: (2024)
OmniTQA: A Cost-Aware System for Hybrid Query Processing over Semi-Structured Data
by: Shahbazi, Nima, et al.
Published: (2026)
by: Shahbazi, Nima, et al.
Published: (2026)
A Dynamic Self-Evolving Extraction System
by: Amin-Naseri, Moin, et al.
Published: (2026)
by: Amin-Naseri, Moin, et al.
Published: (2026)
Towards Probabilistic Question Answering Over Tabular Data
by: Shen, Chen, et al.
Published: (2025)
by: Shen, Chen, et al.
Published: (2025)
RECAP: REwriting Conversations for Intent Understanding in Agentic Planning
by: Mitra, Kushan, et al.
Published: (2025)
by: Mitra, Kushan, et al.
Published: (2025)
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-$k$
by: Taguchi, Chihiro, et al.
Published: (2025)
by: Taguchi, Chihiro, et al.
Published: (2025)
FactLens: Benchmarking Fine-Grained Fact Verification
by: Mitra, Kushan, et al.
Published: (2024)
by: Mitra, Kushan, et al.
Published: (2024)
Same Content, Different Representations: A Controlled Study for Table QA
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Orchestrating Agents and Data for Enterprise: A Blueprint Architecture for Compound AI
by: Kandogan, Eser, et al.
Published: (2025)
by: Kandogan, Eser, et al.
Published: (2025)
Verification-Aware Planning for Multi-Agent Systems
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
Characterizing Large Language Models as Rationalizers of Knowledge-intensive Tasks
by: Mishra, Aditi, et al.
Published: (2023)
by: Mishra, Aditi, et al.
Published: (2023)
Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure
by: Tan, Chenchen, et al.
Published: (2026)
by: Tan, Chenchen, et al.
Published: (2026)
Geometry-Aware Decoding with Wasserstein-Regularized Truncation and Mass Penalties for Large Language Models
by: Davoodi, Arash Gholami, et al.
Published: (2026)
by: Davoodi, Arash Gholami, et al.
Published: (2026)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
by: Hassell, Jackson, et al.
Published: (2025)
by: Hassell, Jackson, et al.
Published: (2025)
More Aligned, Less Diverse? Analyzing the Grammar and Lexicon of Two Generations of LLMs
by: Gude, Adrián, et al.
Published: (2026)
by: Gude, Adrián, et al.
Published: (2026)
Similar Items
-
Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education
by: Iso, Hayate, et al.
Published: (2025) -
From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization
by: Belem, Catarina G., et al.
Published: (2024) -
Insight-RAG: Enhancing LLMs with Insight-Driven Augmentation
by: Pezeshkpour, Pouya, et al.
Published: (2025) -
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
by: Pezeshkpour, Pouya, et al.
Published: (2025) -
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026)