Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yixin, Fabbri, Alexander R., Chen, Jiawen, Zhao, Yilun, Han, Simeng, Joty, Shafiq, Liu, Pengfei, Radev, Dragomir, Wu, Chien-Sheng, Cohan, Arman |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Learning to Summarize with Large Language Models as References
by: Liu, Yixin, et al.
Published: (2023)
by: Liu, Yixin, et al.
Published: (2023)
ReIFE: Re-evaluating Instruction-Following Evaluation
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
References Improve LLM Alignment in Non-Verifiable Domains
by: Shi, Kejian, et al.
Published: (2026)
by: Shi, Kejian, et al.
Published: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
by: Han, Simeng, et al.
Published: (2024)
by: Han, Simeng, et al.
Published: (2024)
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
by: Huang, Kung-Hsiang, et al.
Published: (2023)
by: Huang, Kung-Hsiang, et al.
Published: (2023)
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
by: Li, Chuhan, et al.
Published: (2024)
by: Li, Chuhan, et al.
Published: (2024)
On Context Utilization in Summarization with Large Language Models
by: Ravaut, Mathieu, et al.
Published: (2023)
by: Ravaut, Mathieu, et al.
Published: (2023)
HYBRIDMIND: Meta Selection of Natural Language and Symbolic Language for Enhanced LLM Reasoning
by: Han, Simeng, et al.
Published: (2024)
by: Han, Simeng, et al.
Published: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
by: Deng, Chunyuan, et al.
Published: (2023)
by: Deng, Chunyuan, et al.
Published: (2023)
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
by: Zhao, Yilun, et al.
Published: (2025)
by: Zhao, Yilun, et al.
Published: (2025)
Fair Abstractive Summarization of Diverse Perspectives
by: Zhang, Yusen, et al.
Published: (2023)
by: Zhang, Yusen, et al.
Published: (2023)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
by: Riddell, Martin, et al.
Published: (2024)
by: Riddell, Martin, et al.
Published: (2024)
Unsupervised Summarization Re-ranking
by: Ravaut, Mathieu, et al.
Published: (2022)
by: Ravaut, Mathieu, et al.
Published: (2022)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
by: Agarwal, Divyansh, et al.
Published: (2024)
by: Agarwal, Divyansh, et al.
Published: (2024)
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
by: Lee, Jinu, et al.
Published: (2025)
by: Lee, Jinu, et al.
Published: (2025)
PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving
by: Feng, Kaiyue, et al.
Published: (2025)
by: Feng, Kaiyue, et al.
Published: (2025)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
Calibrating Long-form Generations from Large Language Models
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
by: Phanse, Rohan, et al.
Published: (2025)
by: Phanse, Rohan, et al.
Published: (2025)
On Positional Bias of Faithfulness for Long-form Summarization
by: Wan, David, et al.
Published: (2024)
by: Wan, David, et al.
Published: (2024)
Variation in Verification: Understanding Verification Dynamics in Large Language Models
by: Zhou, Yefan, et al.
Published: (2025)
by: Zhou, Yefan, et al.
Published: (2025)
NAACL2025 Tutorial: Adaptation of Large Language Models
by: Ke, Zixuan, et al.
Published: (2025)
by: Ke, Zixuan, et al.
Published: (2025)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
by: Liu, Hongjun, et al.
Published: (2025)
by: Liu, Hongjun, et al.
Published: (2025)
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
by: Hu, Tiansheng, et al.
Published: (2025)
by: Hu, Tiansheng, et al.
Published: (2025)
Evaluating Psychological Safety of Large Language Models
by: Li, Xingxuan, et al.
Published: (2022)
by: Li, Xingxuan, et al.
Published: (2022)
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
by: Li, Chuyuan, et al.
Published: (2025)
by: Li, Chuyuan, et al.
Published: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
by: Hu, Tiansheng, et al.
Published: (2026)
by: Hu, Tiansheng, et al.
Published: (2026)
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
On the Benefits of Fine-Grained Loss Truncation: A Case Study on Factuality in Summarization
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2024)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2024)
FOLIO: Natural Language Reasoning with First-Order Logic
by: Han, Simeng, et al.
Published: (2022)
by: Han, Simeng, et al.
Published: (2022)
YaleNLP @ PerAnsSumm 2025: Multi-Perspective Integration via Mixture-of-Agents for Enhanced Healthcare QA Summarization
by: Jang, Dongsuk, et al.
Published: (2025)
by: Jang, Dongsuk, et al.
Published: (2025)
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
by: Wang, Chengye, et al.
Published: (2025)
by: Wang, Chengye, et al.
Published: (2025)
BootPIG: Bootstrapping Zero-shot Personalized Image Generation Capabilities in Pretrained Diffusion Models
by: Purushwalkam, Senthil, et al.
Published: (2024)
by: Purushwalkam, Senthil, et al.
Published: (2024)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
DnA-Eval: Enhancing Large Language Model Evaluation through Decomposition and Aggregation
by: Li, Minzhi, et al.
Published: (2024)
by: Li, Minzhi, et al.
Published: (2024)
Similar Items
-
On Learning to Summarize with Large Language Models as References
by: Liu, Yixin, et al.
Published: (2023) -
ReIFE: Re-evaluating Instruction-Following Evaluation
by: Liu, Yixin, et al.
Published: (2024) -
References Improve LLM Alignment in Non-Verifiable Domains
by: Shi, Kejian, et al.
Published: (2026) -
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025) -
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
by: Han, Simeng, et al.
Published: (2024)