LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Huyen, Zhang, Haoxuan, Zhang, Yang, Chen, Haihua, Ding, Junhua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation
by: Nguyen, Huyen, et al.
Published: (2026)
by: Nguyen, Huyen, et al.
Published: (2026)
Scaling Laws for State Dynamics in Large Language Models
by: Li, Jacob X, et al.
Published: (2025)
by: Li, Jacob X, et al.
Published: (2025)
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
by: Idahl, Maximilian, et al.
Published: (2026)
by: Idahl, Maximilian, et al.
Published: (2026)
AskSport: Web Application for Sports Question-Answering
by: Onofre, Enzo B, et al.
Published: (2025)
by: Onofre, Enzo B, et al.
Published: (2025)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
by: Dayarathne, Ranul, et al.
Published: (2025)
by: Dayarathne, Ranul, et al.
Published: (2025)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
A Graph-based RAG for Energy Efficiency Question Answering
by: Campi, Riccardo, et al.
Published: (2025)
by: Campi, Riccardo, et al.
Published: (2025)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
by: Han, Lifeng, et al.
Published: (2016)
by: Han, Lifeng, et al.
Published: (2016)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
by: Zhang, Luyan, et al.
Published: (2025)
by: Zhang, Luyan, et al.
Published: (2025)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
by: Drchal, Jan, et al.
Published: (2023)
by: Drchal, Jan, et al.
Published: (2023)
Automating Clinical Information Retrieval from Finnish Electronic Health Records Using Large Language Models
by: Saukkoriipi, Mikko, et al.
Published: (2026)
by: Saukkoriipi, Mikko, et al.
Published: (2026)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
by: Hashemi, Helia, et al.
Published: (2024)
by: Hashemi, Helia, et al.
Published: (2024)
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
by: Hildebrand, Samuel, et al.
Published: (2025)
by: Hildebrand, Samuel, et al.
Published: (2025)
CR-LT-KGQA: A Knowledge Graph Question Answering Dataset Requiring Commonsense Reasoning and Long-Tail Knowledge
by: Guo, Willis, et al.
Published: (2024)
by: Guo, Willis, et al.
Published: (2024)
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
by: Zhang, Yiqing, et al.
Published: (2026)
by: Zhang, Yiqing, et al.
Published: (2026)
QuAnTS: Question Answering on Time Series
by: Divo, Felix, et al.
Published: (2025)
by: Divo, Felix, et al.
Published: (2025)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
by: Karpurapu, Shanthi, et al.
Published: (2024)
by: Karpurapu, Shanthi, et al.
Published: (2024)
Towards Conditioning Clinical Text Generation for User Control
by: Koraş, Osman Alperen, et al.
Published: (2025)
by: Koraş, Osman Alperen, et al.
Published: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
by: Zhu, Qian, et al.
Published: (2026)
by: Zhu, Qian, et al.
Published: (2026)
Introducing Brain-like Concepts to Embodied Hand-crafted Dialog Management System
by: Joublin, Frank, et al.
Published: (2024)
by: Joublin, Frank, et al.
Published: (2024)
Open-TI: Open Traffic Intelligence with Augmented Language Model
by: Da, Longchao, et al.
Published: (2023)
by: Da, Longchao, et al.
Published: (2023)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
by: Wu, Dekun, et al.
Published: (2023)
by: Wu, Dekun, et al.
Published: (2023)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
by: Li, Jianing, et al.
Published: (2024)
by: Li, Jianing, et al.
Published: (2024)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026)
by: Kuric, Eduard, et al.
Published: (2026)
From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
by: Gill, Gurbinder, et al.
Published: (2025)
by: Gill, Gurbinder, et al.
Published: (2025)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
by: Ghandi, Taraneh, et al.
Published: (2026)
by: Ghandi, Taraneh, et al.
Published: (2026)
elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings
by: Osborne, Philip, et al.
Published: (2025)
by: Osborne, Philip, et al.
Published: (2025)
Calibrated Confidence Estimation for Tabular Question Answering
by: Voss, Lukas
Published: (2026)
by: Voss, Lukas
Published: (2026)
How to Evaluate Medical AI
by: Kopanichuk, Ilia, et al.
Published: (2025)
by: Kopanichuk, Ilia, et al.
Published: (2025)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
by: Rodriguez, David, et al.
Published: (2025)
by: Rodriguez, David, et al.
Published: (2025)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
by: Borisov, Vadim
Published: (2026)
by: Borisov, Vadim
Published: (2026)
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
by: Hosseini, Peyman, et al.
Published: (2025)
by: Hosseini, Peyman, et al.
Published: (2025)
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
by: Wright, Devin R., et al.
Published: (2026)
by: Wright, Devin R., et al.
Published: (2026)
Complementarity, Augmentation, or Substitutivity? The Impact of Generative Artificial Intelligence on the U.S. Federal Workforce
by: Resh, William G., et al.
Published: (2025)
by: Resh, William G., et al.
Published: (2025)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
by: Chien, Sheng-You, et al.
Published: (2026)
by: Chien, Sheng-You, et al.
Published: (2026)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
by: Hasan, Mohammed Rakibul
Published: (2026)
by: Hasan, Mohammed Rakibul
Published: (2026)
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
by: Pandey, Anupam, et al.
Published: (2025)
by: Pandey, Anupam, et al.
Published: (2025)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
by: Fan, Lin, et al.
Published: (2026)
by: Fan, Lin, et al.
Published: (2026)
Similar Items
-
LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation
by: Nguyen, Huyen, et al.
Published: (2026) -
Scaling Laws for State Dynamics in Large Language Models
by: Li, Jacob X, et al.
Published: (2025) -
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
by: Idahl, Maximilian, et al.
Published: (2026) -
AskSport: Web Application for Sports Question-Answering
by: Onofre, Enzo B, et al.
Published: (2025) -
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
by: Dayarathne, Ranul, et al.
Published: (2025)