Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Min, Hyangsuk, Lee, Yuho, Ban, Minjeong, Deng, Jiaqi, Kim, Nicole Hee-Yeon, Yun, Taewon, Su, Hang, Cai, Jason, Song, Hwanjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
by: Deng, Jiaqi, et al.
Published: (2025)
by: Deng, Jiaqi, et al.
Published: (2025)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024)
by: Lee, Yuho, et al.
Published: (2024)
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
by: Yun, Taewon, et al.
Published: (2025)
by: Yun, Taewon, et al.
Published: (2025)
Learning to Summarize from LLM-generated Feedback
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
by: Ban, Minjeong, et al.
Published: (2026)
by: Ban, Minjeong, et al.
Published: (2026)
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
by: Oh, Jihwan, et al.
Published: (2024)
by: Oh, Jihwan, et al.
Published: (2024)
What Makes a Sale? Rethinking End-to-End Seller--Buyer Retail Dynamics with LLM Agents
by: Choi, Jeonghwan, et al.
Published: (2026)
by: Choi, Jeonghwan, et al.
Published: (2026)
FineSurE: Fine-grained Summarization Evaluation using LLMs
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection
by: He, Jianfeng, et al.
Published: (2024)
by: He, Jianfeng, et al.
Published: (2024)
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding
by: Yun, Taewon, et al.
Published: (2026)
by: Yun, Taewon, et al.
Published: (2026)
LLM-based User Profile Management for Recommender System
by: Bang, Seunghwan, et al.
Published: (2025)
by: Bang, Seunghwan, et al.
Published: (2025)
Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines
by: Song, Hwanjun
Published: (2026)
by: Song, Hwanjun
Published: (2026)
Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence
by: Bang, Seunghwan, et al.
Published: (2026)
by: Bang, Seunghwan, et al.
Published: (2026)
Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights
by: Mao, Shunqi, et al.
Published: (2024)
by: Mao, Shunqi, et al.
Published: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
Rethinking LLM-Based Recommendations: A Personalized Query-Driven Parallel Integration
by: Han, Donghee, et al.
Published: (2025)
by: Han, Donghee, et al.
Published: (2025)
Relationship between perceived human resource management system strength, thriving at work, and employee's turnover intention in Chinese hotels
by: Jing Guo, et al.
Published: (2024)
by: Jing Guo, et al.
Published: (2024)
Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent Encoders
by: Zhang, Yuwei, et al.
Published: (2024)
by: Zhang, Yuwei, et al.
Published: (2024)
Ageism and Voting Behaviour in a Fictitious Election Situation in Japan
by: Yuho Shimizu
Published: (2025)
by: Yuho Shimizu
Published: (2025)
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
by: Song, Hwanjun, et al.
Published: (2025)
by: Song, Hwanjun, et al.
Published: (2025)
Quantile-Free Uncertainty Quantification in Graph Neural Networks
by: park, Soyoung, et al.
Published: (2026)
by: park, Soyoung, et al.
Published: (2026)
Factors Influencing Attitudes Toward Dementia Among Elderly Individuals in the Community
by: Ka Hee Yoo, et al.
Published: (2025)
by: Ka Hee Yoo, et al.
Published: (2025)
Green Supply Chain Management: uma revisão sistemática integrativa dos estudos publicados
by: Gustavo Yuho Endo
Published: (2021)
by: Gustavo Yuho Endo
Published: (2021)
IDENTIFICAÇÃO DO PERFIL DE POTENCIAIS CLIENTES DE SERVIÇOS AMBIENTALMENTE CORRETOS DE UMA OFICINA MECÂNICA
by: Gustavo Yuho Endo
Published: (2016)
by: Gustavo Yuho Endo
Published: (2016)
Self‐efficacy of genetic counselors and genetic counseling students using the Korean genetic counseling self‐efficacy scale
by: In Hee Choi, et al.
Published: (2026)
by: In Hee Choi, et al.
Published: (2026)
Validity and Reliability of the Korean Version of the Self‐Care Self‐Efficacy Scale for Patients With Heart Failure: A Psychometric Evaluation
by: JinShil Kim, et al.
Published: (2025)
by: JinShil Kim, et al.
Published: (2025)
Trajectory-Guided Diffusion for Foreground-Preserving Background Generation in Multi-Layer Documents
by: Kang, Taewon
Published: (2026)
by: Kang, Taewon
Published: (2026)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
by: Yuan, Dong, et al.
Published: (2024)
by: Yuan, Dong, et al.
Published: (2024)
Bipartite quantum states admitting a causal explanation
by: Song, Minjeong, et al.
Published: (2025)
by: Song, Minjeong, et al.
Published: (2025)
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
by: Liu, Yinhong, et al.
Published: (2025)
by: Liu, Yinhong, et al.
Published: (2025)
In-Context Learning with Noisy Labels
by: Kang, Junyong, et al.
Published: (2024)
by: Kang, Junyong, et al.
Published: (2024)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
Consistent truncation and generalized duality based on exceptional generalized cosets
by: Hassler, Falk, et al.
Published: (2025)
by: Hassler, Falk, et al.
Published: (2025)
On quantum Poisson-Lie T-duality of WZNW models
by: Sakatani, Yuho, et al.
Published: (2023)
by: Sakatani, Yuho, et al.
Published: (2023)
All maximal gauged supergravities with uplift
by: Hassler, Falk, et al.
Published: (2022)
by: Hassler, Falk, et al.
Published: (2022)
Exact and approximate conditions of tabletop reversibility: when is Petz recovery cost-free?
by: Song, Minjeong, et al.
Published: (2025)
by: Song, Minjeong, et al.
Published: (2025)
CREMA: A Contrastive Regularized Masked Autoencoder for Robust ECG Diagnostics across Clinical Domains
by: Song, Junho, et al.
Published: (2024)
by: Song, Junho, et al.
Published: (2024)
Analysis of Log Data from an International Online Educational Assessment System: A Multi-state Survival Modeling Approach to Reaction Time between and across Action Sequence
by: Park, Jina, et al.
Published: (2024)
by: Park, Jina, et al.
Published: (2024)
Survey of Query-based Text Summarization
by: Yu, Hang, et al.
Published: (2022)
by: Yu, Hang, et al.
Published: (2022)
GLEAN: Active Generalized Category Discovery with Diverse LLM Feedback
by: Zou, Henry Peng, et al.
Published: (2025)
by: Zou, Henry Peng, et al.
Published: (2025)
Similar Items
-
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
by: Deng, Jiaqi, et al.
Published: (2025) -
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024) -
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
by: Yun, Taewon, et al.
Published: (2025) -
Learning to Summarize from LLM-generated Feedback
by: Song, Hwanjun, et al.
Published: (2024) -
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
by: Ban, Minjeong, et al.
Published: (2026)