DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramnath, Sahana, Chitsazan, Nima, Zhou, Mingyang, Lee, Chia-Hsuan, Zhang, Shi-Xiong, Rawls, Stephen, Sahu, Sambit, Cho, Sangwoo, Ren, Xiang, Winata, Genta Indra, Veldanda, Akshaj Kumar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024)
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
von: Horoi, Stefan, et al.
Veröffentlicht: (2025)
von: Horoi, Stefan, et al.
Veröffentlicht: (2025)
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
von: Zhao, Hanyang, et al.
Veröffentlicht: (2024)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2024)
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
von: Chakraborty, Amartya, et al.
Veröffentlicht: (2025)
von: Chakraborty, Amartya, et al.
Veröffentlicht: (2025)
MINERS: Multilingual Language Models as Semantic Retrievers
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
von: Mircea, Andrei, et al.
Veröffentlicht: (2025)
von: Mircea, Andrei, et al.
Veröffentlicht: (2025)
SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
von: Kapusuzoglu, Berkcan, et al.
Veröffentlicht: (2025)
von: Kapusuzoglu, Berkcan, et al.
Veröffentlicht: (2025)
What Causes Knowledge Loss in Multilingual Language Models?
von: Khelli, Maria, et al.
Veröffentlicht: (2025)
von: Khelli, Maria, et al.
Veröffentlicht: (2025)
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
von: Niu, Tianyi, et al.
Veröffentlicht: (2026)
von: Niu, Tianyi, et al.
Veröffentlicht: (2026)
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
von: Kuwanto, Garry, et al.
Veröffentlicht: (2024)
von: Kuwanto, Garry, et al.
Veröffentlicht: (2024)
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
von: Wiyono, Vanessa Rebecca, et al.
Veröffentlicht: (2025)
von: Wiyono, Vanessa Rebecca, et al.
Veröffentlicht: (2025)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
von: Hudi, Frederikus, et al.
Veröffentlicht: (2025)
von: Hudi, Frederikus, et al.
Veröffentlicht: (2025)
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
von: Merin, Adril Putra, et al.
Veröffentlicht: (2026)
von: Merin, Adril Putra, et al.
Veröffentlicht: (2026)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2024)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2024)
SPARC-RAG: Adaptive Sequential-Parallel Scaling with Context Management for Retrieval-Augmented Generation
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
Hyper-parameter Tuning for Fair Classification without Sensitive Attribute Access
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2023)
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2023)
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
von: Anugraha, David, et al.
Veröffentlicht: (2024)
von: Anugraha, David, et al.
Veröffentlicht: (2024)
Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups
von: Veldanda, Gautam
Veröffentlicht: (2026)
von: Veldanda, Gautam
Veröffentlicht: (2026)
Lessons from the Field: An Adaptable Lifecycle Approach to Applied Dialogue Summarization
von: Chawla, Kushal, et al.
Veröffentlicht: (2026)
von: Chawla, Kushal, et al.
Veröffentlicht: (2026)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
von: Adilazuarda, Muhammad Farid, et al.
Veröffentlicht: (2024)
von: Adilazuarda, Muhammad Farid, et al.
Veröffentlicht: (2024)
Critique-Guided Distillation for Robust Reasoning via Refinement
von: Kapusuzoglu, Berkcan, et al.
Veröffentlicht: (2025)
von: Kapusuzoglu, Berkcan, et al.
Veröffentlicht: (2025)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
von: Anugraha, David, et al.
Veröffentlicht: (2025)
von: Anugraha, David, et al.
Veröffentlicht: (2025)
ProxyLM: Predicting Language Model Performance on Multilingual Tasks via Proxy Models
von: Anugraha, David, et al.
Veröffentlicht: (2024)
von: Anugraha, David, et al.
Veröffentlicht: (2024)
PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues
von: Farhansyah, Mohammad Rifqi, et al.
Veröffentlicht: (2026)
von: Farhansyah, Mohammad Rifqi, et al.
Veröffentlicht: (2026)
DIAL: Direct Iterative Adversarial Learning for Realistic Multi-Turn Dialogue Simulation
von: Zhu, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhu, Ziyi, et al.
Veröffentlicht: (2025)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?
von: Lee, Chia-Hsuan, et al.
Veröffentlicht: (2026)
von: Lee, Chia-Hsuan, et al.
Veröffentlicht: (2026)
Do Language Models Understand Honorific Systems in Javanese?
von: Farhansyah, Mohammad Rifqi, et al.
Veröffentlicht: (2025)
von: Farhansyah, Mohammad Rifqi, et al.
Veröffentlicht: (2025)
CAVE: Controllable Authorship Verification Explanations
von: Ramnath, Sahana, et al.
Veröffentlicht: (2024)
von: Ramnath, Sahana, et al.
Veröffentlicht: (2024)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
von: Anugraha, David, et al.
Veröffentlicht: (2025)
von: Anugraha, David, et al.
Veröffentlicht: (2025)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
von: Elsetohy, Alaa, et al.
Veröffentlicht: (2026)
von: Elsetohy, Alaa, et al.
Veröffentlicht: (2026)
Continual Pre-training of MoEs: How robust is your router?
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
A Multi-Agent Dual Dialogue System to Support Mental Health Care Providers
von: Kampman, Onno P., et al.
Veröffentlicht: (2024)
von: Kampman, Onno P., et al.
Veröffentlicht: (2024)
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
von: Zhao, Hanyang, et al.
Veröffentlicht: (2025)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2025)
Vision Language Models are Confused Tourists
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2025)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2025)
SUMMER READERS.
von: DOWNING, VIRGINIA
Veröffentlicht: (1967)
von: DOWNING, VIRGINIA
Veröffentlicht: (1967)
Ähnliche Einträge
-
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024) -
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
von: Horoi, Stefan, et al.
Veröffentlicht: (2025) -
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
von: Zhao, Bo, et al.
Veröffentlicht: (2025) -
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
von: Zhao, Hanyang, et al.
Veröffentlicht: (2024) -
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)