Rethinking Scientific Summarization Evaluation: Grounding Explainable Metrics on Facet-aware Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xiuying, Wang, Tairan, Zhu, Qingqing, Guo, Taicheng, Gao, Shen, Lu, Zhiyong, Gao, Xin, Zhang, Xiangliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Write Summary Step-by-Step: A Pilot Study of Stepwise Summarization
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
Flexible and Adaptable Summarization via Expertise Separation
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
Evaluating and Mitigating Bias in AI-Based Medical Text Generation
von: Chen, Xiuying, et al.
Veröffentlicht: (2025)
von: Chen, Xiuying, et al.
Veröffentlicht: (2025)
ScholarChemQA: Unveiling the Power of Language Models in Chemical Research Question Answering
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
SenWave: A Fine-Grained Multi-Language Sentiment Analysis Dataset Sourced from COVID-19 Tweets
von: Yang, Qiang, et al.
Veröffentlicht: (2025)
von: Yang, Qiang, et al.
Veröffentlicht: (2025)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
von: Gao, Lang, et al.
Veröffentlicht: (2024)
von: Gao, Lang, et al.
Veröffentlicht: (2024)
Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports
von: Zhu, Qingqing, et al.
Veröffentlicht: (2024)
von: Zhu, Qingqing, et al.
Veröffentlicht: (2024)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
von: Guo, Taicheng, et al.
Veröffentlicht: (2026)
von: Guo, Taicheng, et al.
Veröffentlicht: (2026)
Large Language Model based Multi-Agents: A Survey of Progress and Challenges
von: Guo, Taicheng, et al.
Veröffentlicht: (2024)
von: Guo, Taicheng, et al.
Veröffentlicht: (2024)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
von: Guo, Yue, et al.
Veröffentlicht: (2023)
von: Guo, Yue, et al.
Veröffentlicht: (2023)
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
von: Feng, Ruixiang, et al.
Veröffentlicht: (2025)
von: Feng, Ruixiang, et al.
Veröffentlicht: (2025)
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization
von: Mondshine, Itai, et al.
Veröffentlicht: (2025)
von: Mondshine, Itai, et al.
Veröffentlicht: (2025)
PerSphere: A Comprehensive Framework for Multi-Faceted Perspective Retrieval and Summarization
von: Luo, Yun, et al.
Veröffentlicht: (2024)
von: Luo, Yun, et al.
Veröffentlicht: (2024)
MTSQL-R1: Towards Long-Horizon Multi-Turn Text-to-SQL via Agentic Training
von: Guo, Taicheng, et al.
Veröffentlicht: (2025)
von: Guo, Taicheng, et al.
Veröffentlicht: (2025)
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
von: Zhuang, Haomin, et al.
Veröffentlicht: (2025)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2025)
Towards Explainable Evaluation Metrics for Machine Translation
von: Leiter, Christoph, et al.
Veröffentlicht: (2023)
von: Leiter, Christoph, et al.
Veröffentlicht: (2023)
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
CCSBench: Evaluating Compositional Controllability in LLMs for Scientific Document Summarization
von: Ding, Yixi, et al.
Veröffentlicht: (2024)
von: Ding, Yixi, et al.
Veröffentlicht: (2024)
What Are They Talking About? A Benchmark of Knowledge-Grounded Discussion Summarization
von: Zhou, Weixiao, et al.
Veröffentlicht: (2025)
von: Zhou, Weixiao, et al.
Veröffentlicht: (2025)
CASPR: Automated Evaluation Metric for Contrastive Summarization
von: Ananthamurugan, Nirupan, et al.
Veröffentlicht: (2024)
von: Ananthamurugan, Nirupan, et al.
Veröffentlicht: (2024)
Calibrating Model-Based Evaluation Metrics for Summarization
von: Liu, Hongye, et al.
Veröffentlicht: (2026)
von: Liu, Hongye, et al.
Veröffentlicht: (2026)
Selecting Query-bag as Pseudo Relevance Feedback for Information-seeking Conversations
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
Scientific Opinion Summarization: Paper Meta-review Generation Dataset, Methods, and Evaluation
von: Zeng, Qi, et al.
Veröffentlicht: (2023)
von: Zeng, Qi, et al.
Veröffentlicht: (2023)
MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
PlainQAFact: Retrieval-augmented Factual Consistency Evaluation Metric for Biomedical Plain Language Summarization
von: You, Zhiwen, et al.
Veröffentlicht: (2025)
von: You, Zhiwen, et al.
Veröffentlicht: (2025)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
How Well Do Multi-modal LLMs Interpret CT Scans? An Auto-Evaluation Framework for Analyses
von: Zhu, Qingqing, et al.
Veröffentlicht: (2024)
von: Zhu, Qingqing, et al.
Veröffentlicht: (2024)
SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research
von: Sun, Liangtai, et al.
Veröffentlicht: (2023)
von: Sun, Liangtai, et al.
Veröffentlicht: (2023)
Anagent For Enhancing Scientific Table & Figure Analysis
von: Guo, Xuehang, et al.
Veröffentlicht: (2026)
von: Guo, Xuehang, et al.
Veröffentlicht: (2026)
Beyond Profile: From Surface-Level Facts to Deep Persona Simulation in LLMs
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
von: Muscato, Benedetta, et al.
Veröffentlicht: (2026)
von: Muscato, Benedetta, et al.
Veröffentlicht: (2026)
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature
von: Shen, Lingdong, et al.
Veröffentlicht: (2024)
von: Shen, Lingdong, et al.
Veröffentlicht: (2024)
Towards Region-aware Bias Evaluation Metrics
von: Borah, Angana, et al.
Veröffentlicht: (2024)
von: Borah, Angana, et al.
Veröffentlicht: (2024)
Large Language Models and Book Summarization: Reading or Remembering, Which Is Better?
von: Fu, Tairan, et al.
Veröffentlicht: (2026)
von: Fu, Tairan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Write Summary Step-by-Step: A Pilot Study of Stepwise Summarization
von: Chen, Xiuying, et al.
Veröffentlicht: (2024) -
Flexible and Adaptable Summarization via Expertise Separation
von: Chen, Xiuying, et al.
Veröffentlicht: (2024) -
Evaluating and Mitigating Bias in AI-Based Medical Text Generation
von: Chen, Xiuying, et al.
Veröffentlicht: (2025) -
ScholarChemQA: Unveiling the Power of Language Models in Chemical Research Question Answering
von: Chen, Xiuying, et al.
Veröffentlicht: (2024) -
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)