Factual consistency evaluation of summarization in the Era of large language models
Fuente:
arXiv
Salvato in:
| Autori principali: | Luo, Zheheng, Xie, Qianqian, Ananiadou, Sophia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Lay Person's Guide to Biomedicine: Orchestrating Large Language Models
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
Are Large Language Models True Healthcare Jacks-of-All-Trades? Benchmarking Across Health Professions Beyond Physician Exams
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
di: Bishop, Jennifer A, et al.
Pubblicazione: (2023)
di: Bishop, Jennifer A, et al.
Pubblicazione: (2023)
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
di: Scirè, Alessandro, et al.
Pubblicazione: (2024)
di: Scirè, Alessandro, et al.
Pubblicazione: (2024)
The current status of large language models in summarizing radiology report impressions
di: Hu, Danqing, et al.
Pubblicazione: (2024)
di: Hu, Danqing, et al.
Pubblicazione: (2024)
EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective Analysis
di: Liu, Zhiwei, et al.
Pubblicazione: (2024)
di: Liu, Zhiwei, et al.
Pubblicazione: (2024)
Closing the gap between open-source and commercial large language models for medical evidence summarization
di: Zhang, Gongbo, et al.
Pubblicazione: (2024)
di: Zhang, Gongbo, et al.
Pubblicazione: (2024)
Leveraging language models for summarizing mental state examinations: A comprehensive evaluation and dataset release
di: Sahu, Nilesh Kumar, et al.
Pubblicazione: (2024)
di: Sahu, Nilesh Kumar, et al.
Pubblicazione: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
di: Aali, Asad, et al.
Pubblicazione: (2024)
di: Aali, Asad, et al.
Pubblicazione: (2024)
FMDLlama: Financial Misinformation Detection based on Large Language Models
di: Liu, Zhiwei, et al.
Pubblicazione: (2024)
di: Liu, Zhiwei, et al.
Pubblicazione: (2024)
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models
di: Yang, Kailai, et al.
Pubblicazione: (2023)
di: Yang, Kailai, et al.
Pubblicazione: (2023)
Towards Interpretable Mental Health Analysis with Large Language Models
di: Yang, Kailai, et al.
Pubblicazione: (2023)
di: Yang, Kailai, et al.
Pubblicazione: (2023)
Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis
di: Yu, Zeping, et al.
Pubblicazione: (2024)
di: Yu, Zeping, et al.
Pubblicazione: (2024)
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
di: Yu, Zeping, et al.
Pubblicazione: (2024)
di: Yu, Zeping, et al.
Pubblicazione: (2024)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
di: Jung, Jee-weon, et al.
Pubblicazione: (2024)
di: Jung, Jee-weon, et al.
Pubblicazione: (2024)
RAEmoLLM: Retrieval Augmented LLMs for Cross-Domain Misinformation Detection Using In-Context Learning Based on Emotional Information
di: Liu, Zhiwei, et al.
Pubblicazione: (2024)
di: Liu, Zhiwei, et al.
Pubblicazione: (2024)
MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models
di: Yang, Kailai, et al.
Pubblicazione: (2024)
di: Yang, Kailai, et al.
Pubblicazione: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
Event-based evaluation of abstractive news summarization
di: You, Huiling, et al.
Pubblicazione: (2025)
di: You, Huiling, et al.
Pubblicazione: (2025)
How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning
di: Yu, Zeping, et al.
Pubblicazione: (2024)
di: Yu, Zeping, et al.
Pubblicazione: (2024)
Neuron-Level Knowledge Attribution in Large Language Models
di: Yu, Zeping, et al.
Pubblicazione: (2023)
di: Yu, Zeping, et al.
Pubblicazione: (2023)
RAAR: Retrieval Augmented Agentic Reasoning for Cross-Domain Misinformation Detection
di: Liu, Zhiwei, et al.
Pubblicazione: (2026)
di: Liu, Zhiwei, et al.
Pubblicazione: (2026)
Selective Preference Optimization via Token-Level Reward Function Estimation
di: Yang, Kailai, et al.
Pubblicazione: (2024)
di: Yang, Kailai, et al.
Pubblicazione: (2024)
ConspEmoLLM-v2: A robust and stable model to detect sentiment-transformed conspiracy theories
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
Disentangled VAD Representations via a Variational Framework for Political Stance Detection
di: Xu, Beiyu, et al.
Pubblicazione: (2025)
di: Xu, Beiyu, et al.
Pubblicazione: (2025)
Re-evaluating Theory of Mind evaluation in large language models
di: Hu, Jennifer, et al.
Pubblicazione: (2025)
di: Hu, Jennifer, et al.
Pubblicazione: (2025)
MisSpans: Fine-Grained False Span Identification in Cross-Domain Fake News
di: Liu, Zhiwei, et al.
Pubblicazione: (2026)
di: Liu, Zhiwei, et al.
Pubblicazione: (2026)
Text2Model: Generating dynamic chemical reactor models using large language models (LLMs)
di: Rupprecht, Sophia, et al.
Pubblicazione: (2025)
di: Rupprecht, Sophia, et al.
Pubblicazione: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images
di: Jiang, Yuechen, et al.
Pubblicazione: (2026)
di: Jiang, Yuechen, et al.
Pubblicazione: (2026)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
di: Wen, Zichen, et al.
Pubblicazione: (2024)
di: Wen, Zichen, et al.
Pubblicazione: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
di: Li, Kunning, et al.
Pubblicazione: (2025)
di: Li, Kunning, et al.
Pubblicazione: (2025)
QFS-Composer: Query-focused summarization pipeline for less resourced languages
di: Đuranović, Vuk, et al.
Pubblicazione: (2026)
di: Đuranović, Vuk, et al.
Pubblicazione: (2026)
Rumor Detection by Multi-task Suffix Learning based on Time-series Dual Sentiments
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
Uncovering inequalities in new knowledge learning by large language models across different languages
di: Wang, Chenglong, et al.
Pubblicazione: (2025)
di: Wang, Chenglong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Lay Person's Guide to Biomedicine: Orchestrating Large Language Models
di: Luo, Zheheng, et al.
Pubblicazione: (2024) -
Are Large Language Models True Healthcare Jacks-of-All-Trades? Benchmarking Across Health Professions Beyond Physician Exams
di: Luo, Zheheng, et al.
Pubblicazione: (2024) -
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
di: Bishop, Jennifer A, et al.
Pubblicazione: (2023) -
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
di: Scirè, Alessandro, et al.
Pubblicazione: (2024) -
The current status of large language models in summarizing radiology report impressions
di: Hu, Danqing, et al.
Pubblicazione: (2024)