Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramprasad, Sanjana, Krishna, Kundan, Lipton, Zachary C, Wallace, Byron C |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
von: Krishna, Kundan, et al.
Veröffentlicht: (2024)
von: Krishna, Kundan, et al.
Veröffentlicht: (2024)
Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Zero-shot Factual Consistency Evaluation Across Domains
von: Agarwal, Raunak
Veröffentlicht: (2024)
von: Agarwal, Raunak
Veröffentlicht: (2024)
LLM-Select: Feature Selection with Large Language Models
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
von: Mujahid, Zain Muhammad, et al.
Veröffentlicht: (2025)
von: Mujahid, Zain Muhammad, et al.
Veröffentlicht: (2025)
Personalized Language Modeling from Personalized Human Feedback
von: Li, Xinyu, et al.
Veröffentlicht: (2024)
von: Li, Xinyu, et al.
Veröffentlicht: (2024)
Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
Entity-level Factual Adaptiveness of Fine-tuning based Abstractive Summarization Models
von: Song, Jongyoon, et al.
Veröffentlicht: (2024)
von: Song, Jongyoon, et al.
Veröffentlicht: (2024)
The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
von: Krishna, Kundan, et al.
Veröffentlicht: (2025)
von: Krishna, Kundan, et al.
Veröffentlicht: (2025)
Less is More for Improving Automatic Evaluation of Factual Consistency
von: Wang, Tong, et al.
Veröffentlicht: (2024)
von: Wang, Tong, et al.
Veröffentlicht: (2024)
Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot
von: Cheng, Xiang, et al.
Veröffentlicht: (2025)
von: Cheng, Xiang, et al.
Veröffentlicht: (2025)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Amrita, et al.
Veröffentlicht: (2024)
The Factuality of Large Language Models in the Legal Domain
von: Hamdani, Rajaa El, et al.
Veröffentlicht: (2024)
von: Hamdani, Rajaa El, et al.
Veröffentlicht: (2024)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
von: Yuan, Jiaqing, et al.
Veröffentlicht: (2024)
von: Yuan, Jiaqing, et al.
Veröffentlicht: (2024)
Policies and Evaluation for Online Meeting Summarization
von: Schneider, Felix, et al.
Veröffentlicht: (2025)
von: Schneider, Felix, et al.
Veröffentlicht: (2025)
DIVERS-Bench: Evaluating Language Identification Across Domain Shifts and Code-Switching
von: Ojo, Jessica, et al.
Veröffentlicht: (2025)
von: Ojo, Jessica, et al.
Veröffentlicht: (2025)
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question Answering
von: Ko, Sungho, et al.
Veröffentlicht: (2024)
von: Ko, Sungho, et al.
Veröffentlicht: (2024)
Evaluating the Factuality of Large Language Models using Large-Scale Knowledge Graphs
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks
von: Labrak, Yanis, et al.
Veröffentlicht: (2023)
von: Labrak, Yanis, et al.
Veröffentlicht: (2023)
From Haystack to Needle: Label Space Reduction for Zero-shot Classification
von: Vandemoortele, Nathan, et al.
Veröffentlicht: (2025)
von: Vandemoortele, Nathan, et al.
Veröffentlicht: (2025)
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
von: Shin, Hagyeong, et al.
Veröffentlicht: (2025)
von: Shin, Hagyeong, et al.
Veröffentlicht: (2025)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
von: Wang, Taowen, et al.
Veröffentlicht: (2024)
von: Wang, Taowen, et al.
Veröffentlicht: (2024)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
PathCoT: Chain-of-Thought Prompting for Zero-shot Pathology Visual Reasoning
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
Towards Reducing Diagnostic Errors with Interpretable Risk Prediction
von: McInerney, Denis Jered, et al.
Veröffentlicht: (2024)
von: McInerney, Denis Jered, et al.
Veröffentlicht: (2024)
Language Models with Conformal Factuality Guarantees
von: Mohri, Christopher, et al.
Veröffentlicht: (2024)
von: Mohri, Christopher, et al.
Veröffentlicht: (2024)
Utilizing GPT to Enhance Text Summarization: A Strategy to Minimize Hallucinations
von: Shakil, Hassan, et al.
Veröffentlicht: (2024)
von: Shakil, Hassan, et al.
Veröffentlicht: (2024)
DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2025)
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2025)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
von: Baidya, Avinash, et al.
Veröffentlicht: (2025)
von: Baidya, Avinash, et al.
Veröffentlicht: (2025)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
von: Wang, Duo, et al.
Veröffentlicht: (2024)
von: Wang, Duo, et al.
Veröffentlicht: (2024)
Factuality Challenges in the Era of Large Language Models
von: Augenstein, Isabelle, et al.
Veröffentlicht: (2023)
von: Augenstein, Isabelle, et al.
Veröffentlicht: (2023)
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
von: Wang, Chao, et al.
Veröffentlicht: (2024)
von: Wang, Chao, et al.
Veröffentlicht: (2024)
Zero-shot data citation function classification using transformer-based large language models (LLMs)
von: Byers, Neil, et al.
Veröffentlicht: (2025)
von: Byers, Neil, et al.
Veröffentlicht: (2025)
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024) -
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
von: Krishna, Kundan, et al.
Veröffentlicht: (2024) -
Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024) -
Zero-shot Factual Consistency Evaluation Across Domains
von: Agarwal, Raunak
Veröffentlicht: (2024) -
LLM-Select: Feature Selection with Large Language Models
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)