FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Scirè, Alessandro, Ghonim, Karim, Navigli, Roberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Truth or Mirage? Towards End-to-End Factuality Evaluation with LLM-Oasis
von: Scirè, Alessandro, et al.
Veröffentlicht: (2024)
von: Scirè, Alessandro, et al.
Veröffentlicht: (2024)
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
von: Perrella, Stefano, et al.
Veröffentlicht: (2024)
von: Perrella, Stefano, et al.
Veröffentlicht: (2024)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
Towards Effective Extraction and Evaluation of Factual Claims
von: Metropolitansky, Dasha, et al.
Veröffentlicht: (2025)
von: Metropolitansky, Dasha, et al.
Veröffentlicht: (2025)
Factual consistency evaluation of summarization in the Era of large language models
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
Merging Facts, Crafting Fallacies: Evaluating the Contradictory Nature of Aggregated Factual Claims in Long-Form Generations
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2024)
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2024)
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
von: Pisano, Raffaele, et al.
Veröffentlicht: (2026)
von: Pisano, Raffaele, et al.
Veröffentlicht: (2026)
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
von: Gatti, Bruno, et al.
Veröffentlicht: (2026)
von: Gatti, Bruno, et al.
Veröffentlicht: (2026)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
von: Bonomo, Tommaso, et al.
Veröffentlicht: (2025)
von: Bonomo, Tommaso, et al.
Veröffentlicht: (2025)
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2025)
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2025)
All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations
von: Wanner, Miriam, et al.
Veröffentlicht: (2025)
von: Wanner, Miriam, et al.
Veröffentlicht: (2025)
The current status of large language models in summarizing radiology report impressions
von: Hu, Danqing, et al.
Veröffentlicht: (2024)
von: Hu, Danqing, et al.
Veröffentlicht: (2024)
FuDoBa: Fusing Document and Knowledge Graph-based Representations with Bayesian Optimisation
von: Koloski, Boshko, et al.
Veröffentlicht: (2025)
von: Koloski, Boshko, et al.
Veröffentlicht: (2025)
Core: Robust Factual Precision with Informative Sub-Claim Identification
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
ReLiK: Retrieve and LinK, Fast and Accurate Entity Linking and Relation Extraction on an Academic Budget
von: Orlando, Riccardo, et al.
Veröffentlicht: (2024)
von: Orlando, Riccardo, et al.
Veröffentlicht: (2024)
Event-based evaluation of abstractive news summarization
von: You, Huiling, et al.
Veröffentlicht: (2025)
von: You, Huiling, et al.
Veröffentlicht: (2025)
AutoML-guided Fusion of Entity and LLM-based Representations for Document Classification
von: Koloski, Boshko, et al.
Veröffentlicht: (2024)
von: Koloski, Boshko, et al.
Veröffentlicht: (2024)
Maverick: Efficient and Accurate Coreference Resolution Defying Recent Trends
von: Martinelli, Giuliano, et al.
Veröffentlicht: (2024)
von: Martinelli, Giuliano, et al.
Veröffentlicht: (2024)
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
von: Huang, Heyuan, et al.
Veröffentlicht: (2025)
von: Huang, Heyuan, et al.
Veröffentlicht: (2025)
QFS-Composer: Query-focused summarization pipeline for less resourced languages
von: Đuranović, Vuk, et al.
Veröffentlicht: (2026)
von: Đuranović, Vuk, et al.
Veröffentlicht: (2026)
JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation
von: Xu, Fan, et al.
Veröffentlicht: (2025)
von: Xu, Fan, et al.
Veröffentlicht: (2025)
Leveraging language models for summarizing mental state examinations: A comprehensive evaluation and dataset release
von: Sahu, Nilesh Kumar, et al.
Veröffentlicht: (2024)
von: Sahu, Nilesh Kumar, et al.
Veröffentlicht: (2024)
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
von: Shin, Hagyeong, et al.
Veröffentlicht: (2025)
von: Shin, Hagyeong, et al.
Veröffentlicht: (2025)
Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency
von: Morbiato, Filippo, et al.
Veröffentlicht: (2025)
von: Morbiato, Filippo, et al.
Veröffentlicht: (2025)
Understanding Finetuning for Factual Knowledge Extraction
von: Ghosal, Gaurav, et al.
Veröffentlicht: (2024)
von: Ghosal, Gaurav, et al.
Veröffentlicht: (2024)
FIBER: A Multilingual Evaluation Resource for Factual Inference Bias
von: Munis, Evren Ayberk, et al.
Veröffentlicht: (2025)
von: Munis, Evren Ayberk, et al.
Veröffentlicht: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
von: Aali, Asad, et al.
Veröffentlicht: (2024)
von: Aali, Asad, et al.
Veröffentlicht: (2024)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
von: Ni, Jingwei, et al.
Veröffentlicht: (2024)
von: Ni, Jingwei, et al.
Veröffentlicht: (2024)
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
von: Huang, Minghui
Veröffentlicht: (2025)
von: Huang, Minghui
Veröffentlicht: (2025)
VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
von: Han, Soyeon Caren, et al.
Veröffentlicht: (2024)
von: Han, Soyeon Caren, et al.
Veröffentlicht: (2024)
ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2024)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2024)
Closing the gap between open-source and commercial large language models for medical evidence summarization
von: Zhang, Gongbo, et al.
Veröffentlicht: (2024)
von: Zhang, Gongbo, et al.
Veröffentlicht: (2024)
FABLES: Evaluating faithfulness and content selection in book-length summarization
von: Kim, Yekyung, et al.
Veröffentlicht: (2024)
von: Kim, Yekyung, et al.
Veröffentlicht: (2024)
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
von: Perrella, Stefano, et al.
Veröffentlicht: (2024)
von: Perrella, Stefano, et al.
Veröffentlicht: (2024)
Document-level Claim Extraction and Decontextualisation for Fact-Checking
von: Deng, Zhenyun, et al.
Veröffentlicht: (2024)
von: Deng, Zhenyun, et al.
Veröffentlicht: (2024)
Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
von: Ye, Yuxuan, et al.
Veröffentlicht: (2026)
von: Ye, Yuxuan, et al.
Veröffentlicht: (2026)
Improving Model Factuality with Fine-grained Critique-based Evaluator
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Truth or Mirage? Towards End-to-End Factuality Evaluation with LLM-Oasis
von: Scirè, Alessandro, et al.
Veröffentlicht: (2024) -
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
von: Perrella, Stefano, et al.
Veröffentlicht: (2024) -
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025) -
Towards Effective Extraction and Evaluation of Factual Claims
von: Metropolitansky, Dasha, et al.
Veröffentlicht: (2025) -
Factual consistency evaluation of summarization in the Era of large language models
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)