Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Dejl, Adam, Barry, James, Pascale, Alessandra, Cano, Javier Carnerero |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
by: Yun, Janghyeon, et al.
Published: (2025)
by: Yun, Janghyeon, et al.
Published: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Argumentative Large Language Models for Explainable and Contestable Claim Verification
by: Freedman, Gabriel, et al.
Published: (2024)
by: Freedman, Gabriel, et al.
Published: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
by: Dugan, Liam, et al.
Published: (2024)
by: Dugan, Liam, et al.
Published: (2024)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
by: Chang, Hoyeon, et al.
Published: (2024)
by: Chang, Hoyeon, et al.
Published: (2024)
Identifying Fairness Issues in Automatically Generated Testing Content
by: Stowe, Kevin, et al.
Published: (2024)
by: Stowe, Kevin, et al.
Published: (2024)
TrustAI at SemEval-2024 Task 8: A Comprehensive Analysis of Multi-domain Machine Generated Text Detection Techniques
by: Urlana, Ashok, et al.
Published: (2024)
by: Urlana, Ashok, et al.
Published: (2024)
Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
by: Iqbal, Hasan, et al.
Published: (2024)
by: Iqbal, Hasan, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
ADAG: Automatically Describing Attribution Graphs
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
Algorithm for Semantic Network Generation from Texts of Low Resource Languages Such as Kiswahili
by: Wanjawa, Barack Wamkaya, et al.
Published: (2025)
by: Wanjawa, Barack Wamkaya, et al.
Published: (2025)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
by: Ge, Danying, et al.
Published: (2025)
by: Ge, Danying, et al.
Published: (2025)
LombardoGraphia: Automatic Classification of Lombard Orthography Variants
by: Signoroni, Edoardo, et al.
Published: (2026)
by: Signoroni, Edoardo, et al.
Published: (2026)
Towards Conditioning Clinical Text Generation for User Control
by: Koraş, Osman Alperen, et al.
Published: (2025)
by: Koraş, Osman Alperen, et al.
Published: (2025)
COMET-poly: Machine Translation Metric Grounded in Other Candidates
by: Züfle, Maike, et al.
Published: (2025)
by: Züfle, Maike, et al.
Published: (2025)
Synthetic Voice Data for Automatic Speech Recognition in African Languages
by: DeRenzi, Brian, et al.
Published: (2025)
by: DeRenzi, Brian, et al.
Published: (2025)
Automatic Generation of Conversational Interfaces for Tabular Data Analysis
by: Gomez-Vazquez, Marcos, et al.
Published: (2023)
by: Gomez-Vazquez, Marcos, et al.
Published: (2023)
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
by: Stowe, Kevin, et al.
Published: (2026)
by: Stowe, Kevin, et al.
Published: (2026)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
by: Cui, Hyang
Published: (2025)
by: Cui, Hyang
Published: (2025)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
by: Zhang, Luyan, et al.
Published: (2025)
by: Zhang, Luyan, et al.
Published: (2025)
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
by: Hawkins, John
Published: (2025)
by: Hawkins, John
Published: (2025)
Enhancing Emotion Prediction in News Headlines: Insights from ChatGPT and Seq2Seq Models for Free-Text Generation
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
Text Summarization With Graph Attention Networks
by: Ardestani, Mohammadreza, et al.
Published: (2026)
by: Ardestani, Mohammadreza, et al.
Published: (2026)
OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
by: Yu, Jinzheng, et al.
Published: (2025)
by: Yu, Jinzheng, et al.
Published: (2025)
Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data
by: Lübbers, Christopher Lee
Published: (2025)
by: Lübbers, Christopher Lee
Published: (2025)
Factual Dialogue Summarization via Learning from Large Language Models
by: Zhu, Rongxin, et al.
Published: (2024)
by: Zhu, Rongxin, et al.
Published: (2024)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
Breaking the HISCO Barrier: Automatic Occupational Standardization with OccCANINE
by: Dahl, Christian Møller, et al.
Published: (2024)
by: Dahl, Christian Møller, et al.
Published: (2024)
Active Few-Shot Learning for Text Classification
by: Ahmadnia, Saeed, et al.
Published: (2025)
by: Ahmadnia, Saeed, et al.
Published: (2025)
Investigating the Impact of Text Summarization on Topic Modeling
by: Khandelwal, Trishia
Published: (2024)
by: Khandelwal, Trishia
Published: (2024)
Normalization of Lithuanian Text Using Regular Expressions
by: Kasparaitis, Pijus
Published: (2023)
by: Kasparaitis, Pijus
Published: (2023)
Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges
by: Becker, Jonas, et al.
Published: (2024)
by: Becker, Jonas, et al.
Published: (2024)
APIO: Automatic Prompt Induction and Optimization for Grammatical Error Correction and Text Simplification
by: Chernodub, Artem, et al.
Published: (2025)
by: Chernodub, Artem, et al.
Published: (2025)
Similar Items
-
Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
by: Costa-jussà, Marta R., et al.
Published: (2024) -
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
by: Yun, Janghyeon, et al.
Published: (2025) -
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024) -
Argumentative Large Language Models for Explainable and Contestable Claim Verification
by: Freedman, Gabriel, et al.
Published: (2024) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)