Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Jafari, Nazanin, Allan, James, Iyyer, Mohit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Robust Claim Verification Through Fact Detection
di: Jafari, Nazanin, et al.
Pubblicazione: (2024)
di: Jafari, Nazanin, et al.
Pubblicazione: (2024)
Literary Evidence Retrieval via Long-Context Language Models
di: Thai, Katherine, et al.
Pubblicazione: (2025)
di: Thai, Katherine, et al.
Pubblicazione: (2025)
Argument Collapse: LLMs Flatten Long-Form Public Debate
di: Kim, Yekyung, et al.
Pubblicazione: (2026)
di: Kim, Yekyung, et al.
Pubblicazione: (2026)
Suri: Multi-constraint Instruction Following for Long-form Text Generation
di: Pham, Chau Minh, et al.
Pubblicazione: (2024)
di: Pham, Chau Minh, et al.
Pubblicazione: (2024)
Target Span Detection for Implicit Harmful Content
di: Jafari, Nazanin, et al.
Pubblicazione: (2024)
di: Jafari, Nazanin, et al.
Pubblicazione: (2024)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
di: Dejl, Adam, et al.
Pubblicazione: (2025)
di: Dejl, Adam, et al.
Pubblicazione: (2025)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
di: Song, Yixiao, et al.
Pubblicazione: (2024)
di: Song, Yixiao, et al.
Pubblicazione: (2024)
Localizing and Mitigating Errors in Long-form Question Answering
di: Sachdeva, Rachneet, et al.
Pubblicazione: (2024)
di: Sachdeva, Rachneet, et al.
Pubblicazione: (2024)
Geometric Factual Recall in Transformers
di: Ravfogel, Shauli, et al.
Pubblicazione: (2026)
di: Ravfogel, Shauli, et al.
Pubblicazione: (2026)
Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation
di: Samarinas, Chris, et al.
Pubblicazione: (2025)
di: Samarinas, Chris, et al.
Pubblicazione: (2025)
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
di: Luo, Wen, et al.
Pubblicazione: (2026)
di: Luo, Wen, et al.
Pubblicazione: (2026)
Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall
di: Qi, Zehan, et al.
Pubblicazione: (2024)
di: Qi, Zehan, et al.
Pubblicazione: (2024)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
di: Ardestani, MohamamdJavad, et al.
Pubblicazione: (2025)
di: Ardestani, MohamamdJavad, et al.
Pubblicazione: (2025)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
di: Samuel, Vinay, et al.
Pubblicazione: (2026)
di: Samuel, Vinay, et al.
Pubblicazione: (2026)
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
di: Srivastava, Alisha, et al.
Pubblicazione: (2025)
di: Srivastava, Alisha, et al.
Pubblicazione: (2025)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
di: Liu, Yihong, et al.
Pubblicazione: (2026)
di: Liu, Yihong, et al.
Pubblicazione: (2026)
StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall
di: Wu, Yerong, et al.
Pubblicazione: (2026)
di: Wu, Yerong, et al.
Pubblicazione: (2026)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
di: Russell, Jenna, et al.
Pubblicazione: (2025)
di: Russell, Jenna, et al.
Pubblicazione: (2025)
CLIPPER: Compression enables long-context synthetic data generation
di: Pham, Chau Minh, et al.
Pubblicazione: (2025)
di: Pham, Chau Minh, et al.
Pubblicazione: (2025)
Merging Facts, Crafting Fallacies: Evaluating the Contradictory Nature of Aggregated Factual Claims in Long-Form Generations
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2024)
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2024)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
di: Yuan, Jiaqing, et al.
Pubblicazione: (2024)
di: Yuan, Jiaqing, et al.
Pubblicazione: (2024)
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
di: Wanner, Miriam, et al.
Pubblicazione: (2024)
di: Wanner, Miriam, et al.
Pubblicazione: (2024)
All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations
di: Wanner, Miriam, et al.
Pubblicazione: (2025)
di: Wanner, Miriam, et al.
Pubblicazione: (2025)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
di: Naseh, Ali, et al.
Pubblicazione: (2024)
di: Naseh, Ali, et al.
Pubblicazione: (2024)
Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline
di: Lu, Meng, et al.
Pubblicazione: (2025)
di: Lu, Meng, et al.
Pubblicazione: (2025)
FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
One ruler to measure them all: Benchmarking multilingual long-context language models
di: Kim, Yekyung, et al.
Pubblicazione: (2025)
di: Kim, Yekyung, et al.
Pubblicazione: (2025)
EditLens: Quantifying the Extent of AI Editing in Text
di: Thai, Katherine, et al.
Pubblicazione: (2025)
di: Thai, Katherine, et al.
Pubblicazione: (2025)
Think Through Uncertainty: Improving Long-Form Generation Factuality via Reasoning Calibration
di: Liu, Xin, et al.
Pubblicazione: (2026)
di: Liu, Xin, et al.
Pubblicazione: (2026)
Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown
di: Tu, Lifu, et al.
Pubblicazione: (2024)
di: Tu, Lifu, et al.
Pubblicazione: (2024)
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
di: Bhalerao, Parth, et al.
Pubblicazione: (2026)
di: Bhalerao, Parth, et al.
Pubblicazione: (2026)
VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts
di: Liu, Xin, et al.
Pubblicazione: (2025)
di: Liu, Xin, et al.
Pubblicazione: (2025)
How Does Response Length Affect Long-Form Factuality
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
di: Yang, Jiayu, et al.
Pubblicazione: (2025)
di: Yang, Jiayu, et al.
Pubblicazione: (2025)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
di: Hastuti, Rochana Prih, et al.
Pubblicazione: (2025)
di: Hastuti, Rochana Prih, et al.
Pubblicazione: (2025)
Precise Information Control in Long-Form Text Generation
di: He, Jacqueline, et al.
Pubblicazione: (2025)
di: He, Jacqueline, et al.
Pubblicazione: (2025)
Frankentext: Stitching random text fragments into long-form narratives
di: Pham, Chau Minh, et al.
Pubblicazione: (2025)
di: Pham, Chau Minh, et al.
Pubblicazione: (2025)
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2024)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2024)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
di: Calderon, Nitay, et al.
Pubblicazione: (2026)
di: Calderon, Nitay, et al.
Pubblicazione: (2026)
Understanding Factual Recall in Transformers via Associative Memories
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Robust Claim Verification Through Fact Detection
di: Jafari, Nazanin, et al.
Pubblicazione: (2024) -
Literary Evidence Retrieval via Long-Context Language Models
di: Thai, Katherine, et al.
Pubblicazione: (2025) -
Argument Collapse: LLMs Flatten Long-Form Public Debate
di: Kim, Yekyung, et al.
Pubblicazione: (2026) -
Suri: Multi-constraint Instruction Following for Long-form Text Generation
di: Pham, Chau Minh, et al.
Pubblicazione: (2024) -
Target Span Detection for Implicit Harmful Content
di: Jafari, Nazanin, et al.
Pubblicazione: (2024)