Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | Godbole, Ameya, Jia, Robin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spiking the training data to correct for test set contamination
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
by: Seo, Wooseok, et al.
Published: (2025)
by: Seo, Wooseok, et al.
Published: (2025)
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025)
by: Hochlehnert, Andreas, et al.
Published: (2025)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
by: Yan, Tianyi Lorena, et al.
Published: (2025)
by: Yan, Tianyi Lorena, et al.
Published: (2025)
Hubble: a Model Suite to Advance the Study of LLM Memorization
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
Understanding Finetuning for Factual Knowledge Extraction
by: Ghosal, Gaurav, et al.
Published: (2024)
by: Ghosal, Gaurav, et al.
Published: (2024)
Conformal Language Model Reasoning with Coherent Factuality
by: Rubin-Toles, Maxon, et al.
Published: (2025)
by: Rubin-Toles, Maxon, et al.
Published: (2025)
Mamba Knockout for Unraveling Factual Information Flow
by: Endy, Nir, et al.
Published: (2025)
by: Endy, Nir, et al.
Published: (2025)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
by: Cui, Guanyu, et al.
Published: (2026)
by: Cui, Guanyu, et al.
Published: (2026)
Persuasion Tokens for Editing Factual Knowledge in LLMs
by: Youssef, Paul, et al.
Published: (2026)
by: Youssef, Paul, et al.
Published: (2026)
Factual Consistency of Multilingual Pretrained Language Models
by: Fierro, Constanza, et al.
Published: (2022)
by: Fierro, Constanza, et al.
Published: (2022)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
by: Anzenberg, Eitan, et al.
Published: (2025)
by: Anzenberg, Eitan, et al.
Published: (2025)
LoFTI: Localization and Factuality Transfer to Indian Locales
by: Simon, Sona Elza, et al.
Published: (2024)
by: Simon, Sona Elza, et al.
Published: (2024)
Temporally Consistent Factuality Probing for Large Language Models
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
by: Mahaut, Matéo, et al.
Published: (2024)
by: Mahaut, Matéo, et al.
Published: (2024)
Zero-shot Factual Consistency Evaluation Across Domains
by: Agarwal, Raunak
Published: (2024)
by: Agarwal, Raunak
Published: (2024)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
by: Roy, Tiasa Singha, et al.
Published: (2025)
by: Roy, Tiasa Singha, et al.
Published: (2025)
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
by: Jiang, Zhengping, et al.
Published: (2025)
by: Jiang, Zhengping, et al.
Published: (2025)
SIFiD: Reassess Summary Factual Inconsistency Detection with LLM
by: Yang, Jiuding, et al.
Published: (2024)
by: Yang, Jiuding, et al.
Published: (2024)
Alexpaca: Learning Factual Clarification Question Generation Without Examples
by: Toles, Matthew, et al.
Published: (2023)
by: Toles, Matthew, et al.
Published: (2023)
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
by: Shin, Hagyeong, et al.
Published: (2025)
by: Shin, Hagyeong, et al.
Published: (2025)
Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization
by: Shin, Sungbin, et al.
Published: (2024)
by: Shin, Sungbin, et al.
Published: (2024)
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
by: Upasani, Shubhangi, et al.
Published: (2026)
by: Upasani, Shubhangi, et al.
Published: (2026)
Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
by: Krishna, Kundan, et al.
Published: (2024)
by: Krishna, Kundan, et al.
Published: (2024)
Alternate Preference Optimization for Unlearning Factual Knowledge in Large Language Models
by: Mekala, Anmol, et al.
Published: (2024)
by: Mekala, Anmol, et al.
Published: (2024)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
by: Chughtai, Bilal, et al.
Published: (2024)
by: Chughtai, Bilal, et al.
Published: (2024)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
by: Qi, Jianing, et al.
Published: (2024)
by: Qi, Jianing, et al.
Published: (2024)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
The Illusionist's Prompt: Exposing the Factual Vulnerabilities of Large Language Models with Linguistic Nuances
by: Wang, Yining, et al.
Published: (2025)
by: Wang, Yining, et al.
Published: (2025)
Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
by: Akpinar, Nil-Jana, et al.
Published: (2025)
by: Akpinar, Nil-Jana, et al.
Published: (2025)
Identifying Factual Inconsistencies in Summaries: Grounding LLM Inference via Task Taxonomy
by: Xu, Liyan, et al.
Published: (2024)
by: Xu, Liyan, et al.
Published: (2024)
Language Models with Conformal Factuality Guarantees
by: Mohri, Christopher, et al.
Published: (2024)
by: Mohri, Christopher, et al.
Published: (2024)
Interrogating LLM design under a fair learning doctrine
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)
When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning
by: Deng, Mengyi, et al.
Published: (2025)
by: Deng, Mengyi, et al.
Published: (2025)
Similar Items
-
Spiking the training data to correct for test set contamination
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026) -
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
by: Seo, Wooseok, et al.
Published: (2025) -
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
by: Fu, Deqing, et al.
Published: (2023) -
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025) -
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
by: Yan, Tianyi Lorena, et al.
Published: (2025)