Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
Fuente:
arXiv
Salvato in:
| Autori principali: | Godbole, Ameya, Jia, Robin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Spiking the training data to correct for test set contamination
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2026)
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2026)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
di: Seo, Wooseok, et al.
Pubblicazione: (2025)
di: Seo, Wooseok, et al.
Pubblicazione: (2025)
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
di: Fu, Deqing, et al.
Pubblicazione: (2023)
di: Fu, Deqing, et al.
Pubblicazione: (2023)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
di: Hochlehnert, Andreas, et al.
Pubblicazione: (2025)
di: Hochlehnert, Andreas, et al.
Pubblicazione: (2025)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)
Hubble: a Model Suite to Advance the Study of LLM Memorization
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2025)
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2025)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
di: Mujahid, Zain Muhammad, et al.
Pubblicazione: (2025)
di: Mujahid, Zain Muhammad, et al.
Pubblicazione: (2025)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
di: Haller, Patrick, et al.
Pubblicazione: (2025)
di: Haller, Patrick, et al.
Pubblicazione: (2025)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
di: Chen, Yi, et al.
Pubblicazione: (2026)
di: Chen, Yi, et al.
Pubblicazione: (2026)
Understanding Finetuning for Factual Knowledge Extraction
di: Ghosal, Gaurav, et al.
Pubblicazione: (2024)
di: Ghosal, Gaurav, et al.
Pubblicazione: (2024)
Conformal Language Model Reasoning with Coherent Factuality
di: Rubin-Toles, Maxon, et al.
Pubblicazione: (2025)
di: Rubin-Toles, Maxon, et al.
Pubblicazione: (2025)
Mamba Knockout for Unraveling Factual Information Flow
di: Endy, Nir, et al.
Pubblicazione: (2025)
di: Endy, Nir, et al.
Pubblicazione: (2025)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
di: Cui, Guanyu, et al.
Pubblicazione: (2026)
di: Cui, Guanyu, et al.
Pubblicazione: (2026)
Persuasion Tokens for Editing Factual Knowledge in LLMs
di: Youssef, Paul, et al.
Pubblicazione: (2026)
di: Youssef, Paul, et al.
Pubblicazione: (2026)
Factual Consistency of Multilingual Pretrained Language Models
di: Fierro, Constanza, et al.
Pubblicazione: (2022)
di: Fierro, Constanza, et al.
Pubblicazione: (2022)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
di: Anzenberg, Eitan, et al.
Pubblicazione: (2025)
di: Anzenberg, Eitan, et al.
Pubblicazione: (2025)
LoFTI: Localization and Factuality Transfer to Indian Locales
di: Simon, Sona Elza, et al.
Pubblicazione: (2024)
di: Simon, Sona Elza, et al.
Pubblicazione: (2024)
Temporally Consistent Factuality Probing for Large Language Models
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2024)
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2024)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
di: Mahaut, Matéo, et al.
Pubblicazione: (2024)
di: Mahaut, Matéo, et al.
Pubblicazione: (2024)
Zero-shot Factual Consistency Evaluation Across Domains
di: Agarwal, Raunak
Pubblicazione: (2024)
di: Agarwal, Raunak
Pubblicazione: (2024)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
di: Jiang, Zhengping, et al.
Pubblicazione: (2025)
di: Jiang, Zhengping, et al.
Pubblicazione: (2025)
SIFiD: Reassess Summary Factual Inconsistency Detection with LLM
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
Alexpaca: Learning Factual Clarification Question Generation Without Examples
di: Toles, Matthew, et al.
Pubblicazione: (2023)
di: Toles, Matthew, et al.
Pubblicazione: (2023)
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
di: Shin, Hagyeong, et al.
Pubblicazione: (2025)
di: Shin, Hagyeong, et al.
Pubblicazione: (2025)
Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization
di: Shin, Sungbin, et al.
Pubblicazione: (2024)
di: Shin, Sungbin, et al.
Pubblicazione: (2024)
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
di: Upasani, Shubhangi, et al.
Pubblicazione: (2026)
di: Upasani, Shubhangi, et al.
Pubblicazione: (2026)
Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines
di: Li, Yuchen, et al.
Pubblicazione: (2024)
di: Li, Yuchen, et al.
Pubblicazione: (2024)
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
di: Krishna, Kundan, et al.
Pubblicazione: (2024)
di: Krishna, Kundan, et al.
Pubblicazione: (2024)
Alternate Preference Optimization for Unlearning Factual Knowledge in Large Language Models
di: Mekala, Anmol, et al.
Pubblicazione: (2024)
di: Mekala, Anmol, et al.
Pubblicazione: (2024)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2026)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2026)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
di: Qi, Jianing, et al.
Pubblicazione: (2024)
di: Qi, Jianing, et al.
Pubblicazione: (2024)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
di: Wang, Qianli, et al.
Pubblicazione: (2025)
di: Wang, Qianli, et al.
Pubblicazione: (2025)
The Illusionist's Prompt: Exposing the Factual Vulnerabilities of Large Language Models with Linguistic Nuances
di: Wang, Yining, et al.
Pubblicazione: (2025)
di: Wang, Yining, et al.
Pubblicazione: (2025)
Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
di: Akpinar, Nil-Jana, et al.
Pubblicazione: (2025)
di: Akpinar, Nil-Jana, et al.
Pubblicazione: (2025)
Identifying Factual Inconsistencies in Summaries: Grounding LLM Inference via Task Taxonomy
di: Xu, Liyan, et al.
Pubblicazione: (2024)
di: Xu, Liyan, et al.
Pubblicazione: (2024)
Language Models with Conformal Factuality Guarantees
di: Mohri, Christopher, et al.
Pubblicazione: (2024)
di: Mohri, Christopher, et al.
Pubblicazione: (2024)
Interrogating LLM design under a fair learning doctrine
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2025)
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2025)
When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning
di: Deng, Mengyi, et al.
Pubblicazione: (2025)
di: Deng, Mengyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Spiking the training data to correct for test set contamination
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2026) -
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
di: Seo, Wooseok, et al.
Pubblicazione: (2025) -
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
di: Fu, Deqing, et al.
Pubblicazione: (2023) -
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
di: Hochlehnert, Andreas, et al.
Pubblicazione: (2025) -
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)