FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Marinescu, Radu, Bhattacharjya, Debarun, Lee, Junkyu, Tchrakian, Tigran, Cano, Javier Carnerero, Hou, Yufang, Daly, Elizabeth, Pascale, Alessandra |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
di: Carnerero-Cano, Javier, et al.
Pubblicazione: (2026)
di: Carnerero-Cano, Javier, et al.
Pubblicazione: (2026)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
di: Hou, Yufang, et al.
Pubblicazione: (2024)
di: Hou, Yufang, et al.
Pubblicazione: (2024)
Foundation Model Sherpas: Guiding Foundation Models through Knowledge and Reasoning
di: Bhattacharjya, Debarun, et al.
Pubblicazione: (2024)
di: Bhattacharjya, Debarun, et al.
Pubblicazione: (2024)
SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models
di: Bhattacharjya, Debarun, et al.
Pubblicazione: (2025)
di: Bhattacharjya, Debarun, et al.
Pubblicazione: (2025)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
di: Dejl, Adam, et al.
Pubblicazione: (2025)
di: Dejl, Adam, et al.
Pubblicazione: (2025)
The Consistency Hypothesis in Uncertainty Quantification for Large Language Models
di: Xiao, Quan, et al.
Pubblicazione: (2025)
di: Xiao, Quan, et al.
Pubblicazione: (2025)
Q-function Decomposition with Intervention Semantics with Factored Action Spaces
di: Lee, Junkyu, et al.
Pubblicazione: (2025)
di: Lee, Junkyu, et al.
Pubblicazione: (2025)
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
di: McCarthy, James, et al.
Pubblicazione: (2025)
di: McCarthy, James, et al.
Pubblicazione: (2025)
VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts
di: Liu, Xin, et al.
Pubblicazione: (2025)
di: Liu, Xin, et al.
Pubblicazione: (2025)
Process Supervision for Chain-of-Thought Reasoning via Monte Carlo Net Information Gain
di: Royer, Corentin, et al.
Pubblicazione: (2026)
di: Royer, Corentin, et al.
Pubblicazione: (2026)
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
di: Li, Junliang, et al.
Pubblicazione: (2025)
di: Li, Junliang, et al.
Pubblicazione: (2025)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
di: Gajcin, Jasmina, et al.
Pubblicazione: (2025)
di: Gajcin, Jasmina, et al.
Pubblicazione: (2025)
MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
di: Ning, Yucheng, et al.
Pubblicazione: (2025)
di: Ning, Yucheng, et al.
Pubblicazione: (2025)
Merging Facts, Crafting Fallacies: Evaluating the Contradictory Nature of Aggregated Factual Claims in Long-Form Generations
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2024)
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2024)
Distilling Event Sequence Knowledge From Large Language Models
di: Wadhwa, Somin, et al.
Pubblicazione: (2024)
di: Wadhwa, Somin, et al.
Pubblicazione: (2024)
The effect of Skyrme--Chern-Simons dynamics on gauged Skyrmions in $2+1$ dimensions
di: Navarro-Lerida, Francisco, et al.
Pubblicazione: (2023)
di: Navarro-Lerida, Francisco, et al.
Pubblicazione: (2023)
Attractive and repulsive Yang-Mills--Higgs magnetic monopoles on $\mathbb{R}^3$
di: Navarro-Lérida, Francisco, et al.
Pubblicazione: (2026)
di: Navarro-Lérida, Francisco, et al.
Pubblicazione: (2026)
What Would an LLM Do? Evaluating Large Language Models for Policymaking to Alleviate Homelessness
di: Coz, Pierre Le, et al.
Pubblicazione: (2025)
di: Coz, Pierre Le, et al.
Pubblicazione: (2025)
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes
di: Shou, Xiao, et al.
Pubblicazione: (2024)
di: Shou, Xiao, et al.
Pubblicazione: (2024)
Gauged Skyrme analogue of Chern-Pontryagin
di: Tchrakian, D. H.
Pubblicazione: (2024)
di: Tchrakian, D. H.
Pubblicazione: (2024)
Think Through Uncertainty: Improving Long-Form Generation Factuality via Reasoning Calibration
di: Liu, Xin, et al.
Pubblicazione: (2026)
di: Liu, Xin, et al.
Pubblicazione: (2026)
FactAlign: Long-form Factuality Alignment of Large Language Models
di: Huang, Chao-Wei, et al.
Pubblicazione: (2024)
di: Huang, Chao-Wei, et al.
Pubblicazione: (2024)
Less is More: Efficient Weight Farcasting with 1-Layer Neural Network
di: Shou, Xiao, et al.
Pubblicazione: (2025)
di: Shou, Xiao, et al.
Pubblicazione: (2025)
MAVEN-Fact: A Large-scale Event Factuality Detection Dataset
di: Li, Chunyang, et al.
Pubblicazione: (2024)
di: Li, Chunyang, et al.
Pubblicazione: (2024)
How Does Response Length Affect Long-Form Factuality
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
di: Baek, Jinheon, et al.
Pubblicazione: (2025)
di: Baek, Jinheon, et al.
Pubblicazione: (2025)
QueryGym: Step-by-Step Interaction with Relational Databases
di: Ananthakrishnan, Haritha, et al.
Pubblicazione: (2025)
di: Ananthakrishnan, Haritha, et al.
Pubblicazione: (2025)
FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
MedFact-R1: Towards Factual Medical Reasoning via Pseudo-Label Augmentation
di: Li, Gengliang, et al.
Pubblicazione: (2025)
di: Li, Gengliang, et al.
Pubblicazione: (2025)
InFact: Informativeness Alignment for Improved LLM Factuality
di: Cohen, Roi, et al.
Pubblicazione: (2025)
di: Cohen, Roi, et al.
Pubblicazione: (2025)
Fact or Facsimile? Evaluating the Factual Robustness of Modern Retrievers
di: Wu, Haoyu, et al.
Pubblicazione: (2025)
di: Wu, Haoyu, et al.
Pubblicazione: (2025)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
di: Carnerero-Cano, Javier, et al.
Pubblicazione: (2026)
di: Carnerero-Cano, Javier, et al.
Pubblicazione: (2026)
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
di: Nie, Fan, et al.
Pubblicazione: (2024)
di: Nie, Fan, et al.
Pubblicazione: (2024)
Long-Form Information Alignment Evaluation Beyond Atomic Facts
di: Zheng, Danna, et al.
Pubblicazione: (2025)
di: Zheng, Danna, et al.
Pubblicazione: (2025)
Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning
di: Pronesti, Massimiliano, et al.
Pubblicazione: (2026)
di: Pronesti, Massimiliano, et al.
Pubblicazione: (2026)
Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
di: Pronesti, Massimiliano, et al.
Pubblicazione: (2025)
di: Pronesti, Massimiliano, et al.
Pubblicazione: (2025)
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
di: Zheng, Tianshi, et al.
Pubblicazione: (2024)
di: Zheng, Tianshi, et al.
Pubblicazione: (2024)
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
di: Chen, Mingda, et al.
Pubblicazione: (2025)
di: Chen, Mingda, et al.
Pubblicazione: (2025)
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
di: Wanner, Miriam, et al.
Pubblicazione: (2024)
di: Wanner, Miriam, et al.
Pubblicazione: (2024)
Multilinear and Linear Programs for Partially Identifiable Queries in Quasi-Markovian Structural Causal Models
di: Arroyo, João P., et al.
Pubblicazione: (2025)
di: Arroyo, João P., et al.
Pubblicazione: (2025)
Documenti analoghi
-
FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
di: Carnerero-Cano, Javier, et al.
Pubblicazione: (2026) -
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
di: Hou, Yufang, et al.
Pubblicazione: (2024) -
Foundation Model Sherpas: Guiding Foundation Models through Knowledge and Reasoning
di: Bhattacharjya, Debarun, et al.
Pubblicazione: (2024) -
SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models
di: Bhattacharjya, Debarun, et al.
Pubblicazione: (2025) -
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
di: Dejl, Adam, et al.
Pubblicazione: (2025)