Language Models with Conformal Factuality Guarantees
Fuente:
arXiv
Salvato in:
| Autori principali: | Mohri, Christopher, Hashimoto, Tatsunori |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Bitter Lesson for Data Filtering
di: Mohri, Christopher, et al.
Pubblicazione: (2026)
di: Mohri, Christopher, et al.
Pubblicazione: (2026)
Observational Scaling Laws and the Predictability of Language Model Performance
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
di: Ruan, Yangjun, et al.
Pubblicazione: (2024)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
di: Jiang, Mingjian, et al.
Pubblicazione: (2024)
di: Jiang, Mingjian, et al.
Pubblicazione: (2024)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
di: Rubashevskii, Aleksandr, et al.
Pubblicazione: (2026)
di: Rubashevskii, Aleksandr, et al.
Pubblicazione: (2026)
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
di: Nie, Fan, et al.
Pubblicazione: (2024)
di: Nie, Fan, et al.
Pubblicazione: (2024)
Linguistic Calibration of Long-Form Generations
di: Band, Neil, et al.
Pubblicazione: (2024)
di: Band, Neil, et al.
Pubblicazione: (2024)
Understanding Finetuning for Factual Knowledge Extraction
di: Ghosal, Gaurav, et al.
Pubblicazione: (2024)
di: Ghosal, Gaurav, et al.
Pubblicazione: (2024)
ConU: Conformal Uncertainty in Large Language Models with Correctness Coverage Guarantees
di: Wang, Zhiyuan, et al.
Pubblicazione: (2024)
di: Wang, Zhiyuan, et al.
Pubblicazione: (2024)
Eliciting Language Model Behaviors with Investigator Agents
di: Li, Xiang Lisa, et al.
Pubblicazione: (2025)
di: Li, Xiang Lisa, et al.
Pubblicazione: (2025)
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
di: Dubois, Yann, et al.
Pubblicazione: (2024)
di: Dubois, Yann, et al.
Pubblicazione: (2024)
Reasoning to Learn from Latent Thoughts
di: Ruan, Yangjun, et al.
Pubblicazione: (2025)
di: Ruan, Yangjun, et al.
Pubblicazione: (2025)
Synthetic continued pretraining
di: Yang, Zitong, et al.
Pubblicazione: (2024)
di: Yang, Zitong, et al.
Pubblicazione: (2024)
Agentic Adversarial QA for Improving Domain-Specific LLMs
di: Grari, Vincent, et al.
Pubblicazione: (2026)
di: Grari, Vincent, et al.
Pubblicazione: (2026)
Factuality Challenges in the Era of Large Language Models
di: Augenstein, Isabelle, et al.
Pubblicazione: (2023)
di: Augenstein, Isabelle, et al.
Pubblicazione: (2023)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
di: Si, Chenglei, et al.
Pubblicazione: (2024)
di: Si, Chenglei, et al.
Pubblicazione: (2024)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
di: Si, Chenglei, et al.
Pubblicazione: (2025)
di: Si, Chenglei, et al.
Pubblicazione: (2025)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
di: Chen, Yi, et al.
Pubblicazione: (2026)
di: Chen, Yi, et al.
Pubblicazione: (2026)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
di: Tamoyan, Hovhannes, et al.
Pubblicazione: (2025)
di: Tamoyan, Hovhannes, et al.
Pubblicazione: (2025)
Towards Execution-Grounded Automated AI Research
di: Si, Chenglei, et al.
Pubblicazione: (2026)
di: Si, Chenglei, et al.
Pubblicazione: (2026)
ConformalNL2LTL: Translating Natural Language Instructions into Temporal Logic Formulas with Conformal Correctness Guarantees
di: Sundarsingh, David Smith, et al.
Pubblicazione: (2025)
di: Sundarsingh, David Smith, et al.
Pubblicazione: (2025)
Synthetic Data for any Differentiable Target
di: Thrush, Tristan, et al.
Pubblicazione: (2026)
di: Thrush, Tristan, et al.
Pubblicazione: (2026)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
Evaluating the Factuality of Large Language Models using Large-Scale Knowledge Graphs
di: Liu, Xiaoze, et al.
Pubblicazione: (2024)
di: Liu, Xiaoze, et al.
Pubblicazione: (2024)
SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language Models
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
di: Yuksekgonul, Mert, et al.
Pubblicazione: (2023)
di: Yuksekgonul, Mert, et al.
Pubblicazione: (2023)
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
di: Chuang, Yung-Sung, et al.
Pubblicazione: (2023)
di: Chuang, Yung-Sung, et al.
Pubblicazione: (2023)
The Factuality of Large Language Models in the Legal Domain
di: Hamdani, Rajaa El, et al.
Pubblicazione: (2024)
di: Hamdani, Rajaa El, et al.
Pubblicazione: (2024)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
di: Wang, Shengyuan, et al.
Pubblicazione: (2025)
di: Wang, Shengyuan, et al.
Pubblicazione: (2025)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)
SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering
di: Li, Xiaopeng, et al.
Pubblicazione: (2024)
di: Li, Xiaopeng, et al.
Pubblicazione: (2024)
Benchmarking Distributional Alignment of Large Language Models
di: Meister, Nicole, et al.
Pubblicazione: (2024)
di: Meister, Nicole, et al.
Pubblicazione: (2024)
Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
di: Qi, Jirui, et al.
Pubblicazione: (2023)
di: Qi, Jirui, et al.
Pubblicazione: (2023)
SConU: Selective Conformal Uncertainty in Large Language Models
di: Wang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Wang, Zhiyuan, et al.
Pubblicazione: (2025)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
di: Ruan, Yangjun, et al.
Pubblicazione: (2023)
di: Ruan, Yangjun, et al.
Pubblicazione: (2023)
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
di: Dubois, Yann, et al.
Pubblicazione: (2023)
di: Dubois, Yann, et al.
Pubblicazione: (2023)
Entity-level Factual Adaptiveness of Fine-tuning based Abstractive Summarization Models
di: Song, Jongyoon, et al.
Pubblicazione: (2024)
di: Song, Jongyoon, et al.
Pubblicazione: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
di: Smith, Matthew L., et al.
Pubblicazione: (2026)
di: Smith, Matthew L., et al.
Pubblicazione: (2026)
Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models
di: Giebisch, Rafael, et al.
Pubblicazione: (2025)
di: Giebisch, Rafael, et al.
Pubblicazione: (2025)
API Is Enough: Conformal Prediction for Large Language Models Without Logit-Access
di: Su, Jiayuan, et al.
Pubblicazione: (2024)
di: Su, Jiayuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Bitter Lesson for Data Filtering
di: Mohri, Christopher, et al.
Pubblicazione: (2026) -
Observational Scaling Laws and the Predictability of Language Model Performance
di: Ruan, Yangjun, et al.
Pubblicazione: (2024) -
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
di: Jiang, Mingjian, et al.
Pubblicazione: (2024) -
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
di: Rubashevskii, Aleksandr, et al.
Pubblicazione: (2026) -
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
di: Nie, Fan, et al.
Pubblicazione: (2024)