Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Xiaoyuan, Truong, Kimberly Le, Fogliato, Riccardo, Swamy, Gokul, Zhang, Weijian, Yang, Minglai, Ye, Longtian, Liu, Bangya, Liu, Minghao, Ilyas, Andrew, Wu, Steven |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
by: Truong, Kimberly Le, et al.
Published: (2025)
by: Truong, Kimberly Le, et al.
Published: (2025)
Confidence Intervals for Error Rates in 1:1 Matching Tasks: Critical Statistical Analysis and Recommendations
by: Fogliato, Riccardo, et al.
Published: (2023)
by: Fogliato, Riccardo, et al.
Published: (2023)
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
by: Bertran, Martin, et al.
Published: (2026)
by: Bertran, Martin, et al.
Published: (2026)
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
by: Liu, Bangya, et al.
Published: (2024)
by: Liu, Bangya, et al.
Published: (2024)
Fzc_Collatz_Final
by: Fogliato, Matteo
Published: (2026)
by: Fogliato, Matteo
Published: (2026)
Resultados CU
by: Fogliato, Matteo
Published: (2026)
by: Fogliato, Matteo
Published: (2026)
Fz Fzc Unified
by: Fogliato, Matteo
Published: (2026)
by: Fogliato, Matteo
Published: (2026)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
by: Si, Chenglei, et al.
Published: (2023)
by: Si, Chenglei, et al.
Published: (2023)
Multi-Agent Imitation Learning: Value is Easy, Regret is Hard
by: Tang, Jingwu, et al.
Published: (2024)
by: Tang, Jingwu, et al.
Published: (2024)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
by: Sutawika, Lintang, et al.
Published: (2026)
by: Sutawika, Lintang, et al.
Published: (2026)
Inverse Reinforcement Learning without Reinforcement Learning
by: Swamy, Gokul, et al.
Published: (2023)
by: Swamy, Gokul, et al.
Published: (2023)
Multicalibration for Confidence Scoring in LLMs
by: Detommaso, Gianluca, et al.
Published: (2024)
by: Detommaso, Gianluca, et al.
Published: (2024)
Stronger Neyman Regret Guarantees for Adaptive Experimental Design
by: Noarov, Georgy, et al.
Published: (2025)
by: Noarov, Georgy, et al.
Published: (2025)
A Framework for Efficient Model Evaluation through Stratification, Sampling, and Estimation
by: Fogliato, Riccardo, et al.
Published: (2024)
by: Fogliato, Riccardo, et al.
Published: (2024)
ICAGC 2024: Inspirational and Convincing Audio Generation Challenge 2024
by: Fu, Ruibo, et al.
Published: (2024)
by: Fu, Ruibo, et al.
Published: (2024)
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
by: Swamy, Gokul, et al.
Published: (2024)
by: Swamy, Gokul, et al.
Published: (2024)
Precise Model Benchmarking with Only a Few Observations
by: Fogliato, Riccardo, et al.
Published: (2024)
by: Fogliato, Riccardo, et al.
Published: (2024)
Indoor Air Quality Compliance for AC Installations in 2025
by: Gokul
Published: (2025)
by: Gokul
Published: (2025)
From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment
by: Wu, Yilin, et al.
Published: (2025)
by: Wu, Yilin, et al.
Published: (2025)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
Chapter Convincing fashion consumers to go green
by: Morais, Ricardo, et al.
Published: (2023)
by: Morais, Ricardo, et al.
Published: (2023)
Finding Convincing Views to Endorse a Claim
by: Agmon, Shunit, et al.
Published: (2024)
by: Agmon, Shunit, et al.
Published: (2024)
Can Language Models Recognize Convincing Arguments?
by: Rescala, Paula, et al.
Published: (2024)
by: Rescala, Paula, et al.
Published: (2024)
Full Proportional Justified Representation
by: Kalayci, Yusuf Hakan, et al.
Published: (2025)
by: Kalayci, Yusuf Hakan, et al.
Published: (2025)
All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning
by: Swamy, Gokul, et al.
Published: (2025)
by: Swamy, Gokul, et al.
Published: (2025)
Hybrid Inverse Reinforcement Learning
by: Ren, Juntao, et al.
Published: (2024)
by: Ren, Juntao, et al.
Published: (2024)
The Virtues of Pessimism in Inverse Reinforcement Learning
by: Wu, David, et al.
Published: (2024)
by: Wu, David, et al.
Published: (2024)
Efficient Imitation under Misspecification
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
Examining Different Placement Strategies for Indoor Environmental Quality Sensors in Office Environments
by: Talami, Riccardo, et al.
Published: (2025)
by: Talami, Riccardo, et al.
Published: (2025)
What Evidence Do Language Models Find Convincing?
by: Wan, Alexander, et al.
Published: (2024)
by: Wan, Alexander, et al.
Published: (2024)
Make a Convincing Case For Increasing Your Dues
Published: (2024)
Published: (2024)
A Survey on LLM-powered Agents for Recommender Systems
by: Peng, Qiyao, et al.
Published: (2025)
by: Peng, Qiyao, et al.
Published: (2025)
Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM
by: Duong, Thang, et al.
Published: (2025)
by: Duong, Thang, et al.
Published: (2025)
Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge
by: Guo, Dake, et al.
Published: (2024)
by: Guo, Dake, et al.
Published: (2024)
MV-S2V: Multi-View Subject-Consistent Video Generation
by: Song, Ziyang, et al.
Published: (2026)
by: Song, Ziyang, et al.
Published: (2026)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Improving LLM Group Fairness on Tabular Data via In-Context Learning
by: Cherepanova, Valeriia, et al.
Published: (2024)
by: Cherepanova, Valeriia, et al.
Published: (2024)
Justifying Transgression
by: Kruijtzer, Gijs
Published: (2024)
by: Kruijtzer, Gijs
Published: (2024)
Similar Items
-
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
by: Truong, Kimberly Le, et al.
Published: (2025) -
Confidence Intervals for Error Rates in 1:1 Matching Tasks: Critical Statistical Analysis and Recommendations
by: Fogliato, Riccardo, et al.
Published: (2023) -
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
by: Bertran, Martin, et al.
Published: (2026) -
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
by: Liu, Bangya, et al.
Published: (2024) -
Fzc_Collatz_Final
by: Fogliato, Matteo
Published: (2026)