Saved in:
| Main Authors: | Devic, Siddartha, Srinivasan, Tejas, Thomason, Jesse, Neiswanger, Willie, Sharan, Vatsal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.07461 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
by: Srinivasan, Tejas, et al.
Published: (2025)
by: Srinivasan, Tejas, et al.
Published: (2025)
WinoViz: Probing Visual Properties of Objects Under Different States
by: Jin, Woojeong, et al.
Published: (2024)
by: Jin, Woojeong, et al.
Published: (2024)
An External Fairness Evaluation of LinkedIn Talent Search
by: Behzad, Tina, et al.
Published: (2025)
by: Behzad, Tina, et al.
Published: (2025)
Stability and Multigroup Fairness in Ranking with Uncertain Predictions
by: Devic, Siddartha, et al.
Published: (2024)
by: Devic, Siddartha, et al.
Published: (2024)
When is Multicalibration Post-Processing Necessary?
by: Hansen, Dutch, et al.
Published: (2024)
by: Hansen, Dutch, et al.
Published: (2024)
Transductive Learning Is Compact
by: Asilis, Julian, et al.
Published: (2024)
by: Asilis, Julian, et al.
Published: (2024)
Are LLM Decisions Faithful to Verbal Confidence?
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025)
by: He, Keyu, et al.
Published: (2025)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024)
by: Liu, Ollie, et al.
Published: (2024)
Auditability and the Landscape of Distance to Multicalibration
by: Derhake, Nathan, et al.
Published: (2025)
by: Derhake, Nathan, et al.
Published: (2025)
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
by: Akgül, Ömer Faruk, et al.
Published: (2026)
by: Akgül, Ömer Faruk, et al.
Published: (2026)
Proper Learnability and the Role of Unlabeled Data
by: Asilis, Julian, et al.
Published: (2025)
by: Asilis, Julian, et al.
Published: (2025)
Regularization and Optimal Multiclass Learning
by: Asilis, Julian, et al.
Published: (2023)
by: Asilis, Julian, et al.
Published: (2023)
LLM Unlearning Without an Expert Curated Dataset
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
by: Srinivasan, Tejas, et al.
Published: (2024)
by: Srinivasan, Tejas, et al.
Published: (2024)
When Parts Are Greater Than Sums: Individual LLM Components Can Outperform Full Models
by: Chang, Ting-Yun, et al.
Published: (2024)
by: Chang, Ting-Yun, et al.
Published: (2024)
Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems
by: İnan, Mert, et al.
Published: (2025)
by: İnan, Mert, et al.
Published: (2025)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
Confidence Should Be Calibrated More Than One Turn Deep
by: Zhang, Zhaohan, et al.
Published: (2026)
by: Zhang, Zhaohan, et al.
Published: (2026)
Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
by: Moore, Kyle, et al.
Published: (2025)
by: Moore, Kyle, et al.
Published: (2025)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
by: Zhou, Tianyi, et al.
Published: (2024)
by: Zhou, Tianyi, et al.
Published: (2024)
Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization
by: Kezar, Lee, et al.
Published: (2025)
by: Kezar, Lee, et al.
Published: (2025)
Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection
by: Anwar, Abrar, et al.
Published: (2025)
by: Anwar, Abrar, et al.
Published: (2025)
Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two Benchmarks
by: Chang, Ting-Yun, et al.
Published: (2023)
by: Chang, Ting-Yun, et al.
Published: (2023)
LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
by: Akgül, Ömer Faruk, et al.
Published: (2025)
by: Akgül, Ömer Faruk, et al.
Published: (2025)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
by: Shojaei, Mostafa Faghih, et al.
Published: (2025)
by: Shojaei, Mostafa Faghih, et al.
Published: (2025)
Limitations on Accurate, Trusted, Human-level Reasoning
by: Panigrahy, Rina, et al.
Published: (2025)
by: Panigrahy, Rina, et al.
Published: (2025)
Uncertainty Quantification for LLM Function-Calling
by: Ye, Zihuiwen, et al.
Published: (2026)
by: Ye, Zihuiwen, et al.
Published: (2026)
Should LLM Safety Be More Than Refusing Harmful Instructions?
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
Words that make SENSE: Sensorimotor Norms in Learned Lexical Token Representations
by: Gupta, Abhinav, et al.
Published: (2026)
by: Gupta, Abhinav, et al.
Published: (2026)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
Compare without Despair: Reliable Preference Evaluation with Generation Separability
by: Ghosh, Sayan, et al.
Published: (2024)
by: Ghosh, Sayan, et al.
Published: (2024)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey
by: Liu, Xiaoou, et al.
Published: (2025)
by: Liu, Xiaoou, et al.
Published: (2025)
Simultaneous Swap Regret Minimization via KL-Calibration
by: Luo, Haipeng, et al.
Published: (2025)
by: Luo, Haipeng, et al.
Published: (2025)
SlimPajama-DC: Understanding Data Combinations for LLM Training
by: Shen, Zhiqiang, et al.
Published: (2023)
by: Shen, Zhiqiang, et al.
Published: (2023)
Can VLMs Recall Factual Associations From Visual References?
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Similar Items
-
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
by: Srinivasan, Tejas, et al.
Published: (2025) -
WinoViz: Probing Visual Properties of Objects Under Different States
by: Jin, Woojeong, et al.
Published: (2024) -
An External Fairness Evaluation of LinkedIn Talent Search
by: Behzad, Tina, et al.
Published: (2025) -
Stability and Multigroup Fairness in Ranking with Uncertain Predictions
by: Devic, Siddartha, et al.
Published: (2024) -
When is Multicalibration Post-Processing Necessary?
by: Hansen, Dutch, et al.
Published: (2024)