Saved in:
| Main Authors: | Lin, Shuyuan, Duan, Lei, Hughes, Philip, Sheng, Yuxuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.16951 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unanswerability Evaluation for Retrieval Augmented Generation
by: Peng, Xiangyu, et al.
Published: (2024)
by: Peng, Xiangyu, et al.
Published: (2024)
Contextual Candor: Enhancing LLM Trustworthiness Through Hierarchical Unanswerability Detection
by: Robinson, Steven, et al.
Published: (2025)
by: Robinson, Steven, et al.
Published: (2025)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
by: He, Xingwei, et al.
Published: (2024)
by: He, Xingwei, et al.
Published: (2024)
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
by: Tan, Chuanyuan, et al.
Published: (2025)
by: Tan, Chuanyuan, et al.
Published: (2025)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)
by: Li, Aaron J., et al.
Published: (2024)
Towards Reliable and Factual Response Generation: Detecting Unanswerable Questions in Information-Seeking Conversations
by: Łajewska, Weronika, et al.
Published: (2024)
by: Łajewska, Weronika, et al.
Published: (2024)
Taming Overconfidence in LLMs: Reward Calibration in RLHF
by: Leng, Jixuan, et al.
Published: (2024)
by: Leng, Jixuan, et al.
Published: (2024)
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF
by: Zheng, Chen, et al.
Published: (2024)
by: Zheng, Chen, et al.
Published: (2024)
RetinaQA: A Robust Knowledge Base Question Answering Model for both Answerable and Unanswerable Questions
by: Faldu, Prayushi, et al.
Published: (2024)
by: Faldu, Prayushi, et al.
Published: (2024)
SynCPKL: Harnessing LLMs to Generate Synthetic Data for Commonsense Persona Knowledge Linking
by: Lin, Kuan-Yen
Published: (2024)
by: Lin, Kuan-Yen
Published: (2024)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
by: Xie, Tengyang, et al.
Published: (2024)
by: Xie, Tengyang, et al.
Published: (2024)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
by: Yu, Tianyu, et al.
Published: (2023)
by: Yu, Tianyu, et al.
Published: (2023)
Towards Federated RLHF with Aggregated Client Preference for LLMs
by: Wu, Feijie, et al.
Published: (2024)
by: Wu, Feijie, et al.
Published: (2024)
I Could've Asked That: Reformulating Unanswerable Questions
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Towards Unbiased Evaluation of Detecting Unanswerable Questions in EHRSQL
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Group Robust Preference Optimization in Reward-free RLHF
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
TrustRAG: Enhancing Robustness and Trustworthiness in Retrieval-Augmented Generation
by: Zhou, Huichi, et al.
Published: (2025)
by: Zhou, Huichi, et al.
Published: (2025)
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
Harmonic LLMs are Trustworthy
by: Kersting, Nicholas S., et al.
Published: (2024)
by: Kersting, Nicholas S., et al.
Published: (2024)
Query Carefully: Detecting the Unanswerables in Text-to-SQL Tasks
by: Saxer, Jasmin, et al.
Published: (2025)
by: Saxer, Jasmin, et al.
Published: (2025)
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
by: Dong, Mingwen, et al.
Published: (2024)
by: Dong, Mingwen, et al.
Published: (2024)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
by: von Recum, Alexander, et al.
Published: (2024)
by: von Recum, Alexander, et al.
Published: (2024)
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
by: Alghamdi, Emad A., et al.
Published: (2024)
by: Alghamdi, Emad A., et al.
Published: (2024)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
by: Sun, Yuhong, et al.
Published: (2024)
by: Sun, Yuhong, et al.
Published: (2024)
Harnessing Collective Intelligence of LLMs for Robust Biomedical QA: A Multi-Model Approach
by: Panou, Dimitra, et al.
Published: (2025)
by: Panou, Dimitra, et al.
Published: (2025)
Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
by: Xiao, Youshao, et al.
Published: (2023)
by: Xiao, Youshao, et al.
Published: (2023)
Iterative Repair with Weak Verifiers for Few-shot Transfer in KBQA with Unanswerability
by: Sawhney, Riya, et al.
Published: (2024)
by: Sawhney, Riya, et al.
Published: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
by: Hong, Junyuan, et al.
Published: (2024)
by: Hong, Junyuan, et al.
Published: (2024)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
by: Srewa, Mahmoud, et al.
Published: (2025)
by: Srewa, Mahmoud, et al.
Published: (2025)
Harnessing Generative LLMs for Enhanced Financial Event Entity Extraction Performance
by: Choi, Soo-joon, et al.
Published: (2025)
by: Choi, Soo-joon, et al.
Published: (2025)
Length Controlled Generation for Black-box LLMs
by: Gu, Yuxuan, et al.
Published: (2024)
by: Gu, Yuxuan, et al.
Published: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
General Exploratory Bonus for Optimistic Exploration in RLHF
by: Li, Wendi, et al.
Published: (2025)
by: Li, Wendi, et al.
Published: (2025)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
Harnessing LLMs for Educational Content-Driven Italian Crossword Generation
by: Zeinalipour, Kamyar, et al.
Published: (2024)
by: Zeinalipour, Kamyar, et al.
Published: (2024)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
Similar Items
-
Unanswerability Evaluation for Retrieval Augmented Generation
by: Peng, Xiangyu, et al.
Published: (2024) -
Contextual Candor: Enhancing LLM Trustworthiness Through Hierarchical Unanswerability Detection
by: Robinson, Steven, et al.
Published: (2025) -
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
by: He, Xingwei, et al.
Published: (2024) -
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024) -
UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
by: Tan, Chuanyuan, et al.
Published: (2025)