Guardado en:
| Autores principales: | Lin, Shuyuan, Duan, Lei, Hughes, Philip, Sheng, Yuxuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2507.16951 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unanswerability Evaluation for Retrieval Augmented Generation
por: Peng, Xiangyu, et al.
Publicado: (2024)
por: Peng, Xiangyu, et al.
Publicado: (2024)
Contextual Candor: Enhancing LLM Trustworthiness Through Hierarchical Unanswerability Detection
por: Robinson, Steven, et al.
Publicado: (2025)
por: Robinson, Steven, et al.
Publicado: (2025)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
por: He, Xingwei, et al.
Publicado: (2024)
por: He, Xingwei, et al.
Publicado: (2024)
Reward-Robust RLHF in LLMs
por: Yan, Yuzi, et al.
Publicado: (2024)
por: Yan, Yuzi, et al.
Publicado: (2024)
UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
por: Tan, Chuanyuan, et al.
Publicado: (2025)
por: Tan, Chuanyuan, et al.
Publicado: (2025)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
por: Li, Aaron J., et al.
Publicado: (2024)
por: Li, Aaron J., et al.
Publicado: (2024)
Towards Reliable and Factual Response Generation: Detecting Unanswerable Questions in Information-Seeking Conversations
por: Łajewska, Weronika, et al.
Publicado: (2024)
por: Łajewska, Weronika, et al.
Publicado: (2024)
Taming Overconfidence in LLMs: Reward Calibration in RLHF
por: Leng, Jixuan, et al.
Publicado: (2024)
por: Leng, Jixuan, et al.
Publicado: (2024)
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF
por: Zheng, Chen, et al.
Publicado: (2024)
por: Zheng, Chen, et al.
Publicado: (2024)
RetinaQA: A Robust Knowledge Base Question Answering Model for both Answerable and Unanswerable Questions
por: Faldu, Prayushi, et al.
Publicado: (2024)
por: Faldu, Prayushi, et al.
Publicado: (2024)
SynCPKL: Harnessing LLMs to Generate Synthetic Data for Commonsense Persona Knowledge Linking
por: Lin, Kuan-Yen
Publicado: (2024)
por: Lin, Kuan-Yen
Publicado: (2024)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
por: Xie, Tengyang, et al.
Publicado: (2024)
por: Xie, Tengyang, et al.
Publicado: (2024)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
por: Yu, Tianyu, et al.
Publicado: (2023)
por: Yu, Tianyu, et al.
Publicado: (2023)
Towards Federated RLHF with Aggregated Client Preference for LLMs
por: Wu, Feijie, et al.
Publicado: (2024)
por: Wu, Feijie, et al.
Publicado: (2024)
I Could've Asked That: Reformulating Unanswerable Questions
por: Zhao, Wenting, et al.
Publicado: (2024)
por: Zhao, Wenting, et al.
Publicado: (2024)
Towards Unbiased Evaluation of Detecting Unanswerable Questions in EHRSQL
por: Yang, Yongjin, et al.
Publicado: (2024)
por: Yang, Yongjin, et al.
Publicado: (2024)
Group Robust Preference Optimization in Reward-free RLHF
por: Ramesh, Shyam Sundhar, et al.
Publicado: (2024)
por: Ramesh, Shyam Sundhar, et al.
Publicado: (2024)
TrustRAG: Enhancing Robustness and Trustworthiness in Retrieval-Augmented Generation
por: Zhou, Huichi, et al.
Publicado: (2025)
por: Zhou, Huichi, et al.
Publicado: (2025)
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
por: Liu, Gabrielle Kaili-May, et al.
Publicado: (2025)
por: Liu, Gabrielle Kaili-May, et al.
Publicado: (2025)
Harmonic LLMs are Trustworthy
por: Kersting, Nicholas S., et al.
Publicado: (2024)
por: Kersting, Nicholas S., et al.
Publicado: (2024)
Query Carefully: Detecting the Unanswerables in Text-to-SQL Tasks
por: Saxer, Jasmin, et al.
Publicado: (2025)
por: Saxer, Jasmin, et al.
Publicado: (2025)
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
por: Dong, Mingwen, et al.
Publicado: (2024)
por: Dong, Mingwen, et al.
Publicado: (2024)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
por: von Recum, Alexander, et al.
Publicado: (2024)
por: von Recum, Alexander, et al.
Publicado: (2024)
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
por: Alghamdi, Emad A., et al.
Publicado: (2024)
por: Alghamdi, Emad A., et al.
Publicado: (2024)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
por: Sun, Yuhong, et al.
Publicado: (2024)
por: Sun, Yuhong, et al.
Publicado: (2024)
Harnessing Collective Intelligence of LLMs for Robust Biomedical QA: A Multi-Model Approach
por: Panou, Dimitra, et al.
Publicado: (2025)
por: Panou, Dimitra, et al.
Publicado: (2025)
Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs
por: Huang, Yue, et al.
Publicado: (2026)
por: Huang, Yue, et al.
Publicado: (2026)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
por: Xiao, Youshao, et al.
Publicado: (2023)
por: Xiao, Youshao, et al.
Publicado: (2023)
Iterative Repair with Weak Verifiers for Few-shot Transfer in KBQA with Unanswerability
por: Sawhney, Riya, et al.
Publicado: (2024)
por: Sawhney, Riya, et al.
Publicado: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
por: Dong, Hanze, et al.
Publicado: (2024)
por: Dong, Hanze, et al.
Publicado: (2024)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
por: Hong, Junyuan, et al.
Publicado: (2024)
por: Hong, Junyuan, et al.
Publicado: (2024)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
por: Srewa, Mahmoud, et al.
Publicado: (2025)
por: Srewa, Mahmoud, et al.
Publicado: (2025)
Harnessing Generative LLMs for Enhanced Financial Event Entity Extraction Performance
por: Choi, Soo-joon, et al.
Publicado: (2025)
por: Choi, Soo-joon, et al.
Publicado: (2025)
Length Controlled Generation for Black-box LLMs
por: Gu, Yuxuan, et al.
Publicado: (2024)
por: Gu, Yuxuan, et al.
Publicado: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
General Exploratory Bonus for Optimistic Exploration in RLHF
por: Li, Wendi, et al.
Publicado: (2025)
por: Li, Wendi, et al.
Publicado: (2025)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
por: Banayeeanzade, Amin, et al.
Publicado: (2025)
por: Banayeeanzade, Amin, et al.
Publicado: (2025)
Harnessing LLMs for Educational Content-Driven Italian Crossword Generation
por: Zeinalipour, Kamyar, et al.
Publicado: (2024)
por: Zeinalipour, Kamyar, et al.
Publicado: (2024)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
por: Xiao, Yuxin, et al.
Publicado: (2024)
por: Xiao, Yuxin, et al.
Publicado: (2024)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
por: Ji, Jiaming, et al.
Publicado: (2024)
por: Ji, Jiaming, et al.
Publicado: (2024)
Ejemplares similares
-
Unanswerability Evaluation for Retrieval Augmented Generation
por: Peng, Xiangyu, et al.
Publicado: (2024) -
Contextual Candor: Enhancing LLM Trustworthiness Through Hierarchical Unanswerability Detection
por: Robinson, Steven, et al.
Publicado: (2025) -
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
por: He, Xingwei, et al.
Publicado: (2024) -
Reward-Robust RLHF in LLMs
por: Yan, Yuzi, et al.
Publicado: (2024) -
UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
por: Tan, Chuanyuan, et al.
Publicado: (2025)