UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions
Fuente:
arXiv
Salvato in:
| Autori principali: | Tan, Chuanyuan, Shao, Wenbiao, Xiong, Hao, Zhu, Tong, Liu, Zhenhua, Shi, Kai, Chen, Wenliang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
Is Fine-Tuning an Effective Solution? Reassessing Knowledge Editing for Unstructured Data
di: Xiong, Hao, et al.
Pubblicazione: (2025)
di: Xiong, Hao, et al.
Pubblicazione: (2025)
Probing Language Models for Pre-training Data Detection
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models
di: Wu, Mengsong, et al.
Pubblicazione: (2025)
di: Wu, Mengsong, et al.
Pubblicazione: (2025)
Seal-Tools: Self-Instruct Tool Learning Dataset for Agent Tuning and Detailed Benchmark
di: Wu, Mengsong, et al.
Pubblicazione: (2024)
di: Wu, Mengsong, et al.
Pubblicazione: (2024)
Towards Reliable and Factual Response Generation: Detecting Unanswerable Questions in Information-Seeking Conversations
di: Łajewska, Weronika, et al.
Pubblicazione: (2024)
di: Łajewska, Weronika, et al.
Pubblicazione: (2024)
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
Unanswerability Evaluation for Retrieval Augmented Generation
di: Peng, Xiangyu, et al.
Pubblicazione: (2024)
di: Peng, Xiangyu, et al.
Pubblicazione: (2024)
Towards Unbiased Evaluation of Detecting Unanswerable Questions in EHRSQL
di: Yang, Yongjin, et al.
Pubblicazione: (2024)
di: Yang, Yongjin, et al.
Pubblicazione: (2024)
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
di: Liu, Zhenhua, et al.
Pubblicazione: (2025)
di: Liu, Zhenhua, et al.
Pubblicazione: (2025)
RetinaQA: A Robust Knowledge Base Question Answering Model for both Answerable and Unanswerable Questions
di: Faldu, Prayushi, et al.
Pubblicazione: (2024)
di: Faldu, Prayushi, et al.
Pubblicazione: (2024)
NesTools: A Dataset for Evaluating Nested Tool Learning Abilities of Large Language Models
di: Han, Han, et al.
Pubblicazione: (2024)
di: Han, Han, et al.
Pubblicazione: (2024)
Improving Factuality in LLMs via Inference-Time Knowledge Graph Construction
di: Wu, Shanglin, et al.
Pubblicazione: (2025)
di: Wu, Shanglin, et al.
Pubblicazione: (2025)
Transparentize the Internal and External Knowledge Utilization in LLMs with Trustworthy Citation
di: Shen, Jiajun, et al.
Pubblicazione: (2025)
di: Shen, Jiajun, et al.
Pubblicazione: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
I Could've Asked That: Reformulating Unanswerable Questions
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2025)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2025)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
di: Yuan, Jiaqing, et al.
Pubblicazione: (2024)
di: Yuan, Jiaqing, et al.
Pubblicazione: (2024)
CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
di: Zong, Qing, et al.
Pubblicazione: (2024)
di: Zong, Qing, et al.
Pubblicazione: (2024)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
di: Wang, Kai, et al.
Pubblicazione: (2026)
di: Wang, Kai, et al.
Pubblicazione: (2026)
Exploring the Generalizability of Factual Hallucination Mitigation via Enhancing Precise Knowledge Utilization
di: Zhang, Siyuan, et al.
Pubblicazione: (2025)
di: Zhang, Siyuan, et al.
Pubblicazione: (2025)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
di: He, Xingwei, et al.
Pubblicazione: (2024)
di: He, Xingwei, et al.
Pubblicazione: (2024)
Persuasion Tokens for Editing Factual Knowledge in LLMs
di: Youssef, Paul, et al.
Pubblicazione: (2026)
di: Youssef, Paul, et al.
Pubblicazione: (2026)
Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs
di: Lin, Shuyuan, et al.
Pubblicazione: (2025)
di: Lin, Shuyuan, et al.
Pubblicazione: (2025)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
di: Gu, Yuzhe, et al.
Pubblicazione: (2025)
di: Gu, Yuzhe, et al.
Pubblicazione: (2025)
FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
di: Liu, Yihong, et al.
Pubblicazione: (2026)
di: Liu, Yihong, et al.
Pubblicazione: (2026)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
di: Hu, Xuming, et al.
Pubblicazione: (2024)
di: Hu, Xuming, et al.
Pubblicazione: (2024)
Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation
di: Dang, Renfei, et al.
Pubblicazione: (2025)
di: Dang, Renfei, et al.
Pubblicazione: (2025)
Rejection Improves Reliability: Training LLMs to Refuse Unknown Questions Using RL from Knowledge Feedback
di: Xu, Hongshen, et al.
Pubblicazione: (2024)
di: Xu, Hongshen, et al.
Pubblicazione: (2024)
AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering
di: Wang, Yuxin, et al.
Pubblicazione: (2026)
di: Wang, Yuxin, et al.
Pubblicazione: (2026)
Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
di: Xu, Derong, et al.
Pubblicazione: (2024)
di: Xu, Derong, et al.
Pubblicazione: (2024)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
di: Gu, Jihao, et al.
Pubblicazione: (2025)
di: Gu, Jihao, et al.
Pubblicazione: (2025)
Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge
di: Liu, Genglin, et al.
Pubblicazione: (2023)
di: Liu, Genglin, et al.
Pubblicazione: (2023)
Knowledgeable In-Context Tuning: Exploring and Exploiting Factual Knowledge for In-Context Learning
di: Wang, Jianing, et al.
Pubblicazione: (2023)
di: Wang, Jianing, et al.
Pubblicazione: (2023)
When Language Shapes Thought: Cross-Lingual Transfer of Factual Knowledge in Question Answering
di: Kang, Eojin, et al.
Pubblicazione: (2025)
di: Kang, Eojin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
di: Liu, Zhenhua, et al.
Pubblicazione: (2024) -
Is Fine-Tuning an Effective Solution? Reassessing Knowledge Editing for Unstructured Data
di: Xiong, Hao, et al.
Pubblicazione: (2025) -
Probing Language Models for Pre-training Data Detection
di: Liu, Zhenhua, et al.
Pubblicazione: (2024) -
Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models
di: Wu, Mengsong, et al.
Pubblicazione: (2025) -
Seal-Tools: Self-Instruct Tool Learning Dataset for Agent Tuning and Detailed Benchmark
di: Wu, Mengsong, et al.
Pubblicazione: (2024)