When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Xinyu, Jin, Chang, Eickhoff, Carsten, Guo, Zhijiang, Bahrainian, Seyed Ali |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Enhancing Retrieval-Augmented Generation: A Study of Best Practices
por: Li, Siran, et al.
Publicado: (2025)
por: Li, Siran, et al.
Publicado: (2025)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
por: Braun, Joschka, et al.
Publicado: (2025)
por: Braun, Joschka, et al.
Publicado: (2025)
Navigating through the hidden embedding space: steering LLMs to improve mental health assessment
por: Ravenda, Federico, et al.
Publicado: (2025)
por: Ravenda, Federico, et al.
Publicado: (2025)
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
por: Machcha, Sravanthi, et al.
Publicado: (2026)
por: Machcha, Sravanthi, et al.
Publicado: (2026)
Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice
por: Ravenda, Federico, et al.
Publicado: (2025)
por: Ravenda, Federico, et al.
Publicado: (2025)
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention
por: Sun, Yuxi, et al.
Publicado: (2025)
por: Sun, Yuxi, et al.
Publicado: (2025)
Teaching LLMs to Abstain via Fine-Grained Semantic Confidence Reward
por: An, Hao, et al.
Publicado: (2025)
por: An, Hao, et al.
Publicado: (2025)
Stable Anisotropic Regularization
por: Rudman, William, et al.
Publicado: (2023)
por: Rudman, William, et al.
Publicado: (2023)
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
por: Nakshatri, Nishanth Sridhar, et al.
Publicado: (2025)
por: Nakshatri, Nishanth Sridhar, et al.
Publicado: (2025)
Outlier Dimensions Encode Task-Specific Knowledge
por: Rudman, William, et al.
Publicado: (2023)
por: Rudman, William, et al.
Publicado: (2023)
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
por: Merullo, Jack, et al.
Publicado: (2024)
por: Merullo, Jack, et al.
Publicado: (2024)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
por: Wei, Shaohang, et al.
Publicado: (2025)
por: Wei, Shaohang, et al.
Publicado: (2025)
A Survey on LLM-Assisted Clinical Trial Recruitment
por: Ghosh, Shrestha, et al.
Publicado: (2025)
por: Ghosh, Shrestha, et al.
Publicado: (2025)
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
por: Liu, Shiqi, et al.
Publicado: (2026)
por: Liu, Shiqi, et al.
Publicado: (2026)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
por: Xiao, Yunpeng, et al.
Publicado: (2025)
por: Xiao, Yunpeng, et al.
Publicado: (2025)
Beyond QA Pairs: Assessing Parameter-Efficient Fine-Tuning for Fact Embedding in LLMs
por: Ratnakar, Shivam, et al.
Publicado: (2025)
por: Ratnakar, Shivam, et al.
Publicado: (2025)
Learning When to Retrieve, What to Rewrite, and How to Respond in Conversational QA
por: Roy, Nirmal, et al.
Publicado: (2024)
por: Roy, Nirmal, et al.
Publicado: (2024)
The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling
por: Zhang, Ruochen, et al.
Publicado: (2024)
por: Zhang, Ruochen, et al.
Publicado: (2024)
Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning
por: Fan, Wei, et al.
Publicado: (2026)
por: Fan, Wei, et al.
Publicado: (2026)
Can LLMs Capture Human Preferences?
por: Goli, Ali, et al.
Publicado: (2023)
por: Goli, Ali, et al.
Publicado: (2023)
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
por: Zhai, Skylar, et al.
Publicado: (2026)
por: Zhai, Skylar, et al.
Publicado: (2026)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
por: Amin, Hasan, et al.
Publicado: (2026)
por: Amin, Hasan, et al.
Publicado: (2026)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
por: Baan, Joris, et al.
Publicado: (2026)
por: Baan, Joris, et al.
Publicado: (2026)
Learning to Correct for QA Reasoning with Black-box LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
por: Bhatia, Gagan, et al.
Publicado: (2024)
por: Bhatia, Gagan, et al.
Publicado: (2024)
LLMs Can Plan Only If We Tell Them
por: Sel, Bilgehan, et al.
Publicado: (2025)
por: Sel, Bilgehan, et al.
Publicado: (2025)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
por: Tie, Guiyao, et al.
Publicado: (2025)
por: Tie, Guiyao, et al.
Publicado: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
por: Wei, Jianhui, et al.
Publicado: (2025)
por: Wei, Jianhui, et al.
Publicado: (2025)
What External Knowledge is Preferred by LLMs? Characterizing and Exploring Chain of Evidence in Imperfect Context for Multi-Hop QA
por: Chang, Zhiyuan, et al.
Publicado: (2024)
por: Chang, Zhiyuan, et al.
Publicado: (2024)
Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
por: Liu, Yesheng, et al.
Publicado: (2025)
por: Liu, Yesheng, et al.
Publicado: (2025)
Enhancing the QA Model through a Multi-domain Debiasing Framework
por: Wang, Yuefeng, et al.
Publicado: (2026)
por: Wang, Yuefeng, et al.
Publicado: (2026)
Less is More: Resource-Efficient Low-Rank Adaptation
por: Tian, Chunlin, et al.
Publicado: (2025)
por: Tian, Chunlin, et al.
Publicado: (2025)
HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-Tuning
por: Tian, Chunlin, et al.
Publicado: (2024)
por: Tian, Chunlin, et al.
Publicado: (2024)
InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?
por: Ka, Keummin, et al.
Publicado: (2025)
por: Ka, Keummin, et al.
Publicado: (2025)
Retracing the Past: LLMs Emit Training Data When They Get Lost
por: Ko, Myeongseob, et al.
Publicado: (2025)
por: Ko, Myeongseob, et al.
Publicado: (2025)
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
por: Song, Maojia, et al.
Publicado: (2025)
por: Song, Maojia, et al.
Publicado: (2025)
What Layers When: Learning to Skip Compute in LLMs with Residual Gates
por: Laitenberger, Filipe, et al.
Publicado: (2025)
por: Laitenberger, Filipe, et al.
Publicado: (2025)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
por: Wang, Minzheng, et al.
Publicado: (2024)
por: Wang, Minzheng, et al.
Publicado: (2024)
Beyond Specialization: Benchmarking LLMs for Transliteration of Indian Languages
por: Azam, Gulfarogh, et al.
Publicado: (2025)
por: Azam, Gulfarogh, et al.
Publicado: (2025)
The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
por: Abdullahi, Tassallah, et al.
Publicado: (2026)
Ejemplares similares
-
Enhancing Retrieval-Augmented Generation: A Study of Best Practices
por: Li, Siran, et al.
Publicado: (2025) -
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
por: Braun, Joschka, et al.
Publicado: (2025) -
Navigating through the hidden embedding space: steering LLMs to improve mental health assessment
por: Ravenda, Federico, et al.
Publicado: (2025) -
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
por: Machcha, Sravanthi, et al.
Publicado: (2026) -
Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice
por: Ravenda, Federico, et al.
Publicado: (2025)