Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive Refinement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Yuxiao, Xu, Lin, Sun, Yang, Li, Wenjun, Shi, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
von: Fastowski, Alina, et al.
Veröffentlicht: (2025)
von: Fastowski, Alina, et al.
Veröffentlicht: (2025)
The Energy of Falsehood: Detecting Hallucinations via Diffusion Model Likelihoods
von: Gautam, Arpit Singh, et al.
Veröffentlicht: (2026)
von: Gautam, Arpit Singh, et al.
Veröffentlicht: (2026)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
von: Cui, Justin, et al.
Veröffentlicht: (2024)
von: Cui, Justin, et al.
Veröffentlicht: (2024)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries
von: Noever, David, et al.
Veröffentlicht: (2025)
von: Noever, David, et al.
Veröffentlicht: (2025)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
Refine Knowledge of Large Language Models via Adaptive Contrastive Learning
von: Li, Yinghui, et al.
Veröffentlicht: (2025)
von: Li, Yinghui, et al.
Veröffentlicht: (2025)
Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2025)
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2025)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
von: He, Jinwen, et al.
Veröffentlicht: (2023)
von: He, Jinwen, et al.
Veröffentlicht: (2023)
Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
von: Li, Jidong, et al.
Veröffentlicht: (2025)
von: Li, Jidong, et al.
Veröffentlicht: (2025)
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
von: Liu, Ying, et al.
Veröffentlicht: (2025)
von: Liu, Ying, et al.
Veröffentlicht: (2025)
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Echo: Learning from Experience Data via User-Driven Refinement
von: Dong, Hande, et al.
Veröffentlicht: (2026)
von: Dong, Hande, et al.
Veröffentlicht: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
PET-SQL: A Prompt-Enhanced Two-Round Refinement of Text-to-SQL with Cross-consistency
von: Li, Zhishuai, et al.
Veröffentlicht: (2024)
von: Li, Zhishuai, et al.
Veröffentlicht: (2024)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
von: Wei, Zhepei, et al.
Veröffentlicht: (2025)
von: Wei, Zhepei, et al.
Veröffentlicht: (2025)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
Learn to Disguise: Avoid Refusal Responses in LLM's Defense via a Multi-agent Attacker-Disguiser Game
von: Xu, Qianqiao, et al.
Veröffentlicht: (2024)
von: Xu, Qianqiao, et al.
Veröffentlicht: (2024)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
von: Chhabra, Vishnu Kabir, et al.
Veröffentlicht: (2025)
Hallucination-Resistant Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
von: Yang, Yupei, et al.
Veröffentlicht: (2025)
von: Yang, Yupei, et al.
Veröffentlicht: (2025)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
von: von Recum, Alexander, et al.
Veröffentlicht: (2024)
von: von Recum, Alexander, et al.
Veröffentlicht: (2024)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
von: Zhang, Haonan, et al.
Veröffentlicht: (2025)
von: Zhang, Haonan, et al.
Veröffentlicht: (2025)
Measuring and Eliminating Refusals in Military Large Language Models
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
DeepRefine: Agent-Compiled Knowledge Refinement via Reinforcement Learning
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
von: Cao, Lang
Veröffentlicht: (2023)
von: Cao, Lang
Veröffentlicht: (2023)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2025)
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
von: Zhai, Skylar, et al.
Veröffentlicht: (2026)
von: Zhai, Skylar, et al.
Veröffentlicht: (2026)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
TruthStance: An Annotated Dataset of Conversations on Truth Social
von: Ameen, Fathima, et al.
Veröffentlicht: (2026)
von: Ameen, Fathima, et al.
Veröffentlicht: (2026)
AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2023)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2023)
Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
von: Lu, Yuxiao, et al.
Veröffentlicht: (2024)
von: Lu, Yuxiao, et al.
Veröffentlicht: (2024)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
von: Fastowski, Alina, et al.
Veröffentlicht: (2025) -
The Energy of Falsehood: Detecting Hallucinations via Diffusion Model Likelihoods
von: Gautam, Arpit Singh, et al.
Veröffentlicht: (2026) -
OR-Bench: An Over-Refusal Benchmark for Large Language Models
von: Cui, Justin, et al.
Veröffentlicht: (2024) -
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025) -
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)