Salvato in:
| Autori principali: | Zhang, Zhehao, Xu, Weijie, Wu, Fanyou, Reddy, Chandan K. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2505.08054 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
di: Si, Shengyun, et al.
Pubblicazione: (2025)
di: Si, Shengyun, et al.
Pubblicazione: (2025)
Sequence-Level Certainty Reduces Hallucination In Knowledge-Grounded Dialogue Generation
di: Wan, Yixin, et al.
Pubblicazione: (2023)
di: Wan, Yixin, et al.
Pubblicazione: (2023)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
di: Yuan, Youliang, et al.
Pubblicazione: (2024)
di: Yuan, Youliang, et al.
Pubblicazione: (2024)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
Synthesizing Conversations from Unlabeled Documents using Automatic Response Segmentation
di: Wu, Fanyou, et al.
Pubblicazione: (2024)
di: Wu, Fanyou, et al.
Pubblicazione: (2024)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
di: Kabra, Sanchit, et al.
Pubblicazione: (2025)
di: Kabra, Sanchit, et al.
Pubblicazione: (2025)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
di: Fang, Xi, et al.
Pubblicazione: (2025)
di: Fang, Xi, et al.
Pubblicazione: (2025)
HR-Agent: A Task-Oriented Dialogue (TOD) LLM Agent Tailored for HR Applications
di: Xu, Weijie, et al.
Pubblicazione: (2024)
di: Xu, Weijie, et al.
Pubblicazione: (2024)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
di: Jiang, Eric Hanchen, et al.
Pubblicazione: (2025)
di: Jiang, Eric Hanchen, et al.
Pubblicazione: (2025)
Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive Refinement
di: Lu, Yuxiao, et al.
Pubblicazione: (2026)
di: Lu, Yuxiao, et al.
Pubblicazione: (2026)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
di: Zhang, Haonan, et al.
Pubblicazione: (2025)
di: Zhang, Haonan, et al.
Pubblicazione: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective
di: Xu, Weijie, et al.
Pubblicazione: (2025)
di: Xu, Weijie, et al.
Pubblicazione: (2025)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
di: Cui, Justin, et al.
Pubblicazione: (2024)
di: Cui, Justin, et al.
Pubblicazione: (2024)
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
di: Abhyankar, Nikhil, et al.
Pubblicazione: (2025)
di: Abhyankar, Nikhil, et al.
Pubblicazione: (2025)
Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries
di: Noever, David, et al.
Pubblicazione: (2025)
di: Noever, David, et al.
Pubblicazione: (2025)
Mitigating Selection Bias with Node Pruning and Auxiliary Options
di: Choi, Hyeong Kyu, et al.
Pubblicazione: (2024)
di: Choi, Hyeong Kyu, et al.
Pubblicazione: (2024)
From Oracle to Noisy Context: Mitigating Contextual Exposure Bias in Speech-LLMs
di: Guo, Xiaoyong, et al.
Pubblicazione: (2026)
di: Guo, Xiaoyong, et al.
Pubblicazione: (2026)
H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables
di: Abhyankar, Nikhil, et al.
Pubblicazione: (2024)
di: Abhyankar, Nikhil, et al.
Pubblicazione: (2024)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
di: Sun, Zhongxiang, et al.
Pubblicazione: (2026)
di: Sun, Zhongxiang, et al.
Pubblicazione: (2026)
SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions
di: Xu, Weijie, et al.
Pubblicazione: (2025)
di: Xu, Weijie, et al.
Pubblicazione: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
di: Chhabra, Vishnu Kabir, et al.
Pubblicazione: (2025)
di: Chhabra, Vishnu Kabir, et al.
Pubblicazione: (2025)
RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
di: Abhyankar, Nikhil, et al.
Pubblicazione: (2025)
di: Abhyankar, Nikhil, et al.
Pubblicazione: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
di: von Recum, Alexander, et al.
Pubblicazione: (2024)
di: von Recum, Alexander, et al.
Pubblicazione: (2024)
Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
di: Hyun, Lee, et al.
Pubblicazione: (2025)
di: Hyun, Lee, et al.
Pubblicazione: (2025)
Where Do Reasoning Models Refuse?
di: Yamaguchi, Kureha, et al.
Pubblicazione: (2025)
di: Yamaguchi, Kureha, et al.
Pubblicazione: (2025)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
di: Xu, Zhangchen, et al.
Pubblicazione: (2025)
di: Xu, Zhangchen, et al.
Pubblicazione: (2025)
When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
di: Anonto, Riad Ahmed, et al.
Pubblicazione: (2025)
di: Anonto, Riad Ahmed, et al.
Pubblicazione: (2025)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
di: Asif, Sadia, et al.
Pubblicazione: (2026)
di: Asif, Sadia, et al.
Pubblicazione: (2026)
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
di: Wu, Di, et al.
Pubblicazione: (2026)
di: Wu, Di, et al.
Pubblicazione: (2026)
Aligned at the Start: Conceptual Groupings in LLM Embeddings
di: Khatir, Mehrdad, et al.
Pubblicazione: (2024)
di: Khatir, Mehrdad, et al.
Pubblicazione: (2024)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
Improving Grammatical Error Correction via Contextual Data Augmentation
di: Wang, Yixuan, et al.
Pubblicazione: (2024)
di: Wang, Yixuan, et al.
Pubblicazione: (2024)
Does Refusal Training in LLMs Generalize to the Past Tense?
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback
di: Pucci, Giulia, et al.
Pubblicazione: (2026)
di: Pucci, Giulia, et al.
Pubblicazione: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
Learn to Disguise: Avoid Refusal Responses in LLM's Defense via a Multi-agent Attacker-Disguiser Game
di: Xu, Qianqiao, et al.
Pubblicazione: (2024)
di: Xu, Qianqiao, et al.
Pubblicazione: (2024)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
di: García-Ferrero, Iker, et al.
Pubblicazione: (2025)
di: García-Ferrero, Iker, et al.
Pubblicazione: (2025)
Mission Impossible: Feedback-Guided Dynamic Interactive Planning for Improving Reasoning on LLMs
di: Yan, Dong, et al.
Pubblicazione: (2025)
di: Yan, Dong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
di: Si, Shengyun, et al.
Pubblicazione: (2025) -
Sequence-Level Certainty Reduces Hallucination In Knowledge-Grounded Dialogue Generation
di: Wan, Yixin, et al.
Pubblicazione: (2023) -
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
di: Yuan, Youliang, et al.
Pubblicazione: (2024) -
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
di: Zhang, Zhehao, et al.
Pubblicazione: (2025) -
Synthesizing Conversations from Unlabeled Documents using Automatic Response Segmentation
di: Wu, Fanyou, et al.
Pubblicazione: (2024)