Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Jin-Seop, Lee, SungJoon, Jung, SeongJun, Li, Boyang, Lee, Jee-Hyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2025)
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2025)
BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
von: Kim, DoYoung, et al.
Veröffentlicht: (2025)
von: Kim, DoYoung, et al.
Veröffentlicht: (2025)
Stabilizing Open-Set Test-Time Adaptation via Primary-Auxiliary Filtering and Knowledge-Integrated Prediction
von: Lee, Byung-Joon, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Joon, et al.
Veröffentlicht: (2025)
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
von: Lee, Joohyeon, et al.
Veröffentlicht: (2025)
von: Lee, Joohyeon, et al.
Veröffentlicht: (2025)
DomCLP: Domain-wise Contrastive Learning with Prototype Mixup for Unsupervised Domain Generalization
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2024)
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2024)
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph
von: Lee, Jihyung, et al.
Veröffentlicht: (2025)
von: Lee, Jihyung, et al.
Veröffentlicht: (2025)
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
von: Na, CheolWon, et al.
Veröffentlicht: (2025)
von: Na, CheolWon, et al.
Veröffentlicht: (2025)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026)
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026)
State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models
von: Lee, TK
Veröffentlicht: (2025)
von: Lee, TK
Veröffentlicht: (2025)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
von: Lan, Wenhao, et al.
Veröffentlicht: (2026)
von: Lan, Wenhao, et al.
Veröffentlicht: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
von: Asif, Sadia, et al.
Veröffentlicht: (2026)
von: Asif, Sadia, et al.
Veröffentlicht: (2026)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
von: Lim, Jiyoung, et al.
Veröffentlicht: (2026)
von: Lim, Jiyoung, et al.
Veröffentlicht: (2026)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
von: Moon, WonJun, et al.
Veröffentlicht: (2023)
von: Moon, WonJun, et al.
Veröffentlicht: (2023)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
Understanding Refusal in Language Models with Sparse Autoencoders
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
von: Bae, Suyoung, et al.
Veröffentlicht: (2026)
von: Bae, Suyoung, et al.
Veröffentlicht: (2026)
Information Density Enhancement Using Lossy Compression in DNA Data Storage
von: Seongjun Seo, et al.
Veröffentlicht: (2024)
von: Seongjun Seo, et al.
Veröffentlicht: (2024)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
von: Dabas, Mahavir, et al.
Veröffentlicht: (2025)
von: Dabas, Mahavir, et al.
Veröffentlicht: (2025)
Refusal Speech
von: Nhatuve, Diocleciano, et al.
Veröffentlicht: (2021)
von: Nhatuve, Diocleciano, et al.
Veröffentlicht: (2021)
Refusals of noncitizenship
von: Peter Nyers
Veröffentlicht: (2024)
von: Peter Nyers
Veröffentlicht: (2024)
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
von: Zhu, Runchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Runchuan, et al.
Veröffentlicht: (2025)
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
von: Kim, Su-Hyeon, et al.
Veröffentlicht: (2026)
von: Kim, Su-Hyeon, et al.
Veröffentlicht: (2026)
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
Signs of the Great Refusal
von: Siegel, Tedd
Veröffentlicht: (2023)
von: Siegel, Tedd
Veröffentlicht: (2023)
Refusal Dark Matter
von: Christopher Love
Veröffentlicht: (2026)
von: Christopher Love
Veröffentlicht: (2026)
Lithuanian Refusals and Politeness
von: Donata Katinaitė
Veröffentlicht: (2021)
von: Donata Katinaitė
Veröffentlicht: (2021)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
von: Jain, Neel, et al.
Veröffentlicht: (2024)
von: Jain, Neel, et al.
Veröffentlicht: (2024)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
von: Collu, Matteo Gioele, et al.
Veröffentlicht: (2026)
von: Collu, Matteo Gioele, et al.
Veröffentlicht: (2026)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
von: Song, Maojia, et al.
Veröffentlicht: (2024)
von: Song, Maojia, et al.
Veröffentlicht: (2024)
Domain-Aware Fine-Tuning: Enhancing Neural Network Adaptability
von: Ha, Seokhyeon, et al.
Veröffentlicht: (2023)
von: Ha, Seokhyeon, et al.
Veröffentlicht: (2023)
RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
von: Nguyen, Tuan T., et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan T., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2025) -
BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
von: Kim, DoYoung, et al.
Veröffentlicht: (2025) -
Stabilizing Open-Set Test-Time Adaptation via Primary-Auxiliary Filtering and Knowledge-Integrated Prediction
von: Lee, Byung-Joon, et al.
Veröffentlicht: (2025) -
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
von: Lee, Joohyeon, et al.
Veröffentlicht: (2025) -
DomCLP: Domain-wise Contrastive Learning with Prototype Mixup for Unsupervised Domain Generalization
von: Lee, Jin-Seop, et al.
Veröffentlicht: (2024)