Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lee, Jin-Seop, Lee, SungJoon, Jung, SeongJun, Li, Boyang, Lee, Jee-Hyong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
par: Lee, Jin-Seop, et autres
Publié: (2025)
par: Lee, Jin-Seop, et autres
Publié: (2025)
BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
par: Kim, DoYoung, et autres
Publié: (2025)
par: Kim, DoYoung, et autres
Publié: (2025)
Stabilizing Open-Set Test-Time Adaptation via Primary-Auxiliary Filtering and Knowledge-Integrated Prediction
par: Lee, Byung-Joon, et autres
Publié: (2025)
par: Lee, Byung-Joon, et autres
Publié: (2025)
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
par: Lee, Joohyeon, et autres
Publié: (2025)
par: Lee, Joohyeon, et autres
Publié: (2025)
DomCLP: Domain-wise Contrastive Learning with Prototype Mixup for Unsupervised Domain Generalization
par: Lee, Jin-Seop, et autres
Publié: (2024)
par: Lee, Jin-Seop, et autres
Publié: (2024)
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph
par: Lee, Jihyung, et autres
Publié: (2025)
par: Lee, Jihyung, et autres
Publié: (2025)
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
par: Na, CheolWon, et autres
Publié: (2025)
par: Na, CheolWon, et autres
Publié: (2025)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
par: Ren, Xuancheng, et autres
Publié: (2026)
par: Ren, Xuancheng, et autres
Publié: (2026)
State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models
par: Lee, TK
Publié: (2025)
par: Lee, TK
Publié: (2025)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
par: Lan, Wenhao, et autres
Publié: (2026)
par: Lan, Wenhao, et autres
Publié: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
par: Muhamed, Aashiq, et autres
Publié: (2025)
par: Muhamed, Aashiq, et autres
Publié: (2025)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
par: DeLeeuw, Caleb
Publié: (2026)
par: DeLeeuw, Caleb
Publié: (2026)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
par: Asif, Sadia, et autres
Publié: (2026)
par: Asif, Sadia, et autres
Publié: (2026)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
par: García-Ferrero, Iker, et autres
Publié: (2025)
par: García-Ferrero, Iker, et autres
Publié: (2025)
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
par: Bae, Suyoung, et autres
Publié: (2025)
par: Bae, Suyoung, et autres
Publié: (2025)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
par: Alagharu, Rishab, et autres
Publié: (2026)
par: Alagharu, Rishab, et autres
Publié: (2026)
Programming Refusal with Conditional Activation Steering
par: Lee, Bruce W., et autres
Publié: (2024)
par: Lee, Bruce W., et autres
Publié: (2024)
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
par: Lim, Jiyoung, et autres
Publié: (2026)
par: Lim, Jiyoung, et autres
Publié: (2026)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
par: Pan, Wenbo, et autres
Publié: (2025)
par: Pan, Wenbo, et autres
Publié: (2025)
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
par: Moon, WonJun, et autres
Publié: (2023)
par: Moon, WonJun, et autres
Publié: (2023)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
par: Wu, Cheng-Kuang, et autres
Publié: (2025)
par: Wu, Cheng-Kuang, et autres
Publié: (2025)
Understanding Refusal in Language Models with Sparse Autoencoders
par: Yeo, Wei Jie, et autres
Publié: (2025)
par: Yeo, Wei Jie, et autres
Publié: (2025)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
par: Bae, Suyoung, et autres
Publié: (2026)
par: Bae, Suyoung, et autres
Publié: (2026)
Information Density Enhancement Using Lossy Compression in DNA Data Storage
par: Seongjun Seo, et autres
Publié: (2024)
par: Seongjun Seo, et autres
Publié: (2024)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
par: Dabas, Mahavir, et autres
Publié: (2025)
par: Dabas, Mahavir, et autres
Publié: (2025)
Refusal Speech
par: Nhatuve, Diocleciano, et autres
Publié: (2021)
par: Nhatuve, Diocleciano, et autres
Publié: (2021)
Refusals of noncitizenship
par: Peter Nyers
Publié: (2024)
par: Peter Nyers
Publié: (2024)
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
par: Zhu, Runchuan, et autres
Publié: (2025)
par: Zhu, Runchuan, et autres
Publié: (2025)
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
par: Kim, Su-Hyeon, et autres
Publié: (2026)
par: Kim, Su-Hyeon, et autres
Publié: (2026)
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
par: Bae, Suyoung, et autres
Publié: (2025)
par: Bae, Suyoung, et autres
Publié: (2025)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
par: Yang, Jiaxi, et autres
Publié: (2026)
par: Yang, Jiaxi, et autres
Publié: (2026)
Signs of the Great Refusal
par: Siegel, Tedd
Publié: (2023)
par: Siegel, Tedd
Publié: (2023)
Refusal Dark Matter
par: Christopher Love
Publié: (2026)
par: Christopher Love
Publié: (2026)
Lithuanian Refusals and Politeness
par: Donata Katinaitė
Publié: (2021)
par: Donata Katinaitė
Publié: (2021)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
par: Hu, Xulin, et autres
Publié: (2026)
par: Hu, Xulin, et autres
Publié: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
par: Jain, Neel, et autres
Publié: (2024)
par: Jain, Neel, et autres
Publié: (2024)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
par: Collu, Matteo Gioele, et autres
Publié: (2026)
par: Collu, Matteo Gioele, et autres
Publié: (2026)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
par: Song, Maojia, et autres
Publié: (2024)
par: Song, Maojia, et autres
Publié: (2024)
Domain-Aware Fine-Tuning: Enhancing Neural Network Adaptability
par: Ha, Seokhyeon, et autres
Publié: (2023)
par: Ha, Seokhyeon, et autres
Publié: (2023)
RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
par: Nguyen, Tuan T., et autres
Publié: (2025)
par: Nguyen, Tuan T., et autres
Publié: (2025)
Documents similaires
-
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
par: Lee, Jin-Seop, et autres
Publié: (2025) -
BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
par: Kim, DoYoung, et autres
Publié: (2025) -
Stabilizing Open-Set Test-Time Adaptation via Primary-Auxiliary Filtering and Knowledge-Integrated Prediction
par: Lee, Byung-Joon, et autres
Publié: (2025) -
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
par: Lee, Joohyeon, et autres
Publié: (2025) -
DomCLP: Domain-wise Contrastive Learning with Prototype Mixup for Unsupervised Domain Generalization
par: Lee, Jin-Seop, et autres
Publié: (2024)