Do Reasoning LLMs Refuse What They Infer in Long Contexts?
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Yu, Shahgir, Haz Sameen, Gong, Huanli, Wei, Zhipeng, Erichson, N. Benjamin, Dong, Yue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks
por: Shahgir, Haz Sameen, et al.
Publicado: (2023)
por: Shahgir, Haz Sameen, et al.
Publicado: (2023)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
por: Zhang, Xinkai, et al.
Publicado: (2026)
por: Zhang, Xinkai, et al.
Publicado: (2026)
Harnessing the Unseen: The Hidden Influence of Intrinsic Knowledge in Long-Context Language Models
por: Fu, Yu, et al.
Publicado: (2025)
por: Fu, Yu, et al.
Publicado: (2025)
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
por: Kim, Sangyeop, et al.
Publicado: (2025)
por: Kim, Sangyeop, et al.
Publicado: (2025)
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
por: Li, Jie, et al.
Publicado: (2024)
por: Li, Jie, et al.
Publicado: (2024)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
por: Zhang, Yuyou, et al.
Publicado: (2025)
por: Zhang, Yuyou, et al.
Publicado: (2025)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
por: Fu, Yu, et al.
Publicado: (2024)
por: Fu, Yu, et al.
Publicado: (2024)
Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy
por: Fu, Yu, et al.
Publicado: (2023)
por: Fu, Yu, et al.
Publicado: (2023)
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
por: Fu, Yu, et al.
Publicado: (2026)
por: Fu, Yu, et al.
Publicado: (2026)
Can LLM Infer Risk Information From MCP Server System Logs?
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
por: Hu, Xulin, et al.
Publicado: (2026)
por: Hu, Xulin, et al.
Publicado: (2026)
Can We Infer Confidential Properties of Training Data from LLMs?
por: Huang, Pengrun, et al.
Publicado: (2025)
por: Huang, Pengrun, et al.
Publicado: (2025)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
por: Zhao, Gejian, et al.
Publicado: (2025)
por: Zhao, Gejian, et al.
Publicado: (2025)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
por: Lan, Wenhao, et al.
Publicado: (2026)
por: Lan, Wenhao, et al.
Publicado: (2026)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
por: Kim, Jinhwa, et al.
Publicado: (2025)
por: Kim, Jinhwa, et al.
Publicado: (2025)
ContextLeak: Auditing Leakage in Private In-Context Learning Methods
por: Choi, Jacob, et al.
Publicado: (2025)
por: Choi, Jacob, et al.
Publicado: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
por: Wang, Yidan, et al.
Publicado: (2025)
por: Wang, Yidan, et al.
Publicado: (2025)
Fingerprinting LLMs via Prompt Injection
por: Hu, Yuepeng, et al.
Publicado: (2025)
por: Hu, Yuepeng, et al.
Publicado: (2025)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
por: Xie, Yueqi, et al.
Publicado: (2024)
por: Xie, Yueqi, et al.
Publicado: (2024)
LockForge: Automating Paper-to-Code for Logic Locking with Multi-Agent Reasoning LLMs
por: Saha, Akashdeep, et al.
Publicado: (2025)
por: Saha, Akashdeep, et al.
Publicado: (2025)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
por: Roh, Jaechul, et al.
Publicado: (2025)
por: Roh, Jaechul, et al.
Publicado: (2025)
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
por: Kumar, Priyanshu, et al.
Publicado: (2024)
por: Kumar, Priyanshu, et al.
Publicado: (2024)
Mitigating Jailbreaks with Intent-Aware LLMs
por: Yeo, Wei Jie, et al.
Publicado: (2025)
por: Yeo, Wei Jie, et al.
Publicado: (2025)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
por: Chae, Kyubyung, et al.
Publicado: (2025)
por: Chae, Kyubyung, et al.
Publicado: (2025)
A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models
por: Linder, Noa, et al.
Publicado: (2026)
por: Linder, Noa, et al.
Publicado: (2026)
LLMs Can Unlearn Refusal with Only 1,000 Benign Samples
por: Guo, Yangyang, et al.
Publicado: (2026)
por: Guo, Yangyang, et al.
Publicado: (2026)
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
por: Chen, Yu, et al.
Publicado: (2026)
por: Chen, Yu, et al.
Publicado: (2026)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
por: Luo, Xuan, et al.
Publicado: (2025)
por: Luo, Xuan, et al.
Publicado: (2025)
BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
por: Luo, Xuan, et al.
Publicado: (2026)
por: Luo, Xuan, et al.
Publicado: (2026)
Towards Understanding the Cognitive Habits of Large Reasoning Models
por: Dong, Jianshuo, et al.
Publicado: (2025)
por: Dong, Jianshuo, et al.
Publicado: (2025)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
por: Geng, Runpeng, et al.
Publicado: (2025)
por: Geng, Runpeng, et al.
Publicado: (2025)
Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs
por: Liu, Jinbo, et al.
Publicado: (2025)
por: Liu, Jinbo, et al.
Publicado: (2025)
A General Pseudonymization Framework for Cloud-Based LLMs: Replacing Privacy Information in Controlled Text Generation
por: Hou, Shilong, et al.
Publicado: (2025)
por: Hou, Shilong, et al.
Publicado: (2025)
What Matters For Safety Alignment?
por: Li, Xing, et al.
Publicado: (2026)
por: Li, Xing, et al.
Publicado: (2026)
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
por: Fu, Jiyuan, et al.
Publicado: (2025)
por: Fu, Jiyuan, et al.
Publicado: (2025)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
por: He, Xuanli, et al.
Publicado: (2024)
por: He, Xuanli, et al.
Publicado: (2024)
Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface
por: Labunets, Andrey, et al.
Publicado: (2025)
por: Labunets, Andrey, et al.
Publicado: (2025)
Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
por: He, Xuanli, et al.
Publicado: (2026)
por: He, Xuanli, et al.
Publicado: (2026)
MPMA: Preference Manipulation Attack Against Model Context Protocol
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
Can LLMs Infer Conversational Agent Users' Personality Traits from Chat History?
por: Cögendez, Derya, et al.
Publicado: (2026)
por: Cögendez, Derya, et al.
Publicado: (2026)
Ejemplares similares
-
Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks
por: Shahgir, Haz Sameen, et al.
Publicado: (2023) -
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
por: Zhang, Xinkai, et al.
Publicado: (2026) -
Harnessing the Unseen: The Hidden Influence of Intrinsic Knowledge in Long-Context Language Models
por: Fu, Yu, et al.
Publicado: (2025) -
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
por: Kim, Sangyeop, et al.
Publicado: (2025) -
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
por: Li, Jie, et al.
Publicado: (2024)