Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khattar, Vanshaj, Rashid, Md Rafi ur, Choudhury, Moumita, Liu, Jing, Koike-Akino, Toshiaki, Jin, Ming, Wang, Ye |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024)
Directional Embedding Smoothing for Robust Vision Language Models
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
Efficient Differentially Private Fine-Tuning of Diffusion Models
von: Liu, Jing, et al.
Veröffentlicht: (2024)
von: Liu, Jing, et al.
Veröffentlicht: (2024)
Smoothed Embeddings for Robust Language Models
von: Hase, Ryo, et al.
Veröffentlicht: (2025)
von: Hase, Ryo, et al.
Veröffentlicht: (2025)
Variational Randomized Smoothing for Sample-Wise Adversarial Robustness
von: Hase, Ryo, et al.
Veröffentlicht: (2024)
von: Hase, Ryo, et al.
Veröffentlicht: (2024)
Analyzing Inference Privacy Risks Through Gradients in Machine Learning
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
Why Does Differential Privacy with Large Epsilon Defend Against Practical Membership Inference Attacks?
von: Lowy, Andrew, et al.
Veröffentlicht: (2024)
von: Lowy, Andrew, et al.
Veröffentlicht: (2024)
Exploring User-level Gradient Inversion with a Diffusion Prior
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
von: Li, Zhuohang, et al.
Veröffentlicht: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
Arbiter PUF: Uniqueness and Reliability Analysis Using Hybrid CMOS-Stanford Memristor Model
von: Rahman, Tanvir, et al.
Veröffentlicht: (2025)
von: Rahman, Tanvir, et al.
Veröffentlicht: (2025)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
von: Wang, Xunguang, et al.
Veröffentlicht: (2026)
von: Wang, Xunguang, et al.
Veröffentlicht: (2026)
Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2023)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2023)
Security Vulnerabilities in Software Supply Chain for Autonomous Vehicles
von: Haque, Md Wasiul, et al.
Veröffentlicht: (2025)
von: Haque, Md Wasiul, et al.
Veröffentlicht: (2025)
Optimization Solution Functions as Deterministic Policies for Offline Reinforcement Learning
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2024)
von: Khattar, Vanshaj, et al.
Veröffentlicht: (2024)
Blockchain Amplification Attack
von: Tsuchiya, Taro, et al.
Veröffentlicht: (2024)
von: Tsuchiya, Taro, et al.
Veröffentlicht: (2024)
Privacy Amplification via Shuffling: Unified, Simplified, and Tightened
von: Wang, Shaowei, et al.
Veröffentlicht: (2023)
von: Wang, Shaowei, et al.
Veröffentlicht: (2023)
Securing Elliptic Curve Cryptocurrencies against Quantum Vulnerabilities: Resource Estimates and Mitigations
von: Babbush, Ryan, et al.
Veröffentlicht: (2026)
von: Babbush, Ryan, et al.
Veröffentlicht: (2026)
Heimdall: Formally Verified Automated Migration of Legacy eBPF Programs to Rust
von: Dasu, Vishnu Asutosh, et al.
Veröffentlicht: (2026)
von: Dasu, Vishnu Asutosh, et al.
Veröffentlicht: (2026)
GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
Validating Threat Modeling Results with the Help of Vulnerable Test Applications
von: Adamov, Oleksandr, et al.
Veröffentlicht: (2026)
von: Adamov, Oleksandr, et al.
Veröffentlicht: (2026)
Evaluating and Enhancing the Vulnerability Reasoning Capabilities of Large Language Models
von: Lu, Li, et al.
Veröffentlicht: (2026)
von: Lu, Li, et al.
Veröffentlicht: (2026)
When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models
von: Hu, Ruihan, et al.
Veröffentlicht: (2026)
von: Hu, Ruihan, et al.
Veröffentlicht: (2026)
State Machine Mutation-based Testing Framework for Wireless Communication Protocols
von: Rashid, Syed Md Mukit, et al.
Veröffentlicht: (2024)
von: Rashid, Syed Md Mukit, et al.
Veröffentlicht: (2024)
Cyber-Physical Security Vulnerabilities Identification and Classification in Smart Manufacturing -- A Defense-in-Depth Driven Framework and Taxonomy
von: Rahman, Md Habibor, et al.
Veröffentlicht: (2024)
von: Rahman, Md Habibor, et al.
Veröffentlicht: (2024)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
von: Gu, Haoran, et al.
Veröffentlicht: (2026)
von: Gu, Haoran, et al.
Veröffentlicht: (2026)
Predicting IoT Device Vulnerability Fix Times with Survival and Failure Time Models
von: A, Carlos A Rivera, et al.
Veröffentlicht: (2025)
von: A, Carlos A Rivera, et al.
Veröffentlicht: (2025)
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
von: Li, Junchen, et al.
Veröffentlicht: (2026)
von: Li, Junchen, et al.
Veröffentlicht: (2026)
randextract: a Reference Library to Test and Validate Privacy Amplification Implementations
von: Veiga, Iyán Méndez, et al.
Veröffentlicht: (2025)
von: Veiga, Iyán Méndez, et al.
Veröffentlicht: (2025)
Keystroke Detection by Exploiting Unintended RF Emission from Repaired USB Keyboards
von: Bari, Md Faizul, et al.
Veröffentlicht: (2025)
von: Bari, Md Faizul, et al.
Veröffentlicht: (2025)
Systematic Assessment of Cache Timing Vulnerabilities on RISC-V Processors
von: Austa, Cédrick, et al.
Veröffentlicht: (2025)
von: Austa, Cédrick, et al.
Veröffentlicht: (2025)
Exposing Vulnerabilities in Counterfeit Prevention Systems Utilizing Physically Unclonable Surface Features
von: Nakra, Anirudh, et al.
Veröffentlicht: (2025)
von: Nakra, Anirudh, et al.
Veröffentlicht: (2025)
Mapping Smarter, Not Harder: A Test-Time Reinforcement Learning Agent That Improves Without Labels or Model Updates
von: Tsao, Wen-Kwang, et al.
Veröffentlicht: (2025)
von: Tsao, Wen-Kwang, et al.
Veröffentlicht: (2025)
Near Exact Privacy Amplification for Matrix Mechanisms
von: Choquette-Choo, Christopher A., et al.
Veröffentlicht: (2024)
von: Choquette-Choo, Christopher A., et al.
Veröffentlicht: (2024)
Vulnerability Mitigation System (VMS): LLM Agent and Evaluation Framework for Autonomous Penetration Testing
von: Abdulzada, Farzana
Veröffentlicht: (2025)
von: Abdulzada, Farzana
Veröffentlicht: (2025)
Boosting Cybersecurity Vulnerability Scanning based on LLM-supported Static Application Security Testing
von: Keltek, Mete, et al.
Veröffentlicht: (2024)
von: Keltek, Mete, et al.
Veröffentlicht: (2024)
TAPFixer: Automatic Detection and Repair of Home Automation Vulnerabilities based on Negated-property Reasoning
von: Yu, Yinbo, et al.
Veröffentlicht: (2024)
von: Yu, Yinbo, et al.
Veröffentlicht: (2024)
Impedance vs. Power Side-channel Vulnerabilities: A Comparative Study
von: Awal, Md Sadik, et al.
Veröffentlicht: (2024)
von: Awal, Md Sadik, et al.
Veröffentlicht: (2024)
IoT-enabled Drowsiness Driver Safety Alert System with Real-Time Monitoring Using Integrated Sensors Technology
von: Muiz, Bakhtiar, et al.
Veröffentlicht: (2025)
von: Muiz, Bakhtiar, et al.
Veröffentlicht: (2025)
Breaking Precision Time: OS Vulnerability Exploits Against IEEE 1588
von: Soomro, Muhammad Abdullah, et al.
Veröffentlicht: (2025)
von: Soomro, Muhammad Abdullah, et al.
Veröffentlicht: (2025)
Impact of Differentials in SIMON32 Algorithm for Lightweight Security of Internet of Things
von: Cook, Jonathan, et al.
Veröffentlicht: (2026)
von: Cook, Jonathan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2024) -
Directional Embedding Smoothing for Robust Vision Language Models
von: Wang, Ye, et al.
Veröffentlicht: (2026) -
Efficient Differentially Private Fine-Tuning of Diffusion Models
von: Liu, Jing, et al.
Veröffentlicht: (2024) -
Smoothed Embeddings for Robust Language Models
von: Hase, Ryo, et al.
Veröffentlicht: (2025) -
Variational Randomized Smoothing for Sample-Wise Adversarial Robustness
von: Hase, Ryo, et al.
Veröffentlicht: (2024)