Before the Last Token: Diagnosing Final-Token Safety Probe Failures
Fuente:
arXiv
Saved in:
| Main Author: | Doda, Shravan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
by: Lelle, Travis
Published: (2026)
by: Lelle, Travis
Published: (2026)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
by: Shravan, Rohan
Published: (2026)
by: Shravan, Rohan
Published: (2026)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
by: Zanbaghi, Shahin, et al.
Published: (2025)
by: Zanbaghi, Shahin, et al.
Published: (2025)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
by: Du, Guanchen, et al.
Published: (2025)
by: Du, Guanchen, et al.
Published: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
by: Wang, Haochuan Kevin, et al.
Published: (2026)
by: Wang, Haochuan Kevin, et al.
Published: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
by: Walker, Nicholas
Published: (2024)
by: Walker, Nicholas
Published: (2024)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
by: Nathanson, Samuel, et al.
Published: (2025)
by: Nathanson, Samuel, et al.
Published: (2025)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
by: Hill, Brennen, et al.
Published: (2025)
by: Hill, Brennen, et al.
Published: (2025)
DiMEx: Breaking the Cold Start Barrier in Data-Free Model Extraction via Latent Diffusion Priors
by: Thesia, Yash, et al.
Published: (2026)
by: Thesia, Yash, et al.
Published: (2026)
ReconXF: Graph Reconstruction Attack via Public Feature Explanations on Privatized Node Features and Labels
by: Sahoo, Rishi Raj, et al.
Published: (2025)
by: Sahoo, Rishi Raj, et al.
Published: (2025)
Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
by: Adjei, Yaw Osei, et al.
Published: (2025)
by: Adjei, Yaw Osei, et al.
Published: (2025)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
by: Sun, Yuhui, et al.
Published: (2025)
by: Sun, Yuhui, et al.
Published: (2025)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
by: He, Yanjin, et al.
Published: (2025)
by: He, Yanjin, et al.
Published: (2025)
Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought
by: Zhang, Yuyi, et al.
Published: (2025)
by: Zhang, Yuyi, et al.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
by: Othman, Refat
Published: (2026)
by: Othman, Refat
Published: (2026)
Binary BPE: A Family of Cross-Platform Tokenizers for Binary Analysis
by: Bommarito II, Michael J.
Published: (2025)
by: Bommarito II, Michael J.
Published: (2025)
Continuous Latent Contexts Enable Efficient Online Learning in Transformers
by: Anand, Emile, et al.
Published: (2026)
by: Anand, Emile, et al.
Published: (2026)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
by: Gu, Yongtong, et al.
Published: (2026)
by: Gu, Yongtong, et al.
Published: (2026)
VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use
by: Santillana, Juan S.
Published: (2026)
by: Santillana, Juan S.
Published: (2026)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
by: Morales, Jaime, et al.
Published: (2026)
by: Morales, Jaime, et al.
Published: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
by: Schesch, Benedikt, et al.
Published: (2026)
by: Schesch, Benedikt, et al.
Published: (2026)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
by: Easley, Eric, et al.
Published: (2026)
by: Easley, Eric, et al.
Published: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026)
by: McCann, Jordan F.
Published: (2026)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
by: Rozonoyer, Benjamin, et al.
Published: (2026)
by: Rozonoyer, Benjamin, et al.
Published: (2026)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
by: Salfati, Samuel
Published: (2026)
by: Salfati, Samuel
Published: (2026)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
by: Atad, Ido Andrew, et al.
Published: (2026)
by: Atad, Ido Andrew, et al.
Published: (2026)
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
by: Moghadasi, Mahdi Naser, et al.
Published: (2026)
by: Moghadasi, Mahdi Naser, et al.
Published: (2026)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
by: Saxena, Udit
Published: (2025)
by: Saxena, Udit
Published: (2025)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
by: Ye, Hua, et al.
Published: (2025)
by: Ye, Hua, et al.
Published: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
by: Bakish, Yarden, et al.
Published: (2025)
by: Bakish, Yarden, et al.
Published: (2025)
QuAnTS: Question Answering on Time Series
by: Divo, Felix, et al.
Published: (2025)
by: Divo, Felix, et al.
Published: (2025)
Similar Items
-
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
by: Lelle, Travis
Published: (2026) -
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
by: Shravan, Rohan
Published: (2026) -
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
by: Zanbaghi, Shahin, et al.
Published: (2025) -
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
by: Du, Guanchen, et al.
Published: (2025) -
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)