Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Hongliang, Li, Tung-Ling, Wu, Yuhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
by: Hejabi, Parsa, et al.
Published: (2025)
by: Hejabi, Parsa, et al.
Published: (2025)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
Aligning Multiple Knowledge Graphs in a Single Pass
by: Yang, Yaming, et al.
Published: (2024)
by: Yang, Yaming, et al.
Published: (2024)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
by: Lyu, Kaifeng, et al.
Published: (2024)
by: Lyu, Kaifeng, et al.
Published: (2024)
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
by: Li, Maximilian, et al.
Published: (2023)
by: Li, Maximilian, et al.
Published: (2023)
Prompt Perturbation Consistency Learning for Robust Language Models
by: Qiang, Yao, et al.
Published: (2024)
by: Qiang, Yao, et al.
Published: (2024)
Aligning (Medical) LLMs for (Counterfactual) Fairness
by: Poulain, Raphael, et al.
Published: (2024)
by: Poulain, Raphael, et al.
Published: (2024)
H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
by: Tekin, Selim Furkan, et al.
Published: (2024)
by: Tekin, Selim Furkan, et al.
Published: (2024)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
by: Liu, Xinyang, et al.
Published: (2023)
by: Liu, Xinyang, et al.
Published: (2023)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
by: Huang, Kaixuan, et al.
Published: (2025)
by: Huang, Kaixuan, et al.
Published: (2025)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs
by: Oliveira, Rafael C. T.
Published: (2026)
by: Oliveira, Rafael C. T.
Published: (2026)
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
by: Ni, Ruikang, et al.
Published: (2024)
by: Ni, Ruikang, et al.
Published: (2024)
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
by: Kang, Yipeng, et al.
Published: (2024)
by: Kang, Yipeng, et al.
Published: (2024)
CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-Tuning
by: Liu, Peiyuan, et al.
Published: (2024)
by: Liu, Peiyuan, et al.
Published: (2024)
Can we Soft Prompt LLMs for Graph Learning Tasks?
by: Liu, Zheyuan, et al.
Published: (2024)
by: Liu, Zheyuan, et al.
Published: (2024)
Selective Prompting Tuning for Personalized Conversations with LLMs
by: Huang, Qiushi, et al.
Published: (2024)
by: Huang, Qiushi, et al.
Published: (2024)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
by: Jiang, Huiqiang, et al.
Published: (2023)
by: Jiang, Huiqiang, et al.
Published: (2023)
Nonsense Helps: Prompt Space Perturbation Broadens Reasoning Exploration
by: Huang, Langlin, et al.
Published: (2026)
by: Huang, Langlin, et al.
Published: (2026)
Large Language Model Unlearning via Embedding-Corrupted Prompts
by: Liu, Chris Yuhao, et al.
Published: (2024)
by: Liu, Chris Yuhao, et al.
Published: (2024)
Aligning LLMs by Predicting Preferences from User Writing Samples
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
by: Aroca-Ouellette, Stéphane, et al.
Published: (2025)
ALIEN: Aligned Entropy Head for Improving Uncertainty Estimation of LLMs
by: Zabolotnyi, Artem, et al.
Published: (2025)
by: Zabolotnyi, Artem, et al.
Published: (2025)
Dialectical Behavior Therapy Approach to LLM Prompting
by: Vitman, Oxana, et al.
Published: (2024)
by: Vitman, Oxana, et al.
Published: (2024)
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying Probes
by: Maiya, Sharan, et al.
Published: (2025)
by: Maiya, Sharan, et al.
Published: (2025)
Probing the Decision Boundaries of In-context Learning in Large Language Models
by: Zhao, Siyan, et al.
Published: (2024)
by: Zhao, Siyan, et al.
Published: (2024)
Dynamic Prompt Fusion for Multi-Task and Cross-Domain Adaptation in LLMs
by: Hu, Xin, et al.
Published: (2025)
by: Hu, Xin, et al.
Published: (2025)
Aligning the Spectrum: Hybrid Graph Pre-training and Prompt Tuning across Homophily and Heterophily
by: Luo, Haitong, et al.
Published: (2025)
by: Luo, Haitong, et al.
Published: (2025)
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
by: Adeseye, Aisvarya, et al.
Published: (2026)
by: Adeseye, Aisvarya, et al.
Published: (2026)
Simplify-This: A Comparative Analysis of Prompt-Based and Fine-Tuned LLMs
by: Cohen, Eilam, et al.
Published: (2026)
by: Cohen, Eilam, et al.
Published: (2026)
Probing to Refine: Reinforcement Distillation of LLMs via Explanatory Inversion
by: Tan, Zhen, et al.
Published: (2026)
by: Tan, Zhen, et al.
Published: (2026)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
by: Liu, Mingyi
Published: (2026)
by: Liu, Mingyi
Published: (2026)
Looking for the Inner Music: Probing LLMs' Understanding of Literary Style
by: Hicke, Rebecca M. M., et al.
Published: (2025)
by: Hicke, Rebecca M. M., et al.
Published: (2025)
StreetMath: Study of LLMs' Approximation Behaviors
by: Tseng, Chiung-Yi, et al.
Published: (2025)
by: Tseng, Chiung-Yi, et al.
Published: (2025)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
Similar Items
-
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025) -
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
by: Li, Tung-Ling, et al.
Published: (2025) -
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
by: Xu, Wanghan, et al.
Published: (2025) -
Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
by: Hejabi, Parsa, et al.
Published: (2025) -
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024)