An Auditing Test To Detect Behavioral Shift in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Richter, Leo, He, Xuanli, Minervini, Pasquale, Kusner, Matt J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)
by: Kaddour, Jean, et al.
Published: (2026)
When Can Proxies Improve the Sample Complexity of Preference Learning?
by: Zhu, Yuchen, et al.
Published: (2024)
by: Zhu, Yuchen, et al.
Published: (2024)
Low-Rank Compression of Language Models via Differentiable Rank Selection
by: Sundrani, Sidhant, et al.
Published: (2025)
by: Sundrani, Sidhant, et al.
Published: (2025)
Setting the Record Straight on Transformer Oversmoothing
by: Dovonon, Gbètondji J-S, et al.
Published: (2024)
by: Dovonon, Gbètondji J-S, et al.
Published: (2024)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
Neurosymbolic Diffusion Models
by: van Krieken, Emile, et al.
Published: (2025)
by: van Krieken, Emile, et al.
Published: (2025)
Causal Machine Learning: A Survey and Open Problems
by: Kaddour, Jean, et al.
Published: (2022)
by: Kaddour, Jean, et al.
Published: (2022)
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
by: Ghazaryan, Gayane, et al.
Published: (2024)
by: Ghazaryan, Gayane, et al.
Published: (2024)
Temporal Smoothness Regularisers for Neural Link Predictors
by: Dileo, Manuel, et al.
Published: (2023)
by: Dileo, Manuel, et al.
Published: (2023)
Universal Properties of Activation Sparsity in Modern Large Language Models
by: Szatkowski, Filip, et al.
Published: (2025)
by: Szatkowski, Filip, et al.
Published: (2025)
Neurosymbolic Reasoning Shortcuts under the Independence Assumption
by: van Krieken, Emile, et al.
Published: (2025)
by: van Krieken, Emile, et al.
Published: (2025)
Probing the Emergence of Cross-lingual Alignment during LLM Training
by: Wang, Hetong, et al.
Published: (2024)
by: Wang, Hetong, et al.
Published: (2024)
Valid Error Bars for Neural Weather Models using Conformal Prediction
by: Gopakumar, Vignesh, et al.
Published: (2024)
by: Gopakumar, Vignesh, et al.
Published: (2024)
Empowering Domain-Specific Language Models with Graph-Oriented Databases: A Paradigm Shift in Performance and Model Maintenance
by: Di Pasquale, Ricardo, et al.
Published: (2024)
by: Di Pasquale, Ricardo, et al.
Published: (2024)
Adaptive Computation Modules: Granular Conditional Computation For Efficient Inference
by: Wójcik, Bartosz, et al.
Published: (2023)
by: Wójcik, Bartosz, et al.
Published: (2023)
Generative Models are Self-Watermarked: Declaring Model Authentication through Re-Generation
by: Desu, Aditya, et al.
Published: (2024)
by: Desu, Aditya, et al.
Published: (2024)
On the Independence Assumption in Neurosymbolic Learning
by: van Krieken, Emile, et al.
Published: (2024)
by: van Krieken, Emile, et al.
Published: (2024)
Using Natural Language Explanations to Improve Robustness of In-context Learning
by: He, Xuanli, et al.
Published: (2023)
by: He, Xuanli, et al.
Published: (2023)
Learning Physical Operators using Neural Operators
by: Gopakumar, Vignesh, et al.
Published: (2026)
by: Gopakumar, Vignesh, et al.
Published: (2026)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023)
by: Gema, Aryo Pradipta, et al.
Published: (2023)
Proxy Methods for Domain Adaptation
by: Tsai, Katherine, et al.
Published: (2024)
by: Tsai, Katherine, et al.
Published: (2024)
Belief Dynamics for Detecting Behavioral Shifts in Safe Collaborative Manipulation
by: Naik, Devashri, et al.
Published: (2026)
by: Naik, Devashri, et al.
Published: (2026)
Conditional computation in neural networks: principles and research trends
by: Scardapane, Simone, et al.
Published: (2024)
by: Scardapane, Simone, et al.
Published: (2024)
PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection
by: Camarato, Steffen J., et al.
Published: (2026)
by: Camarato, Steffen J., et al.
Published: (2026)
From Fragile to Certified: Wasserstein Audits of Group Fairness Under Distribution Shift
by: Ehyaei, Ahmad-Reza, et al.
Published: (2025)
by: Ehyaei, Ahmad-Reza, et al.
Published: (2025)
Mixtures of In-Context Learners
by: Hong, Giwon, et al.
Published: (2024)
by: Hong, Giwon, et al.
Published: (2024)
Self-Training Large Language Models for Tool-Use Without Demonstrations
by: Luo, Ne, et al.
Published: (2025)
by: Luo, Ne, et al.
Published: (2025)
Matcha: Mitigating Graph Structure Shifts with Test-Time Adaptation
by: Bao, Wenxuan, et al.
Published: (2024)
by: Bao, Wenxuan, et al.
Published: (2024)
Targeted Tests for LLM Reasoning: An Audit-Constrained Protocol
by: Li, Hongmin
Published: (2026)
by: Li, Hongmin
Published: (2026)
Probabilistic Forecasting of Radiation Exposure for Spaceflight
by: Gurav, Rutuja, et al.
Published: (2024)
by: Gurav, Rutuja, et al.
Published: (2024)
Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
by: Vaugrante, Laurène, et al.
Published: (2026)
by: Vaugrante, Laurène, et al.
Published: (2026)
Multiple Distribution Shift -- Aerial (MDS-A): A Dataset for Test-Time Error Detection and Model Adaptation
by: Ngu, Noel, et al.
Published: (2025)
by: Ngu, Noel, et al.
Published: (2025)
FLARE: Faithful Logic-Aided Reasoning and Exploration
by: Arakelyan, Erik, et al.
Published: (2024)
by: Arakelyan, Erik, et al.
Published: (2024)
Testing For Distribution Shifts with Conditional Conformal Test Martingales
by: Shaer, Shalev, et al.
Published: (2026)
by: Shaer, Shalev, et al.
Published: (2026)
Stress-Testing Alignment Audits With Prompt-Level Strategic Deception
by: Daniels, Oliver, et al.
Published: (2026)
by: Daniels, Oliver, et al.
Published: (2026)
Is Complex Query Answering Really Complex?
by: Gregucci, Cosimo, et al.
Published: (2024)
by: Gregucci, Cosimo, et al.
Published: (2024)
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Calibrated Physics-Informed Uncertainty Quantification
by: Gopakumar, Vignesh, et al.
Published: (2025)
by: Gopakumar, Vignesh, et al.
Published: (2025)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024)
by: Sui, Elaine, et al.
Published: (2024)
Auditing Language Model Unlearning via Information Decomposition
by: Goel, Anmol, et al.
Published: (2026)
by: Goel, Anmol, et al.
Published: (2026)
Similar Items
-
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026) -
When Can Proxies Improve the Sample Complexity of Preference Learning?
by: Zhu, Yuchen, et al.
Published: (2024) -
Low-Rank Compression of Language Models via Differentiable Rank Selection
by: Sundrani, Sidhant, et al.
Published: (2025) -
Setting the Record Straight on Transformer Oversmoothing
by: Dovonon, Gbètondji J-S, et al.
Published: (2024) -
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)