Learning the Signature of Memorization in Autoregressive Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ilić, David, Cvejoski, Kostadin, Stanojević, David, Grigorenko, Evgeny |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Powerful Training-Free Membership Inference Against Autoregressive Language Models
by: Ilić, David, et al.
Published: (2026)
by: Ilić, David, et al.
Published: (2026)
Scalable and Verifiable Federated Learning for Cross-Institution Financial Fraud Detection
by: Panth, Prajwal, et al.
Published: (2026)
by: Panth, Prajwal, et al.
Published: (2026)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
by: Rosenblatt, Lucas, et al.
Published: (2026)
by: Rosenblatt, Lucas, et al.
Published: (2026)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
How Worrying Are Privacy Attacks Against Machine Learning?
by: Domingo-Ferrer, Josep
Published: (2025)
by: Domingo-Ferrer, Josep
Published: (2025)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026)
by: Leonesi, Matteo, et al.
Published: (2026)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
by: DeLeeuw, Caleb
Published: (2026)
by: DeLeeuw, Caleb
Published: (2026)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
by: Cohen, Liran, et al.
Published: (2025)
by: Cohen, Liran, et al.
Published: (2025)
Exponential-Family Membership Inference: From LiRA and RMIA to BaVarIA
by: Brännvall, Rickard
Published: (2026)
by: Brännvall, Rickard
Published: (2026)
Comparing Fairness of Generative Mobility Models
by: Wang, Daniel, et al.
Published: (2024)
by: Wang, Daniel, et al.
Published: (2024)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
by: Bercovich, Ivan, et al.
Published: (2026)
by: Bercovich, Ivan, et al.
Published: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
PAC-DP: Personalized Adaptive Clipping for Differentially Private Federated Learning
by: Zhou, Hao, et al.
Published: (2026)
by: Zhou, Hao, et al.
Published: (2026)
A High-Recall Cost-Sensitive Machine Learning Framework for Real-Time Online Banking Transaction Fraud Detection
by: R., Karthikeyan V., et al.
Published: (2026)
by: R., Karthikeyan V., et al.
Published: (2026)
Mitigating Disparate Impact of Differentially Private Learning through Bounded Adaptive Clipping
by: Zhao, Linzh, et al.
Published: (2025)
by: Zhao, Linzh, et al.
Published: (2025)
Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
by: Arora, Sunil, et al.
Published: (2025)
by: Arora, Sunil, et al.
Published: (2025)
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
by: Chen, Renmiao, et al.
Published: (2025)
by: Chen, Renmiao, et al.
Published: (2025)
SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do
by: Dhodapkar, Aditya, et al.
Published: (2026)
by: Dhodapkar, Aditya, et al.
Published: (2026)
Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures
by: Mahi, Ishrak Hamim, et al.
Published: (2026)
by: Mahi, Ishrak Hamim, et al.
Published: (2026)
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
Measuring the Authority Stack of AI Systems: Empirical Analysis of 366,120 Forced-Choice Responses Across 8 AI Models
by: Lee, Seulki
Published: (2026)
by: Lee, Seulki
Published: (2026)
Privacy in the Age of AI: A Taxonomy of Data Risks
by: Billiris, Grace, et al.
Published: (2025)
by: Billiris, Grace, et al.
Published: (2025)
Toward Individual Fairness Without Centralized Data: Selective Counterfactual Consistency for Vertical Federated Learning
by: Wasif, Dawood, et al.
Published: (2026)
by: Wasif, Dawood, et al.
Published: (2026)
Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
Readout-Side Bypass for Residual Hybrid Quantum-Classical Models
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
Whisper Leak: a side-channel attack on Large Language Models
by: McDonald, Geoff, et al.
Published: (2025)
by: McDonald, Geoff, et al.
Published: (2025)
HybridVFL: Disentangled Feature Learning for Edge-Enabled Vertical Federated Multimodal Classification
by: Anoosha, Mostafa, et al.
Published: (2025)
by: Anoosha, Mostafa, et al.
Published: (2025)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
by: Beltoft, Stine, et al.
Published: (2025)
by: Beltoft, Stine, et al.
Published: (2025)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening
by: Webster, Kevin T
Published: (2025)
by: Webster, Kevin T
Published: (2025)
$\mathsf{OPA}$: One-shot Private Aggregation with Single Client Interaction and its Applications to Federated Learning
by: Karthikeyan, Harish, et al.
Published: (2024)
by: Karthikeyan, Harish, et al.
Published: (2024)
Cross-Border Data Security and Privacy Risks in Large Language Models and IoT Systems
by: Handapangoda, Chalitha
Published: (2026)
by: Handapangoda, Chalitha
Published: (2026)
Membership Inference Attacks against Large Audio Language Models
by: Dong, Jia-Kai, et al.
Published: (2026)
by: Dong, Jia-Kai, et al.
Published: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
Protecting Private Code in IDE Autocomplete using Differential Privacy
by: Grigorenko, Evgeny, et al.
Published: (2026)
by: Grigorenko, Evgeny, et al.
Published: (2026)
Selecting for Less Discriminatory Algorithms: A Relational Search Framework for Navigating Fairness-Accuracy Trade-offs in Practice
by: Samad, Hana, et al.
Published: (2025)
by: Samad, Hana, et al.
Published: (2025)
The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence
by: Slattery, Peter, et al.
Published: (2024)
by: Slattery, Peter, et al.
Published: (2024)
Automated Hardware Trojan Insertion in Industrial-Scale Designs
by: Popryho, Yaroslav, et al.
Published: (2025)
by: Popryho, Yaroslav, et al.
Published: (2025)
Similar Items
-
Powerful Training-Free Membership Inference Against Autoregressive Language Models
by: Ilić, David, et al.
Published: (2026) -
Scalable and Verifiable Federated Learning for Cross-Institution Financial Fraud Detection
by: Panth, Prajwal, et al.
Published: (2026) -
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
by: Rosenblatt, Lucas, et al.
Published: (2026) -
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
by: Blanco-Justicia, Alberto, et al.
Published: (2024) -
How Worrying Are Privacy Attacks Against Machine Learning?
by: Domingo-Ferrer, Josep
Published: (2025)