Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
Fuente:
arXiv
Saved in:
| Main Authors: | Brzozowski, Michał, Dubanowska, Zuzanna, Cassano, Enrico, Chung, Neo Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning
by: Mandica, Paolo, et al.
Published: (2026)
by: Mandica, Paolo, et al.
Published: (2026)
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
by: Brzozowski, Michał, et al.
Published: (2026)
by: Brzozowski, Michał, et al.
Published: (2026)
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
by: Brzozowski, Michał, et al.
Published: (2026)
by: Brzozowski, Michał, et al.
Published: (2026)
Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
by: Dubanowska, Zuzanna, et al.
Published: (2025)
by: Dubanowska, Zuzanna, et al.
Published: (2025)
Transcoder Adapters for Reasoning-Model Diffing
by: Hu, Nathan, et al.
Published: (2026)
by: Hu, Nathan, et al.
Published: (2026)
Demystifying Verbatim Memorization in Large Language Models
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
Regularizing Attention Scores with Bootstrapping
by: Chung, Neo Christopher, et al.
Published: (2026)
by: Chung, Neo Christopher, et al.
Published: (2026)
Simple LLM Baselines are Competitive for Model Diffing
by: Kempf, Elias, et al.
Published: (2026)
by: Kempf, Elias, et al.
Published: (2026)
CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions
by: Wagner, Laurin, et al.
Published: (2024)
by: Wagner, Laurin, et al.
Published: (2024)
Class-Discriminative Attention Maps for Vision Transformers
by: Brocki, Lennart, et al.
Published: (2023)
by: Brocki, Lennart, et al.
Published: (2023)
Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection
by: Smith, Griffin Dietz, et al.
Published: (2025)
by: Smith, Griffin Dietz, et al.
Published: (2025)
Large Language Models Can Verbatim Reproduce Long Malicious Sequences
by: Lin, Sharon, et al.
Published: (2025)
by: Lin, Sharon, et al.
Published: (2025)
From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning
by: Sun, Zhanyi, et al.
Published: (2026)
by: Sun, Zhanyi, et al.
Published: (2026)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
by: Liu, Ken Ziyu, et al.
Published: (2025)
by: Liu, Ken Ziyu, et al.
Published: (2025)
EXALT: EXplainable ALgorithmic Tools for Optimization Problems
by: Bączek, Zuzanna, et al.
Published: (2025)
by: Bączek, Zuzanna, et al.
Published: (2025)
Finetuning a Weather Foundation Model with Lightweight Decoders for Unseen Physical Processes
by: Lehmann, Fanny, et al.
Published: (2025)
by: Lehmann, Fanny, et al.
Published: (2025)
FedDifRC: Unlocking the Potential of Text-to-Image Diffusion Models in Heterogeneous Federated Learning
by: Wang, Huan, et al.
Published: (2025)
by: Wang, Huan, et al.
Published: (2025)
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
by: Kassem, Aly, et al.
Published: (2026)
by: Kassem, Aly, et al.
Published: (2026)
DifCluE: Generating Counterfactual Explanations with Diffusion Autoencoders and modal clustering
by: Jain, Suparshva, et al.
Published: (2025)
by: Jain, Suparshva, et al.
Published: (2025)
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
by: Kassem, Aly M., et al.
Published: (2025)
by: Kassem, Aly M., et al.
Published: (2025)
DSAC-C: Constrained Maximum Entropy for Robust Discrete Soft-Actor Critic
by: Neo, Dexter, et al.
Published: (2023)
by: Neo, Dexter, et al.
Published: (2023)
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
by: Geng, Saibo, et al.
Published: (2023)
by: Geng, Saibo, et al.
Published: (2023)
DifFaiRec: Generative Fair Recommender with Conditional Diffusion Model
by: Jiang, Zhenhao, et al.
Published: (2024)
by: Jiang, Zhenhao, et al.
Published: (2024)
The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data
by: Baek, Christina, et al.
Published: (2026)
by: Baek, Christina, et al.
Published: (2026)
UCD: Unlearning in LLMs via Contrastive Decoding
by: Suriyakumar, Vinith M., et al.
Published: (2025)
by: Suriyakumar, Vinith M., et al.
Published: (2025)
Clustering with minimum spanning trees: How good can it be?
by: Gagolewski, Marek, et al.
Published: (2023)
by: Gagolewski, Marek, et al.
Published: (2023)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Safeguarding Generative AI Applications in Preclinical Imaging through Hybrid Anomaly Detection
by: Binda, Jakub, et al.
Published: (2025)
by: Binda, Jakub, et al.
Published: (2025)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
by: Phan, Phuc, et al.
Published: (2024)
by: Phan, Phuc, et al.
Published: (2024)
DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching
by: Aiello, Emanuele, et al.
Published: (2024)
by: Aiello, Emanuele, et al.
Published: (2024)
Logits-Based Finetuning
by: Li, Jingyao, et al.
Published: (2025)
by: Li, Jingyao, et al.
Published: (2025)
Invisible Safety Threat: Malicious Finetuning for LLM via Steganography
by: Wan, Guangnian, et al.
Published: (2026)
by: Wan, Guangnian, et al.
Published: (2026)
Exact Unlearning of Finetuning Data via Model Merging at Scale
by: Kuo, Kevin, et al.
Published: (2025)
by: Kuo, Kevin, et al.
Published: (2025)
Real-Time Energy Measurement for Non-Intrusive Well-Being Monitoring of Elderly People -- a Case Study
by: Brzozowski, Mateusz, et al.
Published: (2024)
by: Brzozowski, Mateusz, et al.
Published: (2024)
MaxEnt Loss: Constrained Maximum Entropy for Calibration under Out-of-Distribution Shift
by: Neo, Dexter, et al.
Published: (2023)
by: Neo, Dexter, et al.
Published: (2023)
Geometric Priors for Generalizable World Models via Vector Symbolic Architecture
by: Chung, William Youngwoo, et al.
Published: (2026)
by: Chung, William Youngwoo, et al.
Published: (2026)
ReFT: Representation Finetuning for Language Models
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Contrastive Representations for Temporal Reasoning
by: Ziarko, Alicja, et al.
Published: (2025)
by: Ziarko, Alicja, et al.
Published: (2025)
Similar Items
-
GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning
by: Mandica, Paolo, et al.
Published: (2026) -
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
by: Brzozowski, Michał, et al.
Published: (2026) -
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
by: Brzozowski, Michał, et al.
Published: (2026) -
Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
by: Dubanowska, Zuzanna, et al.
Published: (2025) -
Transcoder Adapters for Reasoning-Model Diffing
by: Hu, Nathan, et al.
Published: (2026)