Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baker, Mohammed Abu, Baroni, Luca, Wilhelm, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large Language Models Often Know When They Are Being Evaluated
von: Needham, Joe, et al.
Veröffentlicht: (2025)
von: Needham, Joe, et al.
Veröffentlicht: (2025)
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026)
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
von: Kostelec, Juan Gabriel, et al.
Veröffentlicht: (2026)
von: Kostelec, Juan Gabriel, et al.
Veröffentlicht: (2026)
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models
von: Liu, Biao, et al.
Veröffentlicht: (2024)
von: Liu, Biao, et al.
Veröffentlicht: (2024)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
Momentum Point-Perplexity Mechanics in Large Language Models
von: Tomaz, Lorenzo, et al.
Veröffentlicht: (2025)
von: Tomaz, Lorenzo, et al.
Veröffentlicht: (2025)
ReFT: Representation Finetuning for Language Models
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
Rethinking GSPO: The Perplexity-Entropy Equivalence
von: Liu, Chi
Veröffentlicht: (2025)
von: Liu, Chi
Veröffentlicht: (2025)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods
von: Sabry, Mohammed, et al.
Veröffentlicht: (2024)
von: Sabry, Mohammed, et al.
Veröffentlicht: (2024)
On Instruction-Finetuning Neural Machine Translation Models
von: Raunak, Vikas, et al.
Veröffentlicht: (2024)
von: Raunak, Vikas, et al.
Veröffentlicht: (2024)
Risk-Averse Finetuning of Large Language Models
von: Chaudhary, Sapana, et al.
Veröffentlicht: (2025)
von: Chaudhary, Sapana, et al.
Veröffentlicht: (2025)
MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning
von: Pavlova, Vera, et al.
Veröffentlicht: (2025)
von: Pavlova, Vera, et al.
Veröffentlicht: (2025)
Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2026)
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2026)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
von: Puerto, Haritz, et al.
Veröffentlicht: (2026)
von: Puerto, Haritz, et al.
Veröffentlicht: (2026)
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
Whisper Finetuning on Nepali Language
von: Rijal, Sanjay, et al.
Veröffentlicht: (2024)
von: Rijal, Sanjay, et al.
Veröffentlicht: (2024)
Perplexity Cannot Always Tell Right from Wrong
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
von: Diao, Shizhe, et al.
Veröffentlicht: (2023)
von: Diao, Shizhe, et al.
Veröffentlicht: (2023)
Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
Slaves to the Law of Large Numbers: An Asymptotic Equipartition Property for Perplexity in Generative Language Models
von: Bell, Tyler, et al.
Veröffentlicht: (2024)
von: Bell, Tyler, et al.
Veröffentlicht: (2024)
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
von: Cui, Yingqian, et al.
Veröffentlicht: (2025)
von: Cui, Yingqian, et al.
Veröffentlicht: (2025)
ClusComp: A Simple Paradigm for Model Compression and Efficient Finetuning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
Evaluation of Finetuned LLMs in AMR Parsing
von: Ho, Shu Han
Veröffentlicht: (2025)
von: Ho, Shu Han
Veröffentlicht: (2025)
Robust Guidance for Unsupervised Data Selection: Capturing Perplexing Named Entities for Domain-Specific Machine Translation
von: Ji, Seunghyun, et al.
Veröffentlicht: (2024)
von: Ji, Seunghyun, et al.
Veröffentlicht: (2024)
RePPL: Recalibrating Perplexity by Uncertainty in Semantic Propagation and Language Generation for Explainable QA Hallucination Detection
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
Large Language Models to Diffusion Finetuning
von: Cetin, Edoardo, et al.
Veröffentlicht: (2025)
von: Cetin, Edoardo, et al.
Veröffentlicht: (2025)
MemoryPrompt: A Light Wrapper to Improve Context Tracking in Pre-trained Language Models
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
von: Zhang, Junjie, et al.
Veröffentlicht: (2025)
von: Zhang, Junjie, et al.
Veröffentlicht: (2025)
Finetuning Generative Large Language Models with Discrimination Instructions for Knowledge Graph Completion
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Continual Learning via Sparse Memory Finetuning
von: Lin, Jessy, et al.
Veröffentlicht: (2025)
von: Lin, Jessy, et al.
Veröffentlicht: (2025)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
Vanishing Gradients in Reinforcement Finetuning of Language Models
von: Razin, Noam, et al.
Veröffentlicht: (2023)
von: Razin, Noam, et al.
Veröffentlicht: (2023)
HowkGPT: Investigating the Detection of ChatGPT-generated University Student Homework through Context-Aware Perplexity Analysis
von: Vasilatos, Christoforos, et al.
Veröffentlicht: (2023)
von: Vasilatos, Christoforos, et al.
Veröffentlicht: (2023)
Zero-Shot Classification of Crisis Tweets Using Instruction-Finetuned Large Language Models
von: McDaniel, Emma, et al.
Veröffentlicht: (2024)
von: McDaniel, Emma, et al.
Veröffentlicht: (2024)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
von: Shivagunde, Namrata, et al.
Veröffentlicht: (2026)
von: Shivagunde, Namrata, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Large Language Models Often Know When They Are Being Evaluated
von: Needham, Joe, et al.
Veröffentlicht: (2025) -
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026) -
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
von: Wang, Haoyu, et al.
Veröffentlicht: (2025) -
When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
von: Kostelec, Juan Gabriel, et al.
Veröffentlicht: (2026) -
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models
von: Liu, Biao, et al.
Veröffentlicht: (2024)