Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Ankner, Zachary, Blakeney, Cody, Sreenivasan, Kartik, Marion, Max, Leavitt, Matthew L., Paul, Mansheej |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
di: Cheng, Letian, et al.
Pubblicazione: (2026)
di: Cheng, Letian, et al.
Pubblicazione: (2026)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
di: Liu, Lei, et al.
Pubblicazione: (2025)
di: Liu, Lei, et al.
Pubblicazione: (2025)
Improving Pretraining Data Using Perplexity Correlations
di: Thrush, Tristan, et al.
Pubblicazione: (2024)
di: Thrush, Tristan, et al.
Pubblicazione: (2024)
What is Wrong with Perplexity for Long-context Language Modeling?
di: Fang, Lizhe, et al.
Pubblicazione: (2024)
di: Fang, Lizhe, et al.
Pubblicazione: (2024)
Rethinking GSPO: The Perplexity-Entropy Equivalence
di: Liu, Chi
Pubblicazione: (2025)
di: Liu, Chi
Pubblicazione: (2025)
Momentum Point-Perplexity Mechanics in Large Language Models
di: Tomaz, Lorenzo, et al.
Pubblicazione: (2025)
di: Tomaz, Lorenzo, et al.
Pubblicazione: (2025)
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
di: Klein, Tassilo, et al.
Pubblicazione: (2024)
di: Klein, Tassilo, et al.
Pubblicazione: (2024)
Low-Perplexity LLM-Generated Sequences and Where To Find Them
di: Wuhrmann, Arthur, et al.
Pubblicazione: (2025)
di: Wuhrmann, Arthur, et al.
Pubblicazione: (2025)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
Perplexity Cannot Always Tell Right from Wrong
di: Veličković, Petar, et al.
Pubblicazione: (2026)
di: Veličković, Petar, et al.
Pubblicazione: (2026)
Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information
di: Hu, Zhengmian, et al.
Pubblicazione: (2023)
di: Hu, Zhengmian, et al.
Pubblicazione: (2023)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
di: Xiao, Yao, et al.
Pubblicazione: (2025)
di: Xiao, Yao, et al.
Pubblicazione: (2025)
Does your data spark joy? Performance gains from domain upsampling at the end of training
di: Blakeney, Cody, et al.
Pubblicazione: (2024)
di: Blakeney, Cody, et al.
Pubblicazione: (2024)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
di: Boreiko, Valentyn, et al.
Pubblicazione: (2024)
di: Boreiko, Valentyn, et al.
Pubblicazione: (2024)
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
di: Cui, Yingqian, et al.
Pubblicazione: (2025)
di: Cui, Yingqian, et al.
Pubblicazione: (2025)
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
di: Dai, Lu, et al.
Pubblicazione: (2025)
di: Dai, Lu, et al.
Pubblicazione: (2025)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
Can Perplexity Predict Fine-tuning Performance? An Investigation of Tokenization Effects on Sequential Language Models for Nepali
di: Luitel, Nishant, et al.
Pubblicazione: (2024)
di: Luitel, Nishant, et al.
Pubblicazione: (2024)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
di: Seegmiller, Parker, et al.
Pubblicazione: (2024)
di: Seegmiller, Parker, et al.
Pubblicazione: (2024)
Hessian of Perplexity for Large Language Models by PyTorch autograd (Open Source)
di: Ilin, Ivan
Pubblicazione: (2025)
di: Ilin, Ivan
Pubblicazione: (2025)
Demystifying Prompts in Language Models via Perplexity Estimation
di: Gonen, Hila, et al.
Pubblicazione: (2022)
di: Gonen, Hila, et al.
Pubblicazione: (2022)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
di: Wu, Siyang, et al.
Pubblicazione: (2025)
di: Wu, Siyang, et al.
Pubblicazione: (2025)
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
di: Hsu, Chan-Jan, et al.
Pubblicazione: (2026)
di: Hsu, Chan-Jan, et al.
Pubblicazione: (2026)
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
di: Mavromatis, Costas, et al.
Pubblicazione: (2024)
di: Mavromatis, Costas, et al.
Pubblicazione: (2024)
destroR: Attacking Transfer Models with Obfuscous Examples to Discard Perplexity
di: Ahmed, Saadat Rafid, et al.
Pubblicazione: (2025)
di: Ahmed, Saadat Rafid, et al.
Pubblicazione: (2025)
Scaling Laws for Precision
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
Confidence, Not Perplexity: A Better Metric for the Creative Era of LLMs
di: Parupudi, V. S. Raghu
Pubblicazione: (2025)
di: Parupudi, V. S. Raghu
Pubblicazione: (2025)
The Serial Perplex.
di: Blackwell, Maree Macon, et al.
Pubblicazione: (1978)
di: Blackwell, Maree Macon, et al.
Pubblicazione: (1978)
A Perplexity and Menger Curvature-Based Approach for Similarity Evaluation of Large Language Models
di: Zhang, Yuantao, et al.
Pubblicazione: (2025)
di: Zhang, Yuantao, et al.
Pubblicazione: (2025)
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models
di: Liu, Biao, et al.
Pubblicazione: (2024)
di: Liu, Biao, et al.
Pubblicazione: (2024)
Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
di: Baker, Mohammed Abu, et al.
Pubblicazione: (2026)
di: Baker, Mohammed Abu, et al.
Pubblicazione: (2026)
When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
di: Kostelec, Juan Gabriel, et al.
Pubblicazione: (2026)
di: Kostelec, Juan Gabriel, et al.
Pubblicazione: (2026)
Can Perplexity Reflect Large Language Model's Ability in Long Text Understanding?
di: Hu, Yutong, et al.
Pubblicazione: (2024)
di: Hu, Yutong, et al.
Pubblicazione: (2024)
Perplexed
di: Tarek Zieneldien
Pubblicazione: (2024)
di: Tarek Zieneldien
Pubblicazione: (2024)
Perplexity
di: Mahmoud, Abdalla
Pubblicazione: (2026)
di: Mahmoud, Abdalla
Pubblicazione: (2026)
Kolmogorov Complexity Bounds for LLM Steganography and a Perplexity-Based Detection Proxy
di: Shportko, Andrii
Pubblicazione: (2026)
di: Shportko, Andrii
Pubblicazione: (2026)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
Efficient Perplexity Bound and Ratio Matching in Discrete Diffusion Language Models
di: Haxholli, Etrit, et al.
Pubblicazione: (2025)
di: Haxholli, Etrit, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
di: Cheng, Letian, et al.
Pubblicazione: (2026) -
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
di: Liu, Lei, et al.
Pubblicazione: (2025) -
Improving Pretraining Data Using Perplexity Correlations
di: Thrush, Tristan, et al.
Pubblicazione: (2024) -
What is Wrong with Perplexity for Long-context Language Modeling?
di: Fang, Lizhe, et al.
Pubblicazione: (2024) -
Rethinking GSPO: The Perplexity-Entropy Equivalence
di: Liu, Chi
Pubblicazione: (2025)