Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Cheng, Letian, Wang, Junyan, Gao, Yan, Wen, Elliott, Dang, Ting, Jia, Hong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
por: Ankner, Zachary, et al.
Publicado: (2024)
por: Ankner, Zachary, et al.
Publicado: (2024)
Rethinking GSPO: The Perplexity-Entropy Equivalence
por: Liu, Chi
Publicado: (2025)
por: Liu, Chi
Publicado: (2025)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
por: Liu, Lei, et al.
Publicado: (2025)
por: Liu, Lei, et al.
Publicado: (2025)
What is Wrong with Perplexity for Long-context Language Modeling?
por: Fang, Lizhe, et al.
Publicado: (2024)
por: Fang, Lizhe, et al.
Publicado: (2024)
Improving Pretraining Data Using Perplexity Correlations
por: Thrush, Tristan, et al.
Publicado: (2024)
por: Thrush, Tristan, et al.
Publicado: (2024)
Low-Perplexity LLM-Generated Sequences and Where To Find Them
por: Wuhrmann, Arthur, et al.
Publicado: (2025)
por: Wuhrmann, Arthur, et al.
Publicado: (2025)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
por: Seegmiller, Parker, et al.
Publicado: (2024)
por: Seegmiller, Parker, et al.
Publicado: (2024)
Perplexity Cannot Always Tell Right from Wrong
por: Veličković, Petar, et al.
Publicado: (2026)
por: Veličković, Petar, et al.
Publicado: (2026)
Momentum Point-Perplexity Mechanics in Large Language Models
por: Tomaz, Lorenzo, et al.
Publicado: (2025)
por: Tomaz, Lorenzo, et al.
Publicado: (2025)
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
por: Klein, Tassilo, et al.
Publicado: (2024)
por: Klein, Tassilo, et al.
Publicado: (2024)
Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information
por: Hu, Zhengmian, et al.
Publicado: (2023)
por: Hu, Zhengmian, et al.
Publicado: (2023)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
por: Gambardella, Andrew, et al.
Publicado: (2025)
por: Gambardella, Andrew, et al.
Publicado: (2025)
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
por: Cui, Yingqian, et al.
Publicado: (2025)
por: Cui, Yingqian, et al.
Publicado: (2025)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
por: Shivagunde, Namrata, et al.
Publicado: (2026)
por: Shivagunde, Namrata, et al.
Publicado: (2026)
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
por: Dai, Lu, et al.
Publicado: (2025)
por: Dai, Lu, et al.
Publicado: (2025)
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
por: Xiao, Yao, et al.
Publicado: (2025)
por: Xiao, Yao, et al.
Publicado: (2025)
Can Perplexity Predict Fine-tuning Performance? An Investigation of Tokenization Effects on Sequential Language Models for Nepali
por: Luitel, Nishant, et al.
Publicado: (2024)
por: Luitel, Nishant, et al.
Publicado: (2024)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
por: Wang, Haoyu, et al.
Publicado: (2025)
por: Wang, Haoyu, et al.
Publicado: (2025)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
por: Boreiko, Valentyn, et al.
Publicado: (2024)
por: Boreiko, Valentyn, et al.
Publicado: (2024)
Confidence, Not Perplexity: A Better Metric for the Creative Era of LLMs
por: Parupudi, V. S. Raghu
Publicado: (2025)
por: Parupudi, V. S. Raghu
Publicado: (2025)
Hessian of Perplexity for Large Language Models by PyTorch autograd (Open Source)
por: Ilin, Ivan
Publicado: (2025)
por: Ilin, Ivan
Publicado: (2025)
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
por: Hsu, Chan-Jan, et al.
Publicado: (2026)
por: Hsu, Chan-Jan, et al.
Publicado: (2026)
Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
por: Baker, Mohammed Abu, et al.
Publicado: (2026)
por: Baker, Mohammed Abu, et al.
Publicado: (2026)
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
por: Mavromatis, Costas, et al.
Publicado: (2024)
por: Mavromatis, Costas, et al.
Publicado: (2024)
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
por: Xu, Zhichao, et al.
Publicado: (2024)
por: Xu, Zhichao, et al.
Publicado: (2024)
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
por: Tian, Changyuan, et al.
Publicado: (2025)
por: Tian, Changyuan, et al.
Publicado: (2025)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
por: Wang, Shaobo, et al.
Publicado: (2025)
por: Wang, Shaobo, et al.
Publicado: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
por: Dobler, Konstantin, et al.
Publicado: (2025)
por: Dobler, Konstantin, et al.
Publicado: (2025)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
por: Cheng, Zicong, et al.
Publicado: (2026)
por: Cheng, Zicong, et al.
Publicado: (2026)
Efficient and Personalized Mobile Health Event Prediction via Small Language Models
por: Wang, Xin, et al.
Publicado: (2024)
por: Wang, Xin, et al.
Publicado: (2024)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
por: Chen, Yen-Shan, et al.
Publicado: (2026)
por: Chen, Yen-Shan, et al.
Publicado: (2026)
Demystifying Prompts in Language Models via Perplexity Estimation
por: Gonen, Hila, et al.
Publicado: (2022)
por: Gonen, Hila, et al.
Publicado: (2022)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
por: Wu, Siyang, et al.
Publicado: (2025)
por: Wu, Siyang, et al.
Publicado: (2025)
RePPL: Recalibrating Perplexity by Uncertainty in Semantic Propagation and Language Generation for Explainable QA Hallucination Detection
por: Huang, Yiming, et al.
Publicado: (2025)
por: Huang, Yiming, et al.
Publicado: (2025)
SELF: Self-Extend the Context Length With Logistic Growth Function
por: Dang, Phat Thanh, et al.
Publicado: (2025)
por: Dang, Phat Thanh, et al.
Publicado: (2025)
The Price of Format: Diversity Collapse in LLMs
por: Yun, Longfei, et al.
Publicado: (2025)
por: Yun, Longfei, et al.
Publicado: (2025)
Intrinsic Entropy of Context Length Scaling in LLMs
por: Shi, Jingzhe, et al.
Publicado: (2025)
por: Shi, Jingzhe, et al.
Publicado: (2025)
Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples
por: Bu, Dake, et al.
Publicado: (2024)
por: Bu, Dake, et al.
Publicado: (2024)
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
por: Hua, Andong, et al.
Publicado: (2025)
por: Hua, Andong, et al.
Publicado: (2025)
Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity
por: Prins, Zoë, et al.
Publicado: (2026)
por: Prins, Zoë, et al.
Publicado: (2026)
Ejemplares similares
-
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
por: Ankner, Zachary, et al.
Publicado: (2024) -
Rethinking GSPO: The Perplexity-Entropy Equivalence
por: Liu, Chi
Publicado: (2025) -
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
por: Liu, Lei, et al.
Publicado: (2025) -
What is Wrong with Perplexity for Long-context Language Modeling?
por: Fang, Lizhe, et al.
Publicado: (2024) -
Improving Pretraining Data Using Perplexity Correlations
por: Thrush, Tristan, et al.
Publicado: (2024)