Guardado en:
| Autor principal: | Parupudi, V. S. Raghu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2510.08596 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Magnitude Matters: a Superior Class of Similarity Metrics for Holistic Semantic Understanding
por: Parupudi, V. S. Raghu
Publicado: (2025)
por: Parupudi, V. S. Raghu
Publicado: (2025)
Systematic Diagnosis of Brittle Reasoning in Large Language Models
por: Parupudi, V. S. Raghu
Publicado: (2025)
por: Parupudi, V. S. Raghu
Publicado: (2025)
Before and After Temperature: A Distributional View of Creative LLM Generation
por: Parupudi, V. S. Raghu, et al.
Publicado: (2026)
por: Parupudi, V. S. Raghu, et al.
Publicado: (2026)
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
por: Cheng, Letian, et al.
Publicado: (2026)
por: Cheng, Letian, et al.
Publicado: (2026)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
por: Ankner, Zachary, et al.
Publicado: (2024)
por: Ankner, Zachary, et al.
Publicado: (2024)
Slaves to the Law of Large Numbers: An Asymptotic Equipartition Property for Perplexity in Generative Language Models
por: Bell, Tyler, et al.
Publicado: (2024)
por: Bell, Tyler, et al.
Publicado: (2024)
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
por: Mavromatis, Costas, et al.
Publicado: (2024)
por: Mavromatis, Costas, et al.
Publicado: (2024)
The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts
por: Johnson, Warren
Publicado: (2026)
por: Johnson, Warren
Publicado: (2026)
Reasoning Models Better Express Their Confidence
por: Yoon, Dongkeun, et al.
Publicado: (2025)
por: Yoon, Dongkeun, et al.
Publicado: (2025)
Lowest Span Confidence: A Zero-Shot Metric for Efficient and Black-Box Hallucination Detection in LLMs
por: Qiao, Yitong, et al.
Publicado: (2026)
por: Qiao, Yitong, et al.
Publicado: (2026)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
por: Wang, Haoyu, et al.
Publicado: (2025)
por: Wang, Haoyu, et al.
Publicado: (2025)
Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity
por: Prins, Zoë, et al.
Publicado: (2026)
por: Prins, Zoë, et al.
Publicado: (2026)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
por: Liu, Lei, et al.
Publicado: (2025)
por: Liu, Lei, et al.
Publicado: (2025)
Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?
por: Zhang, Junyan, et al.
Publicado: (2025)
por: Zhang, Junyan, et al.
Publicado: (2025)
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
por: Tian, Changyuan, et al.
Publicado: (2025)
por: Tian, Changyuan, et al.
Publicado: (2025)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
por: Seegmiller, Parker, et al.
Publicado: (2024)
por: Seegmiller, Parker, et al.
Publicado: (2024)
On Verbalized Confidence Scores for LLMs
por: Yang, Daniel, et al.
Publicado: (2024)
por: Yang, Daniel, et al.
Publicado: (2024)
Writing in Symbiosis: Mapping Human Creative Agency in the AI Era
por: Doshi, Vivan, et al.
Publicado: (2025)
por: Doshi, Vivan, et al.
Publicado: (2025)
Demystifying Prompts in Language Models via Perplexity Estimation
por: Gonen, Hila, et al.
Publicado: (2022)
por: Gonen, Hila, et al.
Publicado: (2022)
CREATE: Testing LLMs for Associative Creativity
por: Wadhwa, Manya, et al.
Publicado: (2026)
por: Wadhwa, Manya, et al.
Publicado: (2026)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
por: Wu, Siyang, et al.
Publicado: (2025)
por: Wu, Siyang, et al.
Publicado: (2025)
Improving Pretraining Data Using Perplexity Correlations
por: Thrush, Tristan, et al.
Publicado: (2024)
por: Thrush, Tristan, et al.
Publicado: (2024)
Rethinking GSPO: The Perplexity-Entropy Equivalence
por: Liu, Chi
Publicado: (2025)
por: Liu, Chi
Publicado: (2025)
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
por: Xu, Zhichao, et al.
Publicado: (2024)
por: Xu, Zhichao, et al.
Publicado: (2024)
Multicalibration for Confidence Scoring in LLMs
por: Detommaso, Gianluca, et al.
Publicado: (2024)
por: Detommaso, Gianluca, et al.
Publicado: (2024)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
Moral Mazes in the Era of LLMs
por: Nguyen, Dang, et al.
Publicado: (2026)
por: Nguyen, Dang, et al.
Publicado: (2026)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
por: Obadinma, Stephen, et al.
Publicado: (2025)
por: Obadinma, Stephen, et al.
Publicado: (2025)
Confidence Estimation for LLMs in Multi-turn Interactions
por: Zhang, Caiqi, et al.
Publicado: (2026)
por: Zhang, Caiqi, et al.
Publicado: (2026)
What is Wrong with Perplexity for Long-context Language Modeling?
por: Fang, Lizhe, et al.
Publicado: (2024)
por: Fang, Lizhe, et al.
Publicado: (2024)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
por: Xiong, Miao, et al.
Publicado: (2023)
por: Xiong, Miao, et al.
Publicado: (2023)
destroR: Attacking Transfer Models with Obfuscous Examples to Discard Perplexity
por: Ahmed, Saadat Rafid, et al.
Publicado: (2025)
por: Ahmed, Saadat Rafid, et al.
Publicado: (2025)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
por: Ge, Huaizhi, et al.
Publicado: (2024)
por: Ge, Huaizhi, et al.
Publicado: (2024)
A Thorough Examination of Decoding Methods in the Era of LLMs
por: Shi, Chufan, et al.
Publicado: (2024)
por: Shi, Chufan, et al.
Publicado: (2024)
Confidence Improves Self-Consistency in LLMs
por: Taubenfeld, Amir, et al.
Publicado: (2025)
por: Taubenfeld, Amir, et al.
Publicado: (2025)
A Perplexity and Menger Curvature-Based Approach for Similarity Evaluation of Large Language Models
por: Zhang, Yuantao, et al.
Publicado: (2025)
por: Zhang, Yuantao, et al.
Publicado: (2025)
Evaluating the Creativity of LLMs in Persian Literary Text Generation
por: Tourajmehr, Armin, et al.
Publicado: (2025)
por: Tourajmehr, Armin, et al.
Publicado: (2025)
IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs
por: Carichon, F., et al.
Publicado: (2026)
por: Carichon, F., et al.
Publicado: (2026)
Supervised Optimism Correction: Be Confident When LLMs Are Sure
por: Zhang, Junjie, et al.
Publicado: (2025)
por: Zhang, Junjie, et al.
Publicado: (2025)
Confidence Estimation in Automatic Short Answer Grading with LLMs
por: Cong, Longwei, et al.
Publicado: (2026)
por: Cong, Longwei, et al.
Publicado: (2026)
Ejemplares similares
-
Magnitude Matters: a Superior Class of Similarity Metrics for Holistic Semantic Understanding
por: Parupudi, V. S. Raghu
Publicado: (2025) -
Systematic Diagnosis of Brittle Reasoning in Large Language Models
por: Parupudi, V. S. Raghu
Publicado: (2025) -
Before and After Temperature: A Distributional View of Creative LLM Generation
por: Parupudi, V. S. Raghu, et al.
Publicado: (2026) -
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
por: Cheng, Letian, et al.
Publicado: (2026) -
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
por: Ankner, Zachary, et al.
Publicado: (2024)