Low-Perplexity LLM-Generated Sequences and Where To Find Them
Fuente:
arXiv
Salvato in:
| Autori principali: | Wuhrmann, Arthur, Kucherenko, Anastasiia, Kucharavy, Andrei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
di: Marinas, Ines Altemir, et al.
Pubblicazione: (2025)
di: Marinas, Ines Altemir, et al.
Pubblicazione: (2025)
Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
di: Marinas, Inés Altemir, et al.
Pubblicazione: (2025)
di: Marinas, Inés Altemir, et al.
Pubblicazione: (2025)
Fantastic Bugs and Where to Find Them in AI Benchmarks
di: Truong, Sang, et al.
Pubblicazione: (2025)
di: Truong, Sang, et al.
Pubblicazione: (2025)
Fantastic Biases (What are They) and Where to Find Them
di: Barriere, Valentin
Pubblicazione: (2024)
di: Barriere, Valentin
Pubblicazione: (2024)
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
LLM Detectors Still Fall Short of Real World: Case of LLM-Generated Short News-Like Posts
di: Gameiro, Henrique Da Silva, et al.
Pubblicazione: (2024)
di: Gameiro, Henrique Da Silva, et al.
Pubblicazione: (2024)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
di: Ankner, Zachary, et al.
Pubblicazione: (2024)
di: Ankner, Zachary, et al.
Pubblicazione: (2024)
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
di: Cheng, Letian, et al.
Pubblicazione: (2026)
di: Cheng, Letian, et al.
Pubblicazione: (2026)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
di: Seegmiller, Parker, et al.
Pubblicazione: (2024)
di: Seegmiller, Parker, et al.
Pubblicazione: (2024)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
di: Liu, Lei, et al.
Pubblicazione: (2025)
di: Liu, Lei, et al.
Pubblicazione: (2025)
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
di: Klein, Tassilo, et al.
Pubblicazione: (2024)
di: Klein, Tassilo, et al.
Pubblicazione: (2024)
Improving Pretraining Data Using Perplexity Correlations
di: Thrush, Tristan, et al.
Pubblicazione: (2024)
di: Thrush, Tristan, et al.
Pubblicazione: (2024)
What is Wrong with Perplexity for Long-context Language Modeling?
di: Fang, Lizhe, et al.
Pubblicazione: (2024)
di: Fang, Lizhe, et al.
Pubblicazione: (2024)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
Low Rank Gradients and Where to Find Them
di: Sonthalia, Rishi, et al.
Pubblicazione: (2025)
di: Sonthalia, Rishi, et al.
Pubblicazione: (2025)
From Model to Breach: Towards Actionable LLM-Generated Vulnerabilities Reporting
di: Vallez, Cyril, et al.
Pubblicazione: (2025)
di: Vallez, Cyril, et al.
Pubblicazione: (2025)
Rethinking GSPO: The Perplexity-Entropy Equivalence
di: Liu, Chi
Pubblicazione: (2025)
di: Liu, Chi
Pubblicazione: (2025)
Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Them
di: Brinner, Marc, et al.
Pubblicazione: (2025)
di: Brinner, Marc, et al.
Pubblicazione: (2025)
Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information
di: Hu, Zhengmian, et al.
Pubblicazione: (2023)
di: Hu, Zhengmian, et al.
Pubblicazione: (2023)
TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code
di: Sternfeld, Alexander, et al.
Pubblicazione: (2025)
di: Sternfeld, Alexander, et al.
Pubblicazione: (2025)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
di: Ravichander, Abhilasha, et al.
Pubblicazione: (2025)
di: Ravichander, Abhilasha, et al.
Pubblicazione: (2025)
Momentum Point-Perplexity Mechanics in Large Language Models
di: Tomaz, Lorenzo, et al.
Pubblicazione: (2025)
di: Tomaz, Lorenzo, et al.
Pubblicazione: (2025)
Perplexity Cannot Always Tell Right from Wrong
di: Veličković, Petar, et al.
Pubblicazione: (2026)
di: Veličković, Petar, et al.
Pubblicazione: (2026)
DiLoCo: Distributed Low-Communication Training of Language Models
di: Douillard, Arthur, et al.
Pubblicazione: (2023)
di: Douillard, Arthur, et al.
Pubblicazione: (2023)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
Can Perplexity Predict Fine-tuning Performance? An Investigation of Tokenization Effects on Sequential Language Models for Nepali
di: Luitel, Nishant, et al.
Pubblicazione: (2024)
di: Luitel, Nishant, et al.
Pubblicazione: (2024)
Transcoders Find Interpretable LLM Feature Circuits
di: Dunefsky, Jacob, et al.
Pubblicazione: (2024)
di: Dunefsky, Jacob, et al.
Pubblicazione: (2024)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
di: Kwon, Omin, et al.
Pubblicazione: (2026)
di: Kwon, Omin, et al.
Pubblicazione: (2026)
DiscoverLLM: From Executing Intents to Discovering Them
di: Kim, Tae Soo, et al.
Pubblicazione: (2026)
di: Kim, Tae Soo, et al.
Pubblicazione: (2026)
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
di: Dai, Lu, et al.
Pubblicazione: (2025)
di: Dai, Lu, et al.
Pubblicazione: (2025)
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
di: Xiao, Yao, et al.
Pubblicazione: (2025)
di: Xiao, Yao, et al.
Pubblicazione: (2025)
A Lightweight Method to Disrupt Memorized Sequences in LLM
di: Prashant, Parjanya Prajakta, et al.
Pubblicazione: (2025)
di: Prashant, Parjanya Prajakta, et al.
Pubblicazione: (2025)
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
di: Kraus, Oliver, et al.
Pubblicazione: (2026)
di: Kraus, Oliver, et al.
Pubblicazione: (2026)
ESQA: Event Sequences Question Answering
di: Abdullaeva, Irina, et al.
Pubblicazione: (2024)
di: Abdullaeva, Irina, et al.
Pubblicazione: (2024)
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
di: Cui, Yingqian, et al.
Pubblicazione: (2025)
di: Cui, Yingqian, et al.
Pubblicazione: (2025)
Fantastic Pretraining Optimizers and Where to Find Them
di: Wen, Kaiyue, et al.
Pubblicazione: (2025)
di: Wen, Kaiyue, et al.
Pubblicazione: (2025)
Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them
di: Bui, Anh, et al.
Pubblicazione: (2025)
di: Bui, Anh, et al.
Pubblicazione: (2025)
Where is the signal in tokenization space?
di: Geh, Renato Lui, et al.
Pubblicazione: (2024)
di: Geh, Renato Lui, et al.
Pubblicazione: (2024)
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
di: Pouransari, Hadi, et al.
Pubblicazione: (2024)
di: Pouransari, Hadi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
di: Marinas, Ines Altemir, et al.
Pubblicazione: (2025) -
Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
di: Marinas, Inés Altemir, et al.
Pubblicazione: (2025) -
Fantastic Bugs and Where to Find Them in AI Benchmarks
di: Truong, Sang, et al.
Pubblicazione: (2025) -
Fantastic Biases (What are They) and Where to Find Them
di: Barriere, Valentin
Pubblicazione: (2024) -
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)