Scale Dependent Data Duplication
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kazdan, Joshua, Levi, Noam, Schaeffer, Rylan, Chudnovsky, Jessica, Puri, Abhay, He, Bo, Donmez, Mehmet, Koyejo, Sanmi, Donoho, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Prediction of Pass@k Scaling in Large Language Models
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
von: Kazdan, Joshua, et al.
Veröffentlicht: (2024)
von: Kazdan, Joshua, et al.
Veröffentlicht: (2024)
Position: Model Collapse Does Not Mean What You Think
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
von: Denisov-Blanch, Yegor, et al.
Veröffentlicht: (2026)
von: Denisov-Blanch, Yegor, et al.
Veröffentlicht: (2026)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
Pretraining Scaling Laws for Generative Evaluations of Language Models
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
How Do Large Language Monkeys Get Their Power (Laws)?
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
von: Obbad, Elyas, et al.
Veröffentlicht: (2024)
von: Obbad, Elyas, et al.
Veröffentlicht: (2024)
What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity from Incidental Causes
von: Lecomte, Victor, et al.
Veröffentlicht: (2023)
von: Lecomte, Victor, et al.
Veröffentlicht: (2023)
Investigating Data Contamination for Pre-training Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
In-Context Learning of Energy Functions
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
Quantifying Variance in Evaluation Benchmarks
von: Madaan, Lovish, et al.
Veröffentlicht: (2024)
von: Madaan, Lovish, et al.
Veröffentlicht: (2024)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
von: Belakaria, Syrine, et al.
Veröffentlicht: (2025)
von: Belakaria, Syrine, et al.
Veröffentlicht: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2026)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2026)
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
von: Levi, Noam
Veröffentlicht: (2026)
von: Levi, Noam
Veröffentlicht: (2026)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2024)
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2024)
Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
A Simple Model of Inference Scaling Laws
von: Levi, Noam
Veröffentlicht: (2024)
von: Levi, Noam
Veröffentlicht: (2024)
Best-of-N Jailbreaking
von: Hughes, John, et al.
Veröffentlicht: (2024)
von: Hughes, John, et al.
Veröffentlicht: (2024)
KGGen: Extracting Knowledge Graphs from Plain Text with Language Models
von: Mo, Belinda, et al.
Veröffentlicht: (2025)
von: Mo, Belinda, et al.
Veröffentlicht: (2025)
The Utility and Complexity of in- and out-of-Distribution Machine Unlearning
von: Allouah, Youssef, et al.
Veröffentlicht: (2024)
von: Allouah, Youssef, et al.
Veröffentlicht: (2024)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2025)
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2025)
A Framework for Objective-Driven Dynamical Stochastic Fields
von: Zhang, Yibo Jacky, et al.
Veröffentlicht: (2025)
von: Zhang, Yibo Jacky, et al.
Veröffentlicht: (2025)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
von: Zhu, Junzhe, et al.
Veröffentlicht: (2023)
von: Zhu, Junzhe, et al.
Veröffentlicht: (2023)
Reliable and Efficient Amortized Model-based Evaluation
von: Truong, Sang, et al.
Veröffentlicht: (2025)
von: Truong, Sang, et al.
Veröffentlicht: (2025)
Why Do Safety Guardrails Degrade Across Languages?
von: Zhang, Max, et al.
Veröffentlicht: (2026)
von: Zhang, Max, et al.
Veröffentlicht: (2026)
Logits are All We Need to Adapt Closed Models
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
von: Stromberg, Nathan, et al.
Veröffentlicht: (2024)
von: Stromberg, Nathan, et al.
Veröffentlicht: (2024)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
von: Chen, Edward, et al.
Veröffentlicht: (2025)
von: Chen, Edward, et al.
Veröffentlicht: (2025)
Universality of the $π^2/6$ Pathway in Avoiding Model Collapse
von: Dey, Apratim, et al.
Veröffentlicht: (2024)
von: Dey, Apratim, et al.
Veröffentlicht: (2024)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
Optimization and Generalization Guarantees for Weight Normalization
von: Cisneros-Velarde, Pedro, et al.
Veröffentlicht: (2024)
von: Cisneros-Velarde, Pedro, et al.
Veröffentlicht: (2024)
Extracting books from production language models
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026)
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026)
Decision from Suboptimal Classifiers: Excess Risk Pre- and Post-Calibration
von: Perez-Lebel, Alexandre, et al.
Veröffentlicht: (2025)
von: Perez-Lebel, Alexandre, et al.
Veröffentlicht: (2025)
Classifying Overlapping Gaussian Mixtures in High Dimensions: From Optimal Classifiers to Neural Nets
von: Cohen, Khen, et al.
Veröffentlicht: (2024)
von: Cohen, Khen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient Prediction of Pass@k Scaling in Large Language Models
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025) -
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
von: Kazdan, Joshua, et al.
Veröffentlicht: (2024) -
Position: Model Collapse Does Not Mean What You Think
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025) -
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
von: Denisov-Blanch, Yegor, et al.
Veröffentlicht: (2026) -
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
von: Gupta, Isha, et al.
Veröffentlicht: (2025)