Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Jiayuan, Feldman, Vitaly, Talwar, Kunal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instance-Optimality for Private KL Distribution Estimation
by: Ye, Jiayuan, et al.
Published: (2025)
by: Ye, Jiayuan, et al.
Published: (2025)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
by: Puvvada, Krishna C., et al.
Published: (2024)
by: Puvvada, Krishna C., et al.
Published: (2024)
Less is More: Improving LLM Alignment via Preference Data Selection
by: Deng, Xun, et al.
Published: (2025)
by: Deng, Xun, et al.
Published: (2025)
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Pruning as a Defense: Reducing Memorization in Large Language Models
by: Gupta, Mansi, et al.
Published: (2025)
by: Gupta, Mansi, et al.
Published: (2025)
Exploring Cross-Client Memorization of Training Data in Large Language Models for Federated Learning
by: Udsa, Tinnakit, et al.
Published: (2025)
by: Udsa, Tinnakit, et al.
Published: (2025)
Trade-offs in Data Memorization via Strong Data Processing Inequalities
by: Feldman, Vitaly, et al.
Published: (2025)
by: Feldman, Vitaly, et al.
Published: (2025)
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
by: Kuratov, Yuri, et al.
Published: (2025)
by: Kuratov, Yuri, et al.
Published: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
Instance-Optimal Private Density Estimation in the Wasserstein Distance
by: Feldman, Vitaly, et al.
Published: (2024)
by: Feldman, Vitaly, et al.
Published: (2024)
Fact Checking Beyond Training Set
by: Karisani, Payam, et al.
Published: (2024)
by: Karisani, Payam, et al.
Published: (2024)
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Local Pan-Privacy for Federated Analytics
by: Feldman, Vitaly, et al.
Published: (2025)
by: Feldman, Vitaly, et al.
Published: (2025)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
by: Schulte, David, et al.
Published: (2024)
by: Schulte, David, et al.
Published: (2024)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
by: Zhuang, Xialie, et al.
Published: (2025)
by: Zhuang, Xialie, et al.
Published: (2025)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
by: Wang, Yihan, et al.
Published: (2022)
by: Wang, Yihan, et al.
Published: (2022)
Supernova: Achieving More with Less in Transformer Architectures
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
Less is KEN: a Universal and Simple Non-Parametric Pruning Algorithm for Large Language Models
by: Mastromattei, Michele, et al.
Published: (2024)
by: Mastromattei, Michele, et al.
Published: (2024)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
by: Lou, Chao, et al.
Published: (2024)
by: Lou, Chao, et al.
Published: (2024)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
by: Liu, Yibai, et al.
Published: (2025)
by: Liu, Yibai, et al.
Published: (2025)
Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency
by: Huang, Rapheal, et al.
Published: (2025)
by: Huang, Rapheal, et al.
Published: (2025)
Does Training on Synthetic Data Make Models Less Robust?
by: Zhang, Lingze, et al.
Published: (2025)
by: Zhang, Lingze, et al.
Published: (2025)
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
by: Wang, Zixin, et al.
Published: (2024)
by: Wang, Zixin, et al.
Published: (2024)
PREAMBLE: Private and Efficient Aggregation via Block Sparse Vectors
by: Asi, Hilal, et al.
Published: (2025)
by: Asi, Hilal, et al.
Published: (2025)
Fingerprinting Codes Meet Geometry: Improved Lower Bounds for Private Query Release and Adaptive Data Analysis
by: Lyu, Xin, et al.
Published: (2024)
by: Lyu, Xin, et al.
Published: (2024)
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
by: Liu, Jinzhe, et al.
Published: (2025)
by: Liu, Jinzhe, et al.
Published: (2025)
SLIM: Let LLM Learn More and Forget Less with Soft LoRA and Identity Mixture
by: Han, Jiayi, et al.
Published: (2024)
by: Han, Jiayi, et al.
Published: (2024)
Elephants Never Forget: Testing Language Models for Memorization of Tabular Data
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Less is More: Denoising Knowledge Graphs For Retrieval Augmented Generation
by: Zheng, Yilun, et al.
Published: (2025)
by: Zheng, Yilun, et al.
Published: (2025)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language Models
by: Barbulescu, George-Octavian, et al.
Published: (2024)
by: Barbulescu, George-Octavian, et al.
Published: (2024)
Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
by: Si, Wai Man, et al.
Published: (2026)
by: Si, Wai Man, et al.
Published: (2026)
Memorization-Compression Cycles Improve Generalization
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
ROSE: Reordered SparseGPT for More Accurate One-Shot Large Language Models Pruning
by: Su, Mingluo, et al.
Published: (2026)
by: Su, Mingluo, et al.
Published: (2026)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
by: Riaz, Haris, et al.
Published: (2025)
by: Riaz, Haris, et al.
Published: (2025)
Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals
by: Opsahl, Tobias A.
Published: (2024)
by: Opsahl, Tobias A.
Published: (2024)
Learning to Detect Language Model Training Data via Active Reconstruction
by: Yin, Junjie Oscar, et al.
Published: (2026)
by: Yin, Junjie Oscar, et al.
Published: (2026)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
by: Barron, Joshua, et al.
Published: (2025)
by: Barron, Joshua, et al.
Published: (2025)
Similar Items
-
Instance-Optimality for Private KL Distribution Estimation
by: Ye, Jiayuan, et al.
Published: (2025) -
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
by: Puvvada, Krishna C., et al.
Published: (2024) -
Less is More: Improving LLM Alignment via Preference Data Selection
by: Deng, Xun, et al.
Published: (2025) -
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024) -
Pruning as a Defense: Reducing Memorization in Large Language Models
by: Gupta, Mansi, et al.
Published: (2025)