Quality over Quantity: An Effective Large-Scale Data Reduction Strategy Based on Pointwise V-Information
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Fei, Zhou, Wenchi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation
by: Xu, Jinda, et al.
Published: (2025)
by: Xu, Jinda, et al.
Published: (2025)
Navigating Data Corruption in Machine Learning: Balancing Quality, Quantity, and Imputation Strategies
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
by: Sajith, Aryan, et al.
Published: (2024)
by: Sajith, Aryan, et al.
Published: (2024)
From Overfitting to Robustness: Quantity, Quality, and Variety Oriented Negative Sample Selection in Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2024)
by: Ali, Adnan, et al.
Published: (2024)
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
by: Yang, Hanlin, et al.
Published: (2024)
by: Yang, Hanlin, et al.
Published: (2024)
Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
A Robust Clustered Federated Learning Approach for Non-IID Data with Quantity Skew
by: Ali, Michael Ben, et al.
Published: (2025)
by: Ali, Michael Ben, et al.
Published: (2025)
Scaling and Transferability of Annealing Strategies in Large Language Model Training
by: Wang, Siqi, et al.
Published: (2025)
by: Wang, Siqi, et al.
Published: (2025)
Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
by: Chen, Zhengyu, et al.
Published: (2025)
by: Chen, Zhengyu, et al.
Published: (2025)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
by: Song, Haonan, et al.
Published: (2026)
by: Song, Haonan, et al.
Published: (2026)
On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling
by: Haas, Moritz, et al.
Published: (2025)
by: Haas, Moritz, et al.
Published: (2025)
RED: Effective Trajectory Representation Learning with Comprehensive Information
by: Zhou, Silin, et al.
Published: (2024)
by: Zhou, Silin, et al.
Published: (2024)
How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws
by: Zhu, Zhitao, et al.
Published: (2026)
by: Zhu, Zhitao, et al.
Published: (2026)
Prediction Is Not Physics: Learning and Evaluating Conserved Quantities in Neural Simulators
by: Bukowski, Andrew, et al.
Published: (2026)
by: Bukowski, Andrew, et al.
Published: (2026)
Effective Exploration Based on the Structural Information Principles
by: Zeng, Xianghua, et al.
Published: (2024)
by: Zeng, Xianghua, et al.
Published: (2024)
A Survey on Data Quality Dimensions and Tools for Machine Learning
by: Zhou, Yuhan, et al.
Published: (2024)
by: Zhou, Yuhan, et al.
Published: (2024)
ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
by: Zhang, Hengrui, et al.
Published: (2025)
by: Zhang, Hengrui, et al.
Published: (2025)
Local Data Quantity-Aware Weighted Averaging for Federated Learning with Dishonest Clients
by: Wu, Leming, et al.
Published: (2025)
by: Wu, Leming, et al.
Published: (2025)
Addressing Data Quality Decompensation in Federated Learning via Dynamic Client Selection
by: Fei, Qinjun, et al.
Published: (2025)
by: Fei, Qinjun, et al.
Published: (2025)
Hierarchical Dataset Selection for High-Quality Data Sharing
by: Zhou, Xiaona, et al.
Published: (2025)
by: Zhou, Xiaona, et al.
Published: (2025)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
by: Zhou, Fan, et al.
Published: (2024)
by: Zhou, Fan, et al.
Published: (2024)
Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
by: Wang, Haohui, et al.
Published: (2025)
by: Wang, Haohui, et al.
Published: (2025)
Representation-Based Data Quality Audits for Audio
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2025)
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2025)
PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models
by: Deng, Fei, et al.
Published: (2024)
by: Deng, Fei, et al.
Published: (2024)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models
by: Shu, Yao, et al.
Published: (2024)
by: Shu, Yao, et al.
Published: (2024)
EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language Models
by: Kim, Jinhee, et al.
Published: (2024)
by: Kim, Jinhee, et al.
Published: (2024)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
by: Zhang, Zishi, et al.
Published: (2026)
by: Zhang, Zishi, et al.
Published: (2026)
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
by: Jian, Ai, et al.
Published: (2025)
by: Jian, Ai, et al.
Published: (2025)
A High-Dimensional Statistical Method for Optimizing Transfer Quantities in Multi-Source Transfer Learning
by: Zhang, Qingyue, et al.
Published: (2025)
by: Zhang, Qingyue, et al.
Published: (2025)
Unified Optimization of Source Weights and Transfer Quantities in Multi-Source Transfer Learning: An Asymptotic Framework
by: Zhang, Qingyue, et al.
Published: (2026)
by: Zhang, Qingyue, et al.
Published: (2026)
Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction
by: Lin, Yong, et al.
Published: (2025)
by: Lin, Yong, et al.
Published: (2025)
Towards Effective Planning Strategies for Dynamic Opinion Networks
by: Muppasani, Bharath, et al.
Published: (2024)
by: Muppasani, Bharath, et al.
Published: (2024)
Wukong: Towards a Scaling Law for Large-Scale Recommendation
by: Zhang, Buyun, et al.
Published: (2024)
by: Zhang, Buyun, et al.
Published: (2024)
Data Can Speak for Itself: Quality-guided Utilization of Wireless Synthetic Data
by: Gong, Chen, et al.
Published: (2025)
by: Gong, Chen, et al.
Published: (2025)
CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation
by: Chen, Yutong, et al.
Published: (2026)
by: Chen, Yutong, et al.
Published: (2026)
Similar Items
-
Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation
by: Xu, Jinda, et al.
Published: (2025) -
Navigating Data Corruption in Machine Learning: Balancing Quality, Quantity, and Imputation Strategies
by: Liu, Qi, et al.
Published: (2024) -
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
by: Sajith, Aryan, et al.
Published: (2024) -
From Overfitting to Robustness: Quantity, Quality, and Variety Oriented Negative Sample Selection in Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2024) -
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
by: Yang, Hanlin, et al.
Published: (2024)