Optimizing Data Curation through Spectral Analysis and Joint Batch Selection (SALN)
Fuente:
arXiv
Saved in:
| Main Author: | Sharifi, Mohammadreza |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
by: Sharifi, Mohammadreza, et al.
Published: (2025)
by: Sharifi, Mohammadreza, et al.
Published: (2025)
Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial Optimization
by: Lee, Deokjae, et al.
Published: (2024)
by: Lee, Deokjae, et al.
Published: (2024)
CliqueParcel: An Approach For Batching LLM Prompts That Jointly Optimizes Efficiency And Faithfulness
by: Liu, Jiayi, et al.
Published: (2024)
by: Liu, Jiayi, et al.
Published: (2024)
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
by: Seedat, Nabeel, et al.
Published: (2023)
by: Seedat, Nabeel, et al.
Published: (2023)
Joint Optimization of Resource Allocation and Data Selection for Fast and Cost-Efficient Federated Edge Learning
by: Jia, Yunjian, et al.
Published: (2024)
by: Jia, Yunjian, et al.
Published: (2024)
Efficient Training of Deep Networks using Guided Spectral Data Selection: A Step Toward Learning What You Need
by: Sharifi, Mohammadreza, et al.
Published: (2025)
by: Sharifi, Mohammadreza, et al.
Published: (2025)
Batch Acquisition Function Evaluations and Decouple Optimizer Updates for Faster Bayesian Optimization
by: Irie, Kaichi, et al.
Published: (2025)
by: Irie, Kaichi, et al.
Published: (2025)
Joint Selection: Adaptively Incorporating Public Information for Private Synthetic Data
by: Fuentes, Miguel, et al.
Published: (2024)
by: Fuentes, Miguel, et al.
Published: (2024)
Pareto Front-Diverse Batch Multi-Objective Bayesian Optimization
by: Ahmadianshalchi, Alaleh, et al.
Published: (2024)
by: Ahmadianshalchi, Alaleh, et al.
Published: (2024)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation
by: Xu, Jinda, et al.
Published: (2025)
by: Xu, Jinda, et al.
Published: (2025)
Automatic Dataset Construction (ADC): Sample Collection, Data Curation, and Beyond
by: Liu, Minghao, et al.
Published: (2024)
by: Liu, Minghao, et al.
Published: (2024)
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
by: Qin, Chongli, et al.
Published: (2025)
by: Qin, Chongli, et al.
Published: (2025)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training
by: Tyagi, Sahil, et al.
Published: (2026)
by: Tyagi, Sahil, et al.
Published: (2026)
Sequential Feature Selection for Efficient Landslide Segmentation from Multi-Spectral Data
by: Ahmad, Arsalaan, et al.
Published: (2026)
by: Ahmad, Arsalaan, et al.
Published: (2026)
Batch Bayesian Active Learning with Partial Batch Label Sampling
by: Hu, Kangping, et al.
Published: (2025)
by: Hu, Kangping, et al.
Published: (2025)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
A Hybrid Computational Intelligence Framework with Metaheuristic Optimization for Drug-Drug Interaction Prediction
by: Shamami, Maryam Abdollahi, et al.
Published: (2025)
by: Shamami, Maryam Abdollahi, et al.
Published: (2025)
Optimizing Predictive AI in Physical Design Flows with Mini Pixel Batch Gradient Descent
by: Yang, Haoyu, et al.
Published: (2024)
by: Yang, Haoyu, et al.
Published: (2024)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024)
by: Filatov, Oleg, et al.
Published: (2024)
Batch-in-Batch: a new adversarial training framework for initial perturbation and sample selection
by: Wu, Yinting, et al.
Published: (2024)
by: Wu, Yinting, et al.
Published: (2024)
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
by: Wang, Jiachen T., et al.
Published: (2025)
by: Wang, Jiachen T., et al.
Published: (2025)
BOTS: Batch Bayesian Optimization of Extended Thompson Sampling for Severely Episode-Limited RL Settings
by: Karine, Karine, et al.
Published: (2024)
by: Karine, Karine, et al.
Published: (2024)
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
by: Falahati, Ali, et al.
Published: (2026)
by: Falahati, Ali, et al.
Published: (2026)
DataRater: Meta-Learned Dataset Curation
by: Calian, Dan A., et al.
Published: (2025)
by: Calian, Dan A., et al.
Published: (2025)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
LLM Data Selection and Utilization via Dynamic Bi-level Optimization
by: Yu, Yang, et al.
Published: (2025)
by: Yu, Yang, et al.
Published: (2025)
Efficient Data Selection for Multimodal Models via Incremental Optimization Utility
by: Jing, Jinhao, et al.
Published: (2026)
by: Jing, Jinhao, et al.
Published: (2026)
BatchTopK Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2024)
by: Bussmann, Bart, et al.
Published: (2024)
Semi-supervised Batch Learning From Logged Data
by: Aminian, Gholamali, et al.
Published: (2022)
by: Aminian, Gholamali, et al.
Published: (2022)
SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Enhancing Machine Learning Model Efficiency through Quantization and Bit Depth Optimization: A Performance Analysis on Healthcare Data
by: Goswami, Mitul, et al.
Published: (2025)
by: Goswami, Mitul, et al.
Published: (2025)
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Curating Demonstrations using Online Experience
by: Chen, Annie S., et al.
Published: (2025)
by: Chen, Annie S., et al.
Published: (2025)
Batch Normalization Amplifies Memorization and Privacy Risks
by: Doan, Ngoc Phu, et al.
Published: (2026)
by: Doan, Ngoc Phu, et al.
Published: (2026)
Riemannian Batch Normalization: A Gyro Approach
by: Chen, Ziheng, et al.
Published: (2025)
by: Chen, Ziheng, et al.
Published: (2025)
Approaching Deep Learning through the Spectral Dynamics of Weights
by: Yunis, David, et al.
Published: (2024)
by: Yunis, David, et al.
Published: (2024)
Similar Items
-
Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
by: Sharifi, Mohammadreza, et al.
Published: (2025) -
Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles
by: Zhang, Yizhou, et al.
Published: (2025) -
Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial Optimization
by: Lee, Deokjae, et al.
Published: (2024) -
CliqueParcel: An Approach For Batching LLM Prompts That Jointly Optimizes Efficiency And Faithfulness
by: Liu, Jiayi, et al.
Published: (2024) -
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
by: Seedat, Nabeel, et al.
Published: (2023)