Mixtraining: A Better Trade-Off Between Compute and Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zexin, Zhang, Jiancheng, Li, Yufei, Zhu, Yinglun, Liu, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Multimodal Active Learning: Efficient Learning with Limited Paired Data
by: Zhang, Jiancheng, et al.
Published: (2025)
by: Zhang, Jiancheng, et al.
Published: (2025)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
by: Li, Yufei, et al.
Published: (2025)
by: Li, Yufei, et al.
Published: (2025)
Interactive Machine Learning: From Theory to Scale
by: Zhu, Yinglun
Published: (2025)
by: Zhu, Yinglun
Published: (2025)
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
by: Zhu, Yinglun, et al.
Published: (2025)
by: Zhu, Yinglun, et al.
Published: (2025)
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
by: Zuo, Bowen, et al.
Published: (2025)
by: Zuo, Bowen, et al.
Published: (2025)
Online Finetuning Decision Transformers with Pure RL Gradients
by: Luo, Junkai, et al.
Published: (2026)
by: Luo, Junkai, et al.
Published: (2026)
MGAS: Multi-Granularity Architecture Search for Trade-Off Between Model Effectiveness and Efficiency
by: Liu, Xiaoyun, et al.
Published: (2023)
by: Liu, Xiaoyun, et al.
Published: (2023)
Active Testing of Large Language Models via Approximate Neyman Allocation
by: Liu, Zeli, et al.
Published: (2026)
by: Liu, Zeli, et al.
Published: (2026)
Regulation of Language Models With Interpretability Will Likely Result In A Performance Trade-Off
by: Kenny, Eoin M., et al.
Published: (2024)
by: Kenny, Eoin M., et al.
Published: (2024)
Multivector Neurons: Better and Faster O(n)-Equivariant Clifford Graph Neural Networks
by: Liu, Cong, et al.
Published: (2024)
by: Liu, Cong, et al.
Published: (2024)
Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries
by: Li, Yanhang, et al.
Published: (2026)
by: Li, Yanhang, et al.
Published: (2026)
Quantifying the Accuracy-Interpretability Trade-Off in Concept-Based Sidechannel Models
by: Debot, David, et al.
Published: (2025)
by: Debot, David, et al.
Published: (2025)
Better World Models Can Lead to Better Post-Training Performance
by: Gupta, Prakhar, et al.
Published: (2025)
by: Gupta, Prakhar, et al.
Published: (2025)
Oversmoothing Alleviation in Graph Neural Networks: A Survey and Unified View
by: Jin, Yufei, et al.
Published: (2024)
by: Jin, Yufei, et al.
Published: (2024)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized Models
by: Shi, Yubin, et al.
Published: (2024)
by: Shi, Yubin, et al.
Published: (2024)
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
Compute-Optimal LLMs Provably Generalize Better With Scale
by: Finzi, Marc, et al.
Published: (2025)
by: Finzi, Marc, et al.
Published: (2025)
Latent Adversarial Regularization for Offline Preference Optimization
by: Jiang, Enyi, et al.
Published: (2026)
by: Jiang, Enyi, et al.
Published: (2026)
Cross-Space Adaptive Filter: Integrating Graph Topology and Node Attributes for Alleviating the Over-smoothing Problem
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
ODRL: A Benchmark for Off-Dynamics Reinforcement Learning
by: Lyu, Jiafei, et al.
Published: (2024)
by: Lyu, Jiafei, et al.
Published: (2024)
Distantly-Supervised Joint Extraction with Noise-Robust Learning
by: Li, Yufei, et al.
Published: (2023)
by: Li, Yufei, et al.
Published: (2023)
Better Generative Replay for Continual Federated Learning
by: Qi, Daiqing, et al.
Published: (2023)
by: Qi, Daiqing, et al.
Published: (2023)
Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
by: Li, Yushu, et al.
Published: (2026)
by: Li, Yushu, et al.
Published: (2026)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
by: Liu, Zeyuan, et al.
Published: (2026)
by: Liu, Zeyuan, et al.
Published: (2026)
Differential Privacy for Anomaly Detection: Analyzing the Trade-off Between Privacy and Explainability
by: Ezzeddine, Fatima, et al.
Published: (2024)
by: Ezzeddine, Fatima, et al.
Published: (2024)
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
by: Li, Yufei, et al.
Published: (2025)
by: Li, Yufei, et al.
Published: (2025)
Demystifying the Accuracy-Interpretability Trade-Off: A Case Study of Inferring Ratings from Reviews
by: Atrey, Pranjal, et al.
Published: (2025)
by: Atrey, Pranjal, et al.
Published: (2025)
Improving Large Models with Small models: Lower Costs and Better Performance
by: Chen, Dong, et al.
Published: (2024)
by: Chen, Dong, et al.
Published: (2024)
Byzantine-Resilient Federated Learning via Distributed Optimization
by: Xia, Yufei, et al.
Published: (2025)
by: Xia, Yufei, et al.
Published: (2025)
Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
by: Liu, Yong, et al.
Published: (2024)
by: Liu, Yong, et al.
Published: (2024)
Exploring the Trade-off Between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales
by: Resck, Lucas E., et al.
Published: (2024)
by: Resck, Lucas E., et al.
Published: (2024)
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
Emissions and Performance Trade-off Between Small and Large Language Models
by: Garg, Anandita, et al.
Published: (2025)
by: Garg, Anandita, et al.
Published: (2025)
The Path Not Taken: RLVR Provably Learns Off the Principals
by: Zhu, Hanqing, et al.
Published: (2025)
by: Zhu, Hanqing, et al.
Published: (2025)
Learning Heterogeneous Performance-Fairness Trade-offs in Federated Learning
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
by: Fan, Chongyu, et al.
Published: (2023)
by: Fan, Chongyu, et al.
Published: (2023)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
by: Amini, Afra, et al.
Published: (2025)
by: Amini, Afra, et al.
Published: (2025)
Similar Items
-
Towards Multimodal Active Learning: Efficient Learning with Limited Paired Data
by: Zhang, Jiancheng, et al.
Published: (2025) -
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
by: Li, Yufei, et al.
Published: (2025) -
Interactive Machine Learning: From Theory to Scale
by: Zhu, Yinglun
Published: (2025) -
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
by: Zhu, Yinglun, et al.
Published: (2025) -
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
by: Zuo, Bowen, et al.
Published: (2025)