Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Dwyer, Joe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
von: Xu, Minghui, et al.
Veröffentlicht: (2026)
von: Xu, Minghui, et al.
Veröffentlicht: (2026)
TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
von: Niu, Chenxu, et al.
Veröffentlicht: (2025)
von: Niu, Chenxu, et al.
Veröffentlicht: (2025)
Measuring the Energy Consumption and Efficiency of Deep Neural Networks: An Empirical Analysis and Design Recommendations
von: Tripp, Charles Edison, et al.
Veröffentlicht: (2024)
von: Tripp, Charles Edison, et al.
Veröffentlicht: (2024)
Green MLOps to Green GenOps: An Empirical Study of Energy Consumption in Discriminative and Generative AI Operations
von: Sánchez-Mompó, Adrián, et al.
Veröffentlicht: (2025)
von: Sánchez-Mompó, Adrián, et al.
Veröffentlicht: (2025)
Energy Consumption in Parallel Neural Network Training
von: Huber, Philipp, et al.
Veröffentlicht: (2025)
von: Huber, Philipp, et al.
Veröffentlicht: (2025)
An Empirical Study on the Power of Future Prediction in Partially Observable Environments
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Parameter Efficient Instruction Tuning: An Empirical Study
von: He, Pengfei
Veröffentlicht: (2024)
von: He, Pengfei
Veröffentlicht: (2024)
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts
von: Chung, Youngseog, et al.
Veröffentlicht: (2024)
von: Chung, Youngseog, et al.
Veröffentlicht: (2024)
The Empirical Impact of Neural Parameter Symmetries, or Lack Thereof
von: Lim, Derek, et al.
Veröffentlicht: (2024)
von: Lim, Derek, et al.
Veröffentlicht: (2024)
Enhancing Parameter Efficiency and Generalization in Large-Scale Models: A Regularized and Masked Low-Rank Adaptation Approach
von: Mao, Yuzhu, et al.
Veröffentlicht: (2024)
von: Mao, Yuzhu, et al.
Veröffentlicht: (2024)
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
Parameter Efficiency Is Not Memory Efficiency: Rethinking Fine-Tuning for On-Device LLM Adaptation
von: Tenison, Irene, et al.
Veröffentlicht: (2026)
von: Tenison, Irene, et al.
Veröffentlicht: (2026)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training
von: Xu, Yongzhong
Veröffentlicht: (2026)
von: Xu, Yongzhong
Veröffentlicht: (2026)
How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
von: Dionisopoulos, Lucas, et al.
Veröffentlicht: (2026)
von: Dionisopoulos, Lucas, et al.
Veröffentlicht: (2026)
Parameter-Efficient Token Embedding Editing for Clinical Class-Level Unlearning
von: Hou, Iyad Ait, et al.
Veröffentlicht: (2026)
von: Hou, Iyad Ait, et al.
Veröffentlicht: (2026)
GateRA: Token-Aware Modulation for Parameter-Efficient Fine-Tuning
von: Ou, Jie, et al.
Veröffentlicht: (2025)
von: Ou, Jie, et al.
Veröffentlicht: (2025)
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
von: Yang, Yongjin, et al.
Veröffentlicht: (2025)
von: Yang, Yongjin, et al.
Veröffentlicht: (2025)
JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation
von: Yang, Yebin, et al.
Veröffentlicht: (2026)
von: Yang, Yebin, et al.
Veröffentlicht: (2026)
Token-Level Prompt Mixture with Parameter-Free Routing for Federated Domain Generalization
von: Gong, Shuai, et al.
Veröffentlicht: (2025)
von: Gong, Shuai, et al.
Veröffentlicht: (2025)
Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity
von: Li, Bojie
Veröffentlicht: (2026)
von: Li, Bojie
Veröffentlicht: (2026)
Revisiting the Relationship between Adversarial and Clean Training: Why Clean Training Can Make Adversarial Training Better
von: Zhou, MingWei, et al.
Veröffentlicht: (2025)
von: Zhou, MingWei, et al.
Veröffentlicht: (2025)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding
von: Zhang, Zhongjian, et al.
Veröffentlicht: (2026)
von: Zhang, Zhongjian, et al.
Veröffentlicht: (2026)
No Prior, No Leakage: Revisiting Reconstruction Attacks in Trained Neural Networks
von: Refael, Yehonatan, et al.
Veröffentlicht: (2025)
von: Refael, Yehonatan, et al.
Veröffentlicht: (2025)
Revisiting LARS for Large Batch Training Generalization of Neural Networks
von: Do, Khoi, et al.
Veröffentlicht: (2023)
von: Do, Khoi, et al.
Veröffentlicht: (2023)
Partial Parameter Updates for Efficient Distributed Training
von: Filippova, Anastasiia, et al.
Veröffentlicht: (2025)
von: Filippova, Anastasiia, et al.
Veröffentlicht: (2025)
Communication Dynamics Neural Networks: FFT-Diagonalized Layers for Improved Hessian Conditioning at Reduced Parameter Count
von: Pan, Lurong
Veröffentlicht: (2026)
von: Pan, Lurong
Veröffentlicht: (2026)
Recurrent Diffusion for Large-Scale Parameter Generation
von: Wang, Kai, et al.
Veröffentlicht: (2025)
von: Wang, Kai, et al.
Veröffentlicht: (2025)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
von: Fu, Yuqian, et al.
Veröffentlicht: (2026)
von: Fu, Yuqian, et al.
Veröffentlicht: (2026)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
von: Wang, Xinglin, et al.
Veröffentlicht: (2025)
von: Wang, Xinglin, et al.
Veröffentlicht: (2025)
Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages
von: Ma, Guozheng, et al.
Veröffentlicht: (2023)
von: Ma, Guozheng, et al.
Veröffentlicht: (2023)
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
von: Du, Xianzhi, et al.
Veröffentlicht: (2024)
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
von: Hedström, Anna, et al.
Veröffentlicht: (2024)
von: Hedström, Anna, et al.
Veröffentlicht: (2024)
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
von: Liu, Tong, et al.
Veröffentlicht: (2026)
von: Liu, Tong, et al.
Veröffentlicht: (2026)
An Empirical Study of Realized GNN Expressiveness
von: Wang, Yanbo, et al.
Veröffentlicht: (2023)
von: Wang, Yanbo, et al.
Veröffentlicht: (2023)
Enhancing Training Efficiency Using Packing with Flash Attention
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
Physics-Informed Machine Learning for Vessel Shaft Power and Fuel Consumption Prediction: Interpretable KAN-based Approach
von: Mohammed, Hamza Haruna, et al.
Veröffentlicht: (2026)
von: Mohammed, Hamza Haruna, et al.
Veröffentlicht: (2026)
Revisiting Privacy, Utility, and Efficiency Trade-offs when Fine-Tuning Large Language Models
von: Das, Soumi, et al.
Veröffentlicht: (2025)
von: Das, Soumi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
von: Tan, Zelin, et al.
Veröffentlicht: (2025) -
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
von: Xu, Minghui, et al.
Veröffentlicht: (2026) -
TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
von: Niu, Chenxu, et al.
Veröffentlicht: (2025) -
Measuring the Energy Consumption and Efficiency of Deep Neural Networks: An Empirical Analysis and Design Recommendations
von: Tripp, Charles Edison, et al.
Veröffentlicht: (2024) -
Green MLOps to Green GenOps: An Empirical Study of Energy Consumption in Discriminative and Generative AI Operations
von: Sánchez-Mompó, Adrián, et al.
Veröffentlicht: (2025)