Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation
Fuente:
arXiv
Guardado en:
| Autores principales: | Kricheli, Joshua Shay, Reid, Alexander Lawrence, Sarkar, Soumajyoti, Gandikota, Venkata, Shakarian, Paulo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge
por: Kricheli, Joshua Shay, et al.
Publicado: (2024)
por: Kricheli, Joshua Shay, et al.
Publicado: (2024)
SNIC bifurcation and its Application to MEMS
por: Kricheli, Joshua Shay
Publicado: (2025)
por: Kricheli, Joshua Shay
Publicado: (2025)
Consistency-based Abductive Reasoning over Perceptual Errors of Multiple Pre-trained Models in Novel Environments
por: Leiva, Mario, et al.
Publicado: (2025)
por: Leiva, Mario, et al.
Publicado: (2025)
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
por: Sarkar, Soumajyoti, et al.
Publicado: (2024)
por: Sarkar, Soumajyoti, et al.
Publicado: (2024)
Diversity Measures: Domain-Independent Proxies for Failure in Language Model Queries
por: Ngu, Noel, et al.
Publicado: (2023)
por: Ngu, Noel, et al.
Publicado: (2023)
Uniform Laws of Large Numbers in Product Spaces
por: Holzman, Ron, et al.
Publicado: (2026)
por: Holzman, Ron, et al.
Publicado: (2026)
Utility Boundary of Dataset Distillation: Scaling and Configuration-Coverage Laws
por: Luo, Zhengquan, et al.
Publicado: (2025)
por: Luo, Zhengquan, et al.
Publicado: (2025)
Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
por: Lee, Dongwoo, et al.
Publicado: (2025)
por: Lee, Dongwoo, et al.
Publicado: (2025)
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
por: Wang, Haiyang, et al.
Publicado: (2024)
por: Wang, Haiyang, et al.
Publicado: (2024)
VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection
por: Taparia, Aditya, et al.
Publicado: (2025)
por: Taparia, Aditya, et al.
Publicado: (2025)
EMC$^2$: Efficient MCMC Negative Sampling for Contrastive Learning with Global Convergence
por: Yau, Chung-Yiu, et al.
Publicado: (2024)
por: Yau, Chung-Yiu, et al.
Publicado: (2024)
Enhancing Quantum Variational Algorithms with Zero Noise Extrapolation via Neural Networks
por: Bhattacharjee, Subhasree, et al.
Publicado: (2024)
por: Bhattacharjee, Subhasree, et al.
Publicado: (2024)
Towards Robust Scaling Laws for Optimizers
por: Volkova, Alexandra, et al.
Publicado: (2026)
por: Volkova, Alexandra, et al.
Publicado: (2026)
Machine Learning Model Integration with Open World Temporal Logic for Process Automation
por: Aditya, Dyuman, et al.
Publicado: (2025)
por: Aditya, Dyuman, et al.
Publicado: (2025)
Rule-Based Error Detection and Correction to Operationalize Movement Trajectory Classification
por: Xi, Bowen, et al.
Publicado: (2023)
por: Xi, Bowen, et al.
Publicado: (2023)
Robust Invariant Representation Learning by Distribution Extrapolation
por: Yoshida, Kotaro, et al.
Publicado: (2025)
por: Yoshida, Kotaro, et al.
Publicado: (2025)
Learning Perturbations to Extrapolate Your LLM
por: Cen, Zetai, et al.
Publicado: (2026)
por: Cen, Zetai, et al.
Publicado: (2026)
Sea-cret Agents: Maritime Abduction for Region Generation to Expose Dark Vessel Trajectories
por: Bavikadi, Divyagna, et al.
Publicado: (2025)
por: Bavikadi, Divyagna, et al.
Publicado: (2025)
JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation
por: Yang, Yebin, et al.
Publicado: (2026)
por: Yang, Yebin, et al.
Publicado: (2026)
PAC Guarantees for Reinforcement Learning: Sample Complexity, Coverage, and Structure
por: Steier, Joshua
Publicado: (2026)
por: Steier, Joshua
Publicado: (2026)
Robustness and Exploration of Variational and Machine Learning Approaches to Inverse Problems: An Overview
por: Auras, Alexander, et al.
Publicado: (2024)
por: Auras, Alexander, et al.
Publicado: (2024)
Abduction of Domain Relationships from Data for VQA
por: Chowdhury, Al Mehdi Saadat, et al.
Publicado: (2025)
por: Chowdhury, Al Mehdi Saadat, et al.
Publicado: (2025)
Softmax Attention with Constant Cost per Token
por: Heinsen, Franz A.
Publicado: (2024)
por: Heinsen, Franz A.
Publicado: (2024)
Targeted Variance Reduction: Robust Bayesian Optimization of Black-Box Simulators with Noise Parameters
por: Miller, John Joshua, et al.
Publicado: (2024)
por: Miller, John Joshua, et al.
Publicado: (2024)
Scaling Laws For Scalable Oversight
por: Engels, Joshua, et al.
Publicado: (2025)
por: Engels, Joshua, et al.
Publicado: (2025)
Robust Gradient Descent via Heavy-Ball Momentum with Predictive Extrapolation
por: Ali, Sarwan
Publicado: (2025)
por: Ali, Sarwan
Publicado: (2025)
Scaling Laws in Linear Regression: Compute, Parameters, and Data
por: Lin, Licong, et al.
Publicado: (2024)
por: Lin, Licong, et al.
Publicado: (2024)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
por: Lin, Zicheng, et al.
Publicado: (2024)
por: Lin, Zicheng, et al.
Publicado: (2024)
Towards Scaling Law Analysis For Spatiotemporal Weather Data
por: Kiefer, Alexander, et al.
Publicado: (2026)
por: Kiefer, Alexander, et al.
Publicado: (2026)
Large-Scale Dataset Pruning in Adversarial Training through Data Importance Extrapolation
por: Nieth, Björn, et al.
Publicado: (2024)
por: Nieth, Björn, et al.
Publicado: (2024)
Learning to Extrapolate to New Tasks: A Relational Approach to Task Extrapolation
por: Ousherovitch, Adam, et al.
Publicado: (2026)
por: Ousherovitch, Adam, et al.
Publicado: (2026)
Why Cannot Neural Networks Master Extrapolation? Insights from Physical Laws
por: Dakhmouche, Ramzi, et al.
Publicado: (2025)
por: Dakhmouche, Ramzi, et al.
Publicado: (2025)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
por: Kamigaito, Hidetaka, et al.
Publicado: (2025)
por: Kamigaito, Hidetaka, et al.
Publicado: (2025)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
por: Liu, Siyuan, et al.
Publicado: (2026)
por: Liu, Siyuan, et al.
Publicado: (2026)
Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency
por: Dwyer, Joe
Publicado: (2026)
por: Dwyer, Joe
Publicado: (2026)
Metal Price Spike Prediction via a Neurosymbolic Ensemble Approach
por: Lee, Nathaniel, et al.
Publicado: (2024)
por: Lee, Nathaniel, et al.
Publicado: (2024)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
por: Abnar, Samira, et al.
Publicado: (2025)
por: Abnar, Samira, et al.
Publicado: (2025)
Scaling Laws are Redundancy Laws
por: Bi, Yuda, et al.
Publicado: (2025)
por: Bi, Yuda, et al.
Publicado: (2025)
Generalized Coverage for More Robust Low-Budget Active Learning
por: Bae, Wonho, et al.
Publicado: (2024)
por: Bae, Wonho, et al.
Publicado: (2024)
Scale-Sensitive Shattering: Learnability and Evaluability at Optimal Scale
por: Aiyer, Shashaank, et al.
Publicado: (2026)
por: Aiyer, Shashaank, et al.
Publicado: (2026)
Ejemplares similares
-
Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge
por: Kricheli, Joshua Shay, et al.
Publicado: (2024) -
SNIC bifurcation and its Application to MEMS
por: Kricheli, Joshua Shay
Publicado: (2025) -
Consistency-based Abductive Reasoning over Perceptual Errors of Multiple Pre-trained Models in Novel Environments
por: Leiva, Mario, et al.
Publicado: (2025) -
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
por: Sarkar, Soumajyoti, et al.
Publicado: (2024) -
Diversity Measures: Domain-Independent Proxies for Failure in Language Model Queries
por: Ngu, Noel, et al.
Publicado: (2023)