Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Houyi, Zheng, Wenzhen, Wang, Qiufeng, Ding, Zhenyu, Wang, Haoying, Wang, Zili, Xuyang, Shijie, Ding, Ning, Zhou, Shuigeng, Zhang, Xiangyu, Jiang, Daxin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
by: Li, Houyi, et al.
Published: (2025)
by: Li, Houyi, et al.
Published: (2025)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
by: Li, Houyi, et al.
Published: (2025)
by: Li, Houyi, et al.
Published: (2025)
Multi-matrix Factorization Attention
by: Hu, Jingcheng, et al.
Published: (2024)
by: Hu, Jingcheng, et al.
Published: (2024)
Is Compression Really Linear with Code Intelligence?
by: Xuyang, Shijie, et al.
Published: (2025)
by: Xuyang, Shijie, et al.
Published: (2025)
Scaling Laws for State Dynamics in Large Language Models
by: Li, Jacob X, et al.
Published: (2025)
by: Li, Jacob X, et al.
Published: (2025)
W2S-AlignTree: Weak-to-Strong Inference-Time Alignment for Large Language Models via Monte Carlo Tree Search
by: Ding, Zhenyu, et al.
Published: (2025)
by: Ding, Zhenyu, et al.
Published: (2025)
Large-Scale LiDAR-Inertial Dataset for Degradation-Robust High-Precision Mapping
by: Jin, Xiaofeng, et al.
Published: (2025)
by: Jin, Xiaofeng, et al.
Published: (2025)
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
by: Sun, Luoyang, et al.
Published: (2026)
by: Sun, Luoyang, et al.
Published: (2026)
Scaling Laws for Neural Material Models
by: Trikha, Akshay, et al.
Published: (2025)
by: Trikha, Akshay, et al.
Published: (2025)
Multi-Scale Simulation of Complex Systems: A Perspective of Integrating Knowledge and Data
by: Wang, Huandong, et al.
Published: (2023)
by: Wang, Huandong, et al.
Published: (2023)
LACE: Controlled Image Prompting and Iterative Refinement with GenAI for Professional Visual Art Creators
by: Huang, Yenkai, et al.
Published: (2025)
by: Huang, Yenkai, et al.
Published: (2025)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
by: Xu, Ruoran, et al.
Published: (2026)
by: Xu, Ruoran, et al.
Published: (2026)
Scaling Laws for Forgetting When Fine-Tuning Large Language Models
by: Kalajdzievski, Damjan
Published: (2024)
by: Kalajdzievski, Damjan
Published: (2024)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
by: Nguyen, Huyen, et al.
Published: (2026)
by: Nguyen, Huyen, et al.
Published: (2026)
Scaling Laws for Code: A More Data-Hungry Regime
by: Luo, Xianzhen, et al.
Published: (2025)
by: Luo, Xianzhen, et al.
Published: (2025)
Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework
by: Chen, Jie, et al.
Published: (2025)
by: Chen, Jie, et al.
Published: (2025)
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
by: Alnemari, Mohammed, et al.
Published: (2026)
by: Alnemari, Mohammed, et al.
Published: (2026)
Adaptive Riemannian Graph Neural Networks
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
Learnable Kernel Density Estimation for Graphs and Its Application to Graph-Level Anomaly Detection
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
by: Ponnock, Jesse
Published: (2025)
by: Ponnock, Jesse
Published: (2025)
FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
by: Wang, Zixing, et al.
Published: (2025)
by: Wang, Zixing, et al.
Published: (2025)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
by: Das, Sourav
Published: (2026)
by: Das, Sourav
Published: (2026)
ContractBench: Can LLM Agents Preserve Observation Contracts?
by: Wang, Jicheng, et al.
Published: (2026)
by: Wang, Jicheng, et al.
Published: (2026)
Neurosymbolic Feature Extraction for Identifying Forced Labor in Supply Chains
by: Wang, Zili, et al.
Published: (2025)
by: Wang, Zili, et al.
Published: (2025)
RefiningGPT: Specialized language Models for Automated Refinery Unit-level Process Diagram Synthesis
by: Liu, Dongxiao, et al.
Published: (2026)
by: Liu, Dongxiao, et al.
Published: (2026)
IDDR-NGP: Incorporating Detectors for Distractor Removal with Instant Neural Radiance Field
by: Huang, Xianliang, et al.
Published: (2026)
by: Huang, Xianliang, et al.
Published: (2026)
Building Large-Scale Drone Defenses from Small-Team Strategies
by: Douglas, Grant, et al.
Published: (2026)
by: Douglas, Grant, et al.
Published: (2026)
WebWorld: A Large-Scale World Model for Web Agent Training
by: Xiao, Zikai, et al.
Published: (2026)
by: Xiao, Zikai, et al.
Published: (2026)
Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization
by: Xu, Ruoran, et al.
Published: (2026)
by: Xu, Ruoran, et al.
Published: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
by: Steele, Brady, et al.
Published: (2026)
by: Steele, Brady, et al.
Published: (2026)
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
by: Hansen-Estruch, Philippe, et al.
Published: (2025)
by: Hansen-Estruch, Philippe, et al.
Published: (2025)
Inference-Time Scaling for Visual AutoRegressive modeling by Searching Representative Samples
by: Tang, Weidong, et al.
Published: (2026)
by: Tang, Weidong, et al.
Published: (2026)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
by: Wei, Kaiwen, et al.
Published: (2025)
by: Wei, Kaiwen, et al.
Published: (2025)
RUQuant: Towards Refining Uniform Quantization for Large Language Models
by: Liu, Han, et al.
Published: (2026)
by: Liu, Han, et al.
Published: (2026)
DynamiX: Large-Scale Dynamic Social Network Simulator
by: Sun, Yanhui, et al.
Published: (2025)
by: Sun, Yanhui, et al.
Published: (2025)
FORGE: An LLM-driven Framework for Large-Scale Smart Contract Vulnerability Dataset Construction
by: Chen, Jiachi, et al.
Published: (2025)
by: Chen, Jiachi, et al.
Published: (2025)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
by: Korolkov, Vasilii
Published: (2025)
by: Korolkov, Vasilii
Published: (2025)
Similar Items
-
Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
by: Li, Houyi, et al.
Published: (2025) -
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
by: Li, Houyi, et al.
Published: (2025) -
Multi-matrix Factorization Attention
by: Hu, Jingcheng, et al.
Published: (2024) -
Is Compression Really Linear with Code Intelligence?
by: Xuyang, Shijie, et al.
Published: (2025) -
Scaling Laws for State Dynamics in Large Language Models
by: Li, Jacob X, et al.
Published: (2025)