Compute Optimal Scaling of Skills: Knowledge vs Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Roberts, Nicholas, Chatterji, Niladri, Narang, Sharan, Lewis, Mike, Hupkes, Dieuwke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
von: Ohmer, Xenia, et al.
Veröffentlicht: (2024)
von: Ohmer, Xenia, et al.
Veröffentlicht: (2024)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Quantifying Variance in Evaluation Benchmarks
von: Madaan, Lovish, et al.
Veröffentlicht: (2024)
von: Madaan, Lovish, et al.
Veröffentlicht: (2024)
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
von: Ramjee, Sharan
Veröffentlicht: (2026)
von: Ramjee, Sharan
Veröffentlicht: (2026)
Interpretability of Language Models via Task Spaces
von: Weber, Lucas, et al.
Veröffentlicht: (2024)
von: Weber, Lucas, et al.
Veröffentlicht: (2024)
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
von: Nathani, Deepak, et al.
Veröffentlicht: (2025)
von: Nathani, Deepak, et al.
Veröffentlicht: (2025)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
Cluster-norm for Unsupervised Probing of Knowledge
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
Pretrained Hybrids with MAD Skills
von: Roberts, Nicholas, et al.
Veröffentlicht: (2024)
von: Roberts, Nicholas, et al.
Veröffentlicht: (2024)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
von: Lewis, Ashley, et al.
Veröffentlicht: (2025)
von: Lewis, Ashley, et al.
Veröffentlicht: (2025)
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
von: Chen, Peter Baile, et al.
Veröffentlicht: (2025)
von: Chen, Peter Baile, et al.
Veröffentlicht: (2025)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases
von: Koretsky, Mathew J., et al.
Veröffentlicht: (2025)
von: Koretsky, Mathew J., et al.
Veröffentlicht: (2025)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2025)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
Evaluating the Robustness of Analogical Reasoning in Large Language Models
von: Lewis, Martha, et al.
Veröffentlicht: (2024)
von: Lewis, Martha, et al.
Veröffentlicht: (2024)
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
von: Singhi, Nishad, et al.
Veröffentlicht: (2025)
von: Singhi, Nishad, et al.
Veröffentlicht: (2025)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
von: Horoi, Stefan, et al.
Veröffentlicht: (2025)
von: Horoi, Stefan, et al.
Veröffentlicht: (2025)
Knowledge Fusion of Large Language Models Via Modular SkillPacks
von: Du, Guodong, et al.
Veröffentlicht: (2025)
von: Du, Guodong, et al.
Veröffentlicht: (2025)
OptScale: Probabilistic Optimality for Inference-time Scaling
von: Wang, Youkang, et al.
Veröffentlicht: (2025)
von: Wang, Youkang, et al.
Veröffentlicht: (2025)
Scaling Reasoning without Attention
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Counterfactual Reasoning with Knowledge Graph Embeddings
von: Zellinger, Lena, et al.
Veröffentlicht: (2024)
von: Zellinger, Lena, et al.
Veröffentlicht: (2024)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
von: Wang, Youkang, et al.
Veröffentlicht: (2026)
von: Wang, Youkang, et al.
Veröffentlicht: (2026)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
von: Zhou, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2026)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
von: Huang, Yuncheng, et al.
Veröffentlicht: (2023)
von: Huang, Yuncheng, et al.
Veröffentlicht: (2023)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
von: Yue, Murong, et al.
Veröffentlicht: (2024)
von: Yue, Murong, et al.
Veröffentlicht: (2024)
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
von: Maiya, Sharan, et al.
Veröffentlicht: (2025)
von: Maiya, Sharan, et al.
Veröffentlicht: (2025)
Laying the Foundation First? Investigating the Generalization from Atomic Skills to Complex Reasoning Tasks
von: Huang, Yuncheng, et al.
Veröffentlicht: (2024)
von: Huang, Yuncheng, et al.
Veröffentlicht: (2024)
Scaling Optimal LR Across Token Horizons
von: Bjorck, Johan, et al.
Veröffentlicht: (2024)
von: Bjorck, Johan, et al.
Veröffentlicht: (2024)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
von: Yang, Ling, et al.
Veröffentlicht: (2025)
von: Yang, Ling, et al.
Veröffentlicht: (2025)
On the Optimal Reasoning Length for RL-Trained Language Models
von: Nohara, Daisuke, et al.
Veröffentlicht: (2026)
von: Nohara, Daisuke, et al.
Veröffentlicht: (2026)
SAC-KG: Exploiting Large Language Models as Skilled Automatic Constructors for Domain Knowledge Graphs
von: Chen, Hanzhu, et al.
Veröffentlicht: (2024)
von: Chen, Hanzhu, et al.
Veröffentlicht: (2024)
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
von: Klila, Jaafer, et al.
Veröffentlicht: (2026)
von: Klila, Jaafer, et al.
Veröffentlicht: (2026)
Knowledge Graph Reasoning with Self-supervised Reinforcement Learning
von: Ma, Ying, et al.
Veröffentlicht: (2024)
von: Ma, Ying, et al.
Veröffentlicht: (2024)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
von: Ohmer, Xenia, et al.
Veröffentlicht: (2024) -
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025) -
Quantifying Variance in Evaluation Benchmarks
von: Madaan, Lovish, et al.
Veröffentlicht: (2024) -
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
von: Ramjee, Sharan
Veröffentlicht: (2026) -
Interpretability of Language Models via Task Spaces
von: Weber, Lucas, et al.
Veröffentlicht: (2024)