Learning to Reason at the Frontier of Learnability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Foster, Thomas, Sims, Anya, Forkel, Johannes, Fellows, Mattie, Foerster, Jakob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SOReL and TOReL: Two Methods for Fully Offline Reinforcement Learning
von: Fellows, Mattie, et al.
Veröffentlicht: (2025)
von: Fellows, Mattie, et al.
Veröffentlicht: (2025)
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
von: Sims, Anya, et al.
Veröffentlicht: (2024)
von: Sims, Anya, et al.
Veröffentlicht: (2024)
Refining Minimax Regret for Unsupervised Environment Design
von: Beukman, Michael, et al.
Veröffentlicht: (2024)
von: Beukman, Michael, et al.
Veröffentlicht: (2024)
The Yokai Learning Environment: Tracking Beliefs Over Space and Time
von: Ruhdorfer, Constantin, et al.
Veröffentlicht: (2025)
von: Ruhdorfer, Constantin, et al.
Veröffentlicht: (2025)
DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
von: Röpke, Willem, et al.
Veröffentlicht: (2026)
von: Röpke, Willem, et al.
Veröffentlicht: (2026)
Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
von: Sabry, Mohammed, et al.
Veröffentlicht: (2025)
von: Sabry, Mohammed, et al.
Veröffentlicht: (2025)
Intent Factored Generation: Unleashing the Diversity in Your Language Model
von: Ahmed, Eltayeb, et al.
Veröffentlicht: (2025)
von: Ahmed, Eltayeb, et al.
Veröffentlicht: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
von: Malek, Alan, et al.
Veröffentlicht: (2025)
von: Malek, Alan, et al.
Veröffentlicht: (2025)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
von: Lu, Chris, et al.
Veröffentlicht: (2024)
von: Lu, Chris, et al.
Veröffentlicht: (2024)
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
von: Fan, Run-Ze, et al.
Veröffentlicht: (2025)
von: Fan, Run-Ze, et al.
Veröffentlicht: (2025)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
von: Cook, Jonathan, et al.
Veröffentlicht: (2025)
von: Cook, Jonathan, et al.
Veröffentlicht: (2025)
Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods
von: Sabry, Mohammed, et al.
Veröffentlicht: (2024)
von: Sabry, Mohammed, et al.
Veröffentlicht: (2024)
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
von: Sabry, Mohammed, et al.
Veröffentlicht: (2026)
von: Sabry, Mohammed, et al.
Veröffentlicht: (2026)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
LESA: Learnable LLM Layer Scaling-Up
von: Yang, Yifei, et al.
Veröffentlicht: (2025)
von: Yang, Yifei, et al.
Veröffentlicht: (2025)
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
von: Faisal, Faizan
Veröffentlicht: (2026)
von: Faisal, Faizan
Veröffentlicht: (2026)
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
von: Yamada, Yutaro, et al.
Veröffentlicht: (2025)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2025)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
von: Zhang, Yike, et al.
Veröffentlicht: (2025)
von: Zhang, Yike, et al.
Veröffentlicht: (2025)
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
von: Lugoloobi, William, et al.
Veröffentlicht: (2026)
von: Lugoloobi, William, et al.
Veröffentlicht: (2026)
Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning Data
von: Agarwal, Ishika, et al.
Veröffentlicht: (2025)
von: Agarwal, Ishika, et al.
Veröffentlicht: (2025)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
Self-Supervised Time-Series Anomaly Detection Using Learnable Data Augmentation
von: Choi, Kukjin, et al.
Veröffentlicht: (2024)
von: Choi, Kukjin, et al.
Veröffentlicht: (2024)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
von: Pei, Zehua, et al.
Veröffentlicht: (2024)
von: Pei, Zehua, et al.
Veröffentlicht: (2024)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
von: Fronsdal, Kai, et al.
Veröffentlicht: (2024)
von: Fronsdal, Kai, et al.
Veröffentlicht: (2024)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
von: Sakarvadia, Mansi, et al.
Veröffentlicht: (2023)
von: Sakarvadia, Mansi, et al.
Veröffentlicht: (2023)
Learning to Reason with Mixture of Tokens
von: Jain, Adit, et al.
Veröffentlicht: (2025)
von: Jain, Adit, et al.
Veröffentlicht: (2025)
Do We Need Frontier Models to Verify Mathematical Proofs?
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
Learnable Privacy Neurons Localization in Language Models
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024)
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024)
Learning to Reason for Hallucination Span Detection
von: Su, Hsuan, et al.
Veröffentlicht: (2025)
von: Su, Hsuan, et al.
Veröffentlicht: (2025)
Reasoning to Learn from Latent Thoughts
von: Ruan, Yangjun, et al.
Veröffentlicht: (2025)
von: Ruan, Yangjun, et al.
Veröffentlicht: (2025)
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
von: Yue, Murong, et al.
Veröffentlicht: (2024)
von: Yue, Murong, et al.
Veröffentlicht: (2024)
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
von: Zheng, Kunhao, et al.
Veröffentlicht: (2026)
von: Zheng, Kunhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SOReL and TOReL: Two Methods for Fully Offline Reinforcement Learning
von: Fellows, Mattie, et al.
Veröffentlicht: (2025) -
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
von: Sims, Anya, et al.
Veröffentlicht: (2024) -
Refining Minimax Regret for Unsupervised Environment Design
von: Beukman, Michael, et al.
Veröffentlicht: (2024) -
The Yokai Learning Environment: Tracking Beliefs Over Space and Time
von: Ruhdorfer, Constantin, et al.
Veröffentlicht: (2025) -
DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
von: Röpke, Willem, et al.
Veröffentlicht: (2026)