Salvato in:
| Autori principali: | Gloeckle, Fabian, Rammal, Ahmad, Arnal, Charles, Munos, Remi, Cabannes, Vivien, Synnaeve, Gabriel, Hayat, Amaury |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2604.03071 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Formalizing Mathematics at Scale
di: Rammal, Ahmad, et al.
Pubblicazione: (2026)
di: Rammal, Ahmad, et al.
Pubblicazione: (2026)
Touring sampling with pushforward maps
di: Cabannes, Vivien, et al.
Pubblicazione: (2023)
di: Cabannes, Vivien, et al.
Pubblicazione: (2023)
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
di: Peyronnet, Antoine, et al.
Pubblicazione: (2026)
di: Peyronnet, Antoine, et al.
Pubblicazione: (2026)
WybeCoder: Verified Imperative Code Generation
di: Gloeckle, Fabian, et al.
Pubblicazione: (2026)
di: Gloeckle, Fabian, et al.
Pubblicazione: (2026)
Provable Benefits of In-Tool Learning for Large Language Models
di: Houliston, Sam, et al.
Pubblicazione: (2025)
di: Houliston, Sam, et al.
Pubblicazione: (2025)
Learning with Hidden Factorial Structure
di: Arnal, Charles, et al.
Pubblicazione: (2024)
di: Arnal, Charles, et al.
Pubblicazione: (2024)
Efficient RL Training for LLMs with Experience Replay
di: Arnal, Charles, et al.
Pubblicazione: (2026)
di: Arnal, Charles, et al.
Pubblicazione: (2026)
Distilling LLM Feedback for Lean Theorem Proving
di: Narozniak, Gaetan, et al.
Pubblicazione: (2026)
di: Narozniak, Gaetan, et al.
Pubblicazione: (2026)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
di: Arnal, Charles, et al.
Pubblicazione: (2025)
di: Arnal, Charles, et al.
Pubblicazione: (2025)
Iteration Head: A Mechanistic Study of Chain-of-Thought
di: Cabannes, Vivien, et al.
Pubblicazione: (2024)
di: Cabannes, Vivien, et al.
Pubblicazione: (2024)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
di: Cohen, Taco, et al.
Pubblicazione: (2025)
di: Cohen, Taco, et al.
Pubblicazione: (2025)
The Galerkin method beats Graph-Based Approaches for Spectral Algorithms
di: Cabannes, Vivien, et al.
Pubblicazione: (2023)
di: Cabannes, Vivien, et al.
Pubblicazione: (2023)
Mode Estimation with Partial Feedback
di: Arnal, Charles, et al.
Pubblicazione: (2024)
di: Arnal, Charles, et al.
Pubblicazione: (2024)
Learning Associative Memories with Gradient Descent
di: Cabannes, Vivien, et al.
Pubblicazione: (2024)
di: Cabannes, Vivien, et al.
Pubblicazione: (2024)
Super-Exponential Regret for UCT, AlphaGo and Variants
di: Orseau, Laurent, et al.
Pubblicazione: (2024)
di: Orseau, Laurent, et al.
Pubblicazione: (2024)
Scaling Laws for Associative Memories
di: Cabannes, Vivien, et al.
Pubblicazione: (2023)
di: Cabannes, Vivien, et al.
Pubblicazione: (2023)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
Short window attention enables long-term memorization
di: Cabannes, Loïc, et al.
Pubblicazione: (2025)
di: Cabannes, Loïc, et al.
Pubblicazione: (2025)
Stochastic activations
di: Lomeli, Maria, et al.
Pubblicazione: (2025)
di: Lomeli, Maria, et al.
Pubblicazione: (2025)
Spectral bandits
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
Positional Encoding via Token-Aware Phase Attention
di: Wang, Yu, et al.
Pubblicazione: (2025)
di: Wang, Yu, et al.
Pubblicazione: (2025)
ProofOptimizer: Training Language Models to Simplify Proofs without Human Demonstrations
di: Gu, Alex, et al.
Pubblicazione: (2025)
di: Gu, Alex, et al.
Pubblicazione: (2025)
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
di: Sancaktar, Cansu, et al.
Pubblicazione: (2026)
di: Sancaktar, Cansu, et al.
Pubblicazione: (2026)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
di: Hassid, Michael, et al.
Pubblicazione: (2025)
di: Hassid, Michael, et al.
Pubblicazione: (2025)
Correlated Quantization for Faster Nonconvex Distributed Optimization
di: Panferov, Andrei, et al.
Pubblicazione: (2024)
di: Panferov, Andrei, et al.
Pubblicazione: (2024)
Temporal Difference Flows
di: Farebrother, Jesse, et al.
Pubblicazione: (2025)
di: Farebrother, Jesse, et al.
Pubblicazione: (2025)
Better & Faster Large Language Models via Multi-token Prediction
di: Gloeckle, Fabian, et al.
Pubblicazione: (2024)
di: Gloeckle, Fabian, et al.
Pubblicazione: (2024)
The KoLMogorov Test: Compression by Code Generation
di: Yoran, Ori, et al.
Pubblicazione: (2025)
di: Yoran, Ori, et al.
Pubblicazione: (2025)
Towards a Neural Debugger for Python
di: Beck, Maximilian, et al.
Pubblicazione: (2026)
di: Beck, Maximilian, et al.
Pubblicazione: (2026)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
di: Chambon, Pierre, et al.
Pubblicazione: (2025)
di: Chambon, Pierre, et al.
Pubblicazione: (2025)
Safety Alignment of LMs via Non-cooperative Games
di: Paulus, Anselm, et al.
Pubblicazione: (2025)
di: Paulus, Anselm, et al.
Pubblicazione: (2025)
Learning Mathematical Rules with Large Language Models
di: Gorceix, Antoine, et al.
Pubblicazione: (2024)
di: Gorceix, Antoine, et al.
Pubblicazione: (2024)
Prompt Selection Matters: Enhancing Text Annotations for Social Sciences with Large Language Models
di: Abraham, Louis, et al.
Pubblicazione: (2024)
di: Abraham, Louis, et al.
Pubblicazione: (2024)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
di: Rammal, Mohamad Rida, et al.
Pubblicazione: (2024)
di: Rammal, Mohamad Rida, et al.
Pubblicazione: (2024)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
di: Donhauser, Konstantin, et al.
Pubblicazione: (2025)
di: Donhauser, Konstantin, et al.
Pubblicazione: (2025)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
di: Gehring, Jonas, et al.
Pubblicazione: (2024)
di: Gehring, Jonas, et al.
Pubblicazione: (2024)
Meta Large Language Model Compiler: Foundation Models of Compiler Optimization
di: Cummins, Chris, et al.
Pubblicazione: (2024)
di: Cummins, Chris, et al.
Pubblicazione: (2024)
An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks
di: Stefan, Gabriel, et al.
Pubblicazione: (2026)
di: Stefan, Gabriel, et al.
Pubblicazione: (2026)
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
di: Zheng, Kunhao, et al.
Pubblicazione: (2026)
di: Zheng, Kunhao, et al.
Pubblicazione: (2026)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
di: Gu, Alex, et al.
Pubblicazione: (2024)
di: Gu, Alex, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Formalizing Mathematics at Scale
di: Rammal, Ahmad, et al.
Pubblicazione: (2026) -
Touring sampling with pushforward maps
di: Cabannes, Vivien, et al.
Pubblicazione: (2023) -
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
di: Peyronnet, Antoine, et al.
Pubblicazione: (2026) -
WybeCoder: Verified Imperative Code Generation
di: Gloeckle, Fabian, et al.
Pubblicazione: (2026) -
Provable Benefits of In-Tool Learning for Large Language Models
di: Houliston, Sam, et al.
Pubblicazione: (2025)