h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Motwani, Sumeet Ramesh, Ivanova, Alesia, Cai, Ziyang, Torr, Philip, Islam, Riashat, Shah, Shital, de Witt, Christian Schroeder, London, Charles |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
MALT: Improving Reasoning with Multi-Agent LLM Training
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification
by: Wang, Erik Y., et al.
Published: (2026)
by: Wang, Erik Y., et al.
Published: (2026)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
Unelicitable Backdoors in Language Models via Cryptographic Transformer Circuits
by: Draguns, Andis, et al.
Published: (2024)
by: Draguns, Andis, et al.
Published: (2024)
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
by: Venhoff, Constantin, et al.
Published: (2024)
by: Venhoff, Constantin, et al.
Published: (2024)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
by: Aljaafari, Tala, et al.
Published: (2025)
by: Aljaafari, Tala, et al.
Published: (2025)
Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap
by: Channing, Georgia, et al.
Published: (2024)
by: Channing, Georgia, et al.
Published: (2024)
STARC: A General Framework For Quantifying Differences Between Reward Functions
by: Skalse, Joar, et al.
Published: (2023)
by: Skalse, Joar, et al.
Published: (2023)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
by: Putta, Pranav, et al.
Published: (2024)
by: Putta, Pranav, et al.
Published: (2024)
Learning Latent Dynamic Robust Representations for World Models
by: Sun, Ruixiang, et al.
Published: (2024)
by: Sun, Ruixiang, et al.
Published: (2024)
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
by: Rose, Aaron, et al.
Published: (2026)
by: Rose, Aaron, et al.
Published: (2026)
Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
by: Mitra, Purbesh, et al.
Published: (2025)
by: Mitra, Purbesh, et al.
Published: (2025)
AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
by: Zhang, Zijing, et al.
Published: (2025)
by: Zhang, Zijing, et al.
Published: (2025)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
by: Naghashyar, Lachin, et al.
Published: (2026)
by: Naghashyar, Lachin, et al.
Published: (2026)
PSyDUCK: Training-Free Steganography for Latent Diffusion
by: Mahfuz, Aqib, et al.
Published: (2025)
by: Mahfuz, Aqib, et al.
Published: (2025)
MAD-Sherlock: Multi-Agent Debate for Visual Misinformation Detection
by: Lakara, Kumud, et al.
Published: (2024)
by: Lakara, Kumud, et al.
Published: (2024)
Topological $(\mathscr{F},\mathscr{G})-$shadowing property
by: Joshi, Shital H., et al.
Published: (2025)
by: Joshi, Shital H., et al.
Published: (2025)
On stronger forms of Devaney chaos
by: Joshi, Shital H., et al.
Published: (2025)
by: Joshi, Shital H., et al.
Published: (2025)
On Stronger Forms of Expansivity
by: Joshi, Shital H., et al.
Published: (2024)
by: Joshi, Shital H., et al.
Published: (2024)
Wavelet Predictive Representations for Non-Stationary Reinforcement Learning
by: Wang, Min, et al.
Published: (2025)
by: Wang, Min, et al.
Published: (2025)
Reinforcement Learning for Sequence Design Leveraging Protein Language Models
by: Subramanian, Jithendaraa, et al.
Published: (2024)
by: Subramanian, Jithendaraa, et al.
Published: (2024)
REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
by: Garg, Divyansh, et al.
Published: (2025)
by: Garg, Divyansh, et al.
Published: (2025)
Mixture of Experts Made Intrinsically Interpretable
by: Yang, Xingyi, et al.
Published: (2025)
by: Yang, Xingyi, et al.
Published: (2025)
When Service Matters: Library Budgets 2010
by: London, Charles
Published: (2010)
by: London, Charles
Published: (2010)
Managing and Evaluating Digital Repositories
by: Zuccala, Alesia, et al.
Published: (2008)
by: Zuccala, Alesia, et al.
Published: (2008)
OpenSanctions Pairs: Large-Scale Entity Matching with LLMs
by: Smith, Chandler, et al.
Published: (2026)
by: Smith, Chandler, et al.
Published: (2026)
Extending Pretrained 10-Second ECG Foundation Models to Longer Horizons
by: Tang, Wei, et al.
Published: (2026)
by: Tang, Wei, et al.
Published: (2026)
Sensitivity Analysis for Climate Science with Generative Flow Models
by: Dobra, Alex, et al.
Published: (2025)
by: Dobra, Alex, et al.
Published: (2025)
Maternal age, perimenopause, and Alzheimer's disease on maternal cognitive and behavioral health
by: Alesia V Prakapenka
Published: (2025)
by: Alesia V Prakapenka
Published: (2025)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
Extending the OWASP Multi-Agentic System Threat Modeling Guide: Insights from Multi-Agent Security Research
by: Krawiecka, Klaudia, et al.
Published: (2025)
by: Krawiecka, Klaudia, et al.
Published: (2025)
Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning
by: Luo, Hongyin, et al.
Published: (2025)
by: Luo, Hongyin, et al.
Published: (2025)
CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs
by: Vaghasiya, Jay, et al.
Published: (2025)
by: Vaghasiya, Jay, et al.
Published: (2025)
Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks
by: Franzmeyer, Tim, et al.
Published: (2022)
by: Franzmeyer, Tim, et al.
Published: (2022)
AnnoCaseLaw: A Richly-Annotated Dataset For Benchmarking Explainable Legal Judgment Prediction
by: Sesodia, Magnus, et al.
Published: (2025)
by: Sesodia, Magnus, et al.
Published: (2025)
Emerging adult siblings' relational entitlement and conflict: The moderating effects of financial dependence on parents
by: Weimiao Zhou, et al.
Published: (2024)
by: Weimiao Zhou, et al.
Published: (2024)
Similar Items
-
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
by: Motwani, Sumeet Ramesh, et al.
Published: (2026) -
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
by: Motwani, Sumeet Ramesh, et al.
Published: (2024) -
MALT: Improving Reasoning with Multi-Agent LLM Training
by: Motwani, Sumeet Ramesh, et al.
Published: (2024) -
HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification
by: Wang, Erik Y., et al.
Published: (2026) -
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)