HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Erik Y., Motwani, Sumeet, Roggeveen, James V., Hodges, Eliot, Jayalath, Dulhan, London, Charles, Ramakrishnan, Kalyan, Cipcigan, Flaviu, Torr, Philip, Abate, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaGFN: Exploring Distant Modes with Adapted Metadynamics for Continuous GFlowNets
by: Phillips, Dominic, et al.
Published: (2024)
by: Phillips, Dominic, et al.
Published: (2024)
MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-Training
by: Jayalath, Dulhan, et al.
Published: (2026)
by: Jayalath, Dulhan, et al.
Published: (2026)
Generalising Multi-Agent Cooperation through Task-Agnostic Communication
by: Jayalath, Dulhan, et al.
Published: (2024)
by: Jayalath, Dulhan, et al.
Published: (2024)
h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
by: Motwani, Sumeet Ramesh, et al.
Published: (2025)
by: Motwani, Sumeet Ramesh, et al.
Published: (2025)
Unlocking Non-Invasive Brain-to-Text
by: Jayalath, Dulhan, et al.
Published: (2025)
by: Jayalath, Dulhan, et al.
Published: (2025)
STARC: A General Framework For Quantifying Differences Between Reward Functions
by: Skalse, Joar, et al.
Published: (2023)
by: Skalse, Joar, et al.
Published: (2023)
Discovery of Novel Reticular Materials for Carbon Dioxide Capture using GFlowNets
by: Cipcigan, Flaviu, et al.
Published: (2023)
by: Cipcigan, Flaviu, et al.
Published: (2023)
Neural Proofs for Sound Verification and Control of Complex Systems
by: Abate, Alessandro
Published: (2025)
by: Abate, Alessandro
Published: (2025)
The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised Learning
by: Jayalath, Dulhan, et al.
Published: (2024)
by: Jayalath, Dulhan, et al.
Published: (2024)
PRISM: Efficient Long-Range Reasoning With Short-Context LLMs
by: Jayalath, Dulhan, et al.
Published: (2024)
by: Jayalath, Dulhan, et al.
Published: (2024)
Implicit Neural Representations for Chemical Reaction Paths
by: Ramakrishnan, Kalyan, et al.
Published: (2025)
by: Ramakrishnan, Kalyan, et al.
Published: (2025)
Stochastic Omega-Regular Verification and Control with Supermartingales
by: Abate, Alessandro, et al.
Published: (2024)
by: Abate, Alessandro, et al.
Published: (2024)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
A parametrization of $3$-class groups of quadratic rings over Dedekind domains
by: Hodges, Eliot, et al.
Published: (2025)
by: Hodges, Eliot, et al.
Published: (2025)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
by: Jayalath, Dulhan, et al.
Published: (2025)
by: Jayalath, Dulhan, et al.
Published: (2025)
A General Framework for Verification and Control of Dynamical Models via Certificate Synthesis
by: Edwards, Alec, et al.
Published: (2023)
by: Edwards, Alec, et al.
Published: (2023)
Fossil 2.0: Formal Certificate Synthesis for the Verification and Control of Dynamical Models
by: Edwards, Alec, et al.
Published: (2023)
by: Edwards, Alec, et al.
Published: (2023)
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency
by: Han, Vernon Toh Yan, et al.
Published: (2023)
by: Han, Vernon Toh Yan, et al.
Published: (2023)
LibriBrain: Over 50 Hours of Within-Subject MEG to Improve Speech Decoding Methods at Scale
by: Özdogan, Miran, et al.
Published: (2025)
by: Özdogan, Miran, et al.
Published: (2025)
Equivariant Interatomic Potentials without Tensor Products
by: Reschützegger, Thiago, et al.
Published: (2026)
by: Reschützegger, Thiago, et al.
Published: (2026)
Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning
by: Zhuang, Wenwen, et al.
Published: (2024)
by: Zhuang, Wenwen, et al.
Published: (2024)
Quantitative Verification with Neural Networks
by: Abate, Alessandro, et al.
Published: (2023)
by: Abate, Alessandro, et al.
Published: (2023)
Viability of Big Bang Nucleosynthesis in Some Generalized Horizon Entropies
by: Phukan, Kajal, et al.
Published: (2026)
by: Phukan, Kajal, et al.
Published: (2026)
EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery
by: Ma, Jicheng, et al.
Published: (2026)
by: Ma, Jicheng, et al.
Published: (2026)
Automated Mathematical Discovery and Verification: Minimizing Pentagons in the Plane
by: Subercaseaux, Bernardo, et al.
Published: (2023)
by: Subercaseaux, Bernardo, et al.
Published: (2023)
Discovery, research and development of axalion® active insecticide: dimpropyridaz†
by: Desirée Hodges
Published: (2024)
by: Desirée Hodges
Published: (2024)
Reinforcement Learning for Long-Horizon Multi-Turn Search Agents
by: Kalyan, Vivek, et al.
Published: (2025)
by: Kalyan, Vivek, et al.
Published: (2025)
Meshless solutions of PDE inverse problems on irregular geometries
by: Roggeveen, James V., et al.
Published: (2025)
by: Roggeveen, James V., et al.
Published: (2025)
Emotions in the Loop: A Survey of Affective Computing for Emotional Support
by: Hegde, Karishma, et al.
Published: (2025)
by: Hegde, Karishma, et al.
Published: (2025)
Measurement of Social Well-being and Progress
by: Thorbecke, Erik
by: Thorbecke, Erik
Scalable Verification of Neural Control Barrier Functions Using Linear Bound Propagation
by: Vertovec, Nikolaus, et al.
Published: (2025)
by: Vertovec, Nikolaus, et al.
Published: (2025)
MALT: Improving Reasoning with Multi-Agent LLM Training
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
Pessimistic Verification for Open Ended Math Questions
by: Huang, Yanxing, et al.
Published: (2025)
by: Huang, Yanxing, et al.
Published: (2025)
Automatic Speech Recognition for Hindi
by: Saha, Anish, et al.
Published: (2024)
by: Saha, Anish, et al.
Published: (2024)
An Study on Progress and new Trends of Indian banking company
by: Mr. Devadhe Uddhav Kalyan
Published: (2025)
by: Mr. Devadhe Uddhav Kalyan
Published: (2025)
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
Networked Communication for Decentralised Agents in Mean-Field Games
by: Benjamin, Patrick, et al.
Published: (2023)
by: Benjamin, Patrick, et al.
Published: (2023)
Networked Communication for Mean-Field Games with Function Approximation and Empirical Mean-Field Estimation
by: Benjamin, Patrick, et al.
Published: (2024)
by: Benjamin, Patrick, et al.
Published: (2024)
Similar Items
-
MetaGFN: Exploring Distant Modes with Adapted Metadynamics for Continuous GFlowNets
by: Phillips, Dominic, et al.
Published: (2024) -
MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-Training
by: Jayalath, Dulhan, et al.
Published: (2026) -
Generalising Multi-Agent Cooperation through Task-Agnostic Communication
by: Jayalath, Dulhan, et al.
Published: (2024) -
h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
by: Motwani, Sumeet Ramesh, et al.
Published: (2025) -
Unlocking Non-Invasive Brain-to-Text
by: Jayalath, Dulhan, et al.
Published: (2025)