Winning Gold at IMO 2025 with a Model-Agnostic Verification-and-Refinement Pipeline
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yichen, Yang, Lin F. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vibe Reasoning: Eliciting Frontier AI Mathematical Capabilities -- A Case Study on IMO 2025 Problem 6
by: Wu, Jiaao, et al.
Published: (2025)
by: Wu, Jiaao, et al.
Published: (2025)
Aristotle: IMO-level Automated Theorem Proving
by: Achim, Tudor, et al.
Published: (2025)
by: Achim, Tudor, et al.
Published: (2025)
Perfect score on IPhO 2025 theory by Gemini agent
by: Huang, Yichen
Published: (2026)
by: Huang, Yichen
Published: (2026)
PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System
by: Yu, Fangchen, et al.
Published: (2025)
by: Yu, Fangchen, et al.
Published: (2025)
Towards Solving More Challenging IMO Problems via Decoupled Reasoning and Proving
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
by: Sinha, Shiven, et al.
Published: (2024)
by: Sinha, Shiven, et al.
Published: (2024)
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
by: Yuan, Xiu, et al.
Published: (2024)
by: Yuan, Xiu, et al.
Published: (2024)
AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing
by: Du, Jianda, et al.
Published: (2026)
by: Du, Jianda, et al.
Published: (2026)
Multilingual Controlled Generation And Gold-Standard-Agnostic Evaluation of Code-Mixed Sentences
by: Gupta, Ayushman, et al.
Published: (2024)
by: Gupta, Ayushman, et al.
Published: (2024)
JOURNALISTS' PERCEPTION OF THE USE OF ARTIFICIAL INTELLIGENCE, AI IN NEWS REPORTAGE IN IMO STATE
by: JULIAN, Chijioke Godswill & OHAEGBULAM, Nwakaego
Published: (2025)
by: JULIAN, Chijioke Godswill & OHAEGBULAM, Nwakaego
Published: (2025)
Joint Verification and Refinement of Language Models for Safety-Constrained Planning
by: Yang, Yunhao, et al.
Published: (2024)
by: Yang, Yunhao, et al.
Published: (2024)
Pessimistic Verification for Open Ended Math Questions
by: Huang, Yanxing, et al.
Published: (2025)
by: Huang, Yanxing, et al.
Published: (2025)
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
by: Singhvi, Arnav, et al.
Published: (2023)
by: Singhvi, Arnav, et al.
Published: (2023)
RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement
by: Jiang, Jinhao, et al.
Published: (2024)
by: Jiang, Jinhao, et al.
Published: (2024)
Mini-Giants: "Small" Language Models and Open Source Win-Win
by: Zhou, Zhengping, et al.
Published: (2023)
by: Zhou, Zhengping, et al.
Published: (2023)
Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuning
by: Cai, Hongyi, et al.
Published: (2025)
by: Cai, Hongyi, et al.
Published: (2025)
Repairing Language Model Pipelines by Meta Self-Refining Competing Constraints at Runtime
by: Eshghie, Mojtaba
Published: (2025)
by: Eshghie, Mojtaba
Published: (2025)
Precise Verification of Transformers through ReLU-Catalyzed Abstraction Refinement
by: Liu, Hengjie, et al.
Published: (2026)
by: Liu, Hengjie, et al.
Published: (2026)
Find A Winning Sign: Sign Is All We Need to Win the Lottery
by: Oh, Junghun, et al.
Published: (2025)
by: Oh, Junghun, et al.
Published: (2025)
Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025
by: Qiu, Jiahao, et al.
Published: (2025)
by: Qiu, Jiahao, et al.
Published: (2025)
Reinforcing VLAs in Task-Agnostic World Models
by: Wang, Yucen, et al.
Published: (2026)
by: Wang, Yucen, et al.
Published: (2026)
STEVE: A Step Verification Pipeline for Computer-use Agent Training
by: Lu, Fanbin, et al.
Published: (2025)
by: Lu, Fanbin, et al.
Published: (2025)
Boosting Efficiency in Task-Agnostic Exploration through Causal Knowledge
by: Yang, Yupei, et al.
Published: (2024)
by: Yang, Yupei, et al.
Published: (2024)
MASE: Interpretable NLP Models via Model-Agnostic Saliency Estimation
by: Yang, Zhou, et al.
Published: (2025)
by: Yang, Zhou, et al.
Published: (2025)
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
by: Hu, Yunhai, et al.
Published: (2026)
by: Hu, Yunhai, et al.
Published: (2026)
Majority Rules: LLM Ensemble is a Winning Approach for Content Categorization
by: Kamen, Ariel, et al.
Published: (2025)
by: Kamen, Ariel, et al.
Published: (2025)
Winning Snake: Design Choices in Multi-Shot ASP
by: Böhl, Elisa, et al.
Published: (2024)
by: Böhl, Elisa, et al.
Published: (2024)
Towards Understanding and Enhancing Security of Proof-of-Training for DNN Model Ownership Verification
by: Chang, Yijia, et al.
Published: (2024)
by: Chang, Yijia, et al.
Published: (2024)
LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models
by: Kim, Yungi, et al.
Published: (2024)
by: Kim, Yungi, et al.
Published: (2024)
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
by: Xie, Yingsha, et al.
Published: (2026)
by: Xie, Yingsha, et al.
Published: (2026)
BoxRL-NNV: Boxed Refinement of Latin Hypercube Samples for Neural Network Verification
by: Das, Sarthak
Published: (2025)
by: Das, Sarthak
Published: (2025)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
by: Wang, Qibin, et al.
Published: (2025)
by: Wang, Qibin, et al.
Published: (2025)
SmartPretrain: Model-Agnostic and Dataset-Agnostic Representation Learning for Motion Prediction
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
AdaSTI: Conditional Diffusion Models with Adaptive Dependency Modeling for Spatio-Temporal Imputation
by: Yang, Yubo, et al.
Published: (2025)
by: Yang, Yubo, et al.
Published: (2025)
Win-k: Improved Membership Inference Attacks on Small Language Models
by: Arkhmammadova, Roya, et al.
Published: (2025)
by: Arkhmammadova, Roya, et al.
Published: (2025)
Beyond Winning: Margin of Victory Relative to Expectation Unlocks Accurate Skill Ratings
by: Shorewala, Shivam, et al.
Published: (2025)
by: Shorewala, Shivam, et al.
Published: (2025)
Physics-Informed Regularization for Domain-Agnostic Dynamical System Modeling
by: Huang, Zijie, et al.
Published: (2024)
by: Huang, Zijie, et al.
Published: (2024)
Progressive Refinement Regulation for Accelerating Diffusion Language Model Decoding
by: Wan, Lipeng, et al.
Published: (2026)
by: Wan, Lipeng, et al.
Published: (2026)
Design a Win-Win Strategy That Is Fair to Both Service Providers and Tasks When Rejection Is Not an Option
by: Trabelsi, Yohai, et al.
Published: (2024)
by: Trabelsi, Yohai, et al.
Published: (2024)
AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models
by: Jia, Tingzheng, et al.
Published: (2026)
by: Jia, Tingzheng, et al.
Published: (2026)
Similar Items
-
Vibe Reasoning: Eliciting Frontier AI Mathematical Capabilities -- A Case Study on IMO 2025 Problem 6
by: Wu, Jiaao, et al.
Published: (2025) -
Aristotle: IMO-level Automated Theorem Proving
by: Achim, Tudor, et al.
Published: (2025) -
Perfect score on IPhO 2025 theory by Gemini agent
by: Huang, Yichen
Published: (2026) -
PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System
by: Yu, Fangchen, et al.
Published: (2025) -
Towards Solving More Challenging IMO Problems via Decoupled Reasoning and Proving
by: Liang, Zhenwen, et al.
Published: (2025)