Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
Fuente:
arXiv
Saved in:
| Main Authors: | Sinha, Shiven, Prabhu, Ameya, Kumaraguru, Ponnurangam, Bhat, Siddharth, Bethge, Matthias |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Random Representations Outperform Online Continually Learned Representations
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
by: Sinha, Shiven, et al.
Published: (2025)
by: Sinha, Shiven, et al.
Published: (2025)
Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2
by: Chervonyi, Yuri, et al.
Published: (2025)
by: Chervonyi, Yuri, et al.
Published: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
Corrective Machine Unlearning
by: Goel, Shashwat, et al.
Published: (2024)
by: Goel, Shashwat, et al.
Published: (2024)
The Topological Dual of a Dataset: A Logic-to-Topology Encoding for AlphaGeometry-Style Data
by: Bordg, Anthony
Published: (2026)
by: Bordg, Anthony
Published: (2026)
Great Models Think Alike and this Undermines AI Oversight
by: Goel, Shashwat, et al.
Published: (2025)
by: Goel, Shashwat, et al.
Published: (2025)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
by: Zheng, Zihan, et al.
Published: (2025)
by: Zheng, Zihan, et al.
Published: (2025)
Beyond the Plane: AlphaGeometry Re-architected
by: Bordg, Anthony, et al.
Published: (2026)
by: Bordg, Anthony, et al.
Published: (2026)
Television Discourse Decoded: Comprehensive Multimodal Analytics at Scale
by: Agarwal, Anmol, et al.
Published: (2024)
by: Agarwal, Anmol, et al.
Published: (2024)
Mapping Post-Training Forgetting in Language Models at Scale
by: Harmon, Jackson, et al.
Published: (2025)
by: Harmon, Jackson, et al.
Published: (2025)
Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025
by: Qiu, Jiahao, et al.
Published: (2025)
by: Qiu, Jiahao, et al.
Published: (2025)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
by: Aneja, Krishak, et al.
Published: (2026)
by: Aneja, Krishak, et al.
Published: (2026)
Newclid: A User-Friendly Replacement for AlphaGeometry
by: Sicca, Vladmir, et al.
Published: (2024)
by: Sicca, Vladmir, et al.
Published: (2024)
Virginia Haviland--1976 Regina Medalist--Presentation and Acceptance
by: Field, Carolyn Wicker, et al.
Published: (1976)
by: Field, Carolyn Wicker, et al.
Published: (1976)
Are We Done with Object-Centric Learning?
by: Rubinstein, Alexander, et al.
Published: (2025)
by: Rubinstein, Alexander, et al.
Published: (2025)
Karen Hesse: From Grade School Writer to Newbery Medalist.
by: Brodie, Carolyn S.
Published: (2001)
by: Brodie, Carolyn S.
Published: (2001)
CiteME: Can Language Models Accurately Cite Scientific Claims?
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
by: Ghosh, Adhiraj, et al.
Published: (2024)
by: Ghosh, Adhiraj, et al.
Published: (2024)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
Be a Limitless Pharmacist: Remarks by the 2025 Paul F. Parker Medalist
by: Rita R. Alloway
Published: (2025)
by: Rita R. Alloway
Published: (2025)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025)
by: Hochlehnert, Andreas, et al.
Published: (2025)
Personal Narratives Empower Politically Disinclined Individuals to Engage in Political Discussions
by: Chebrolu, Tejasvi, et al.
Published: (2025)
by: Chebrolu, Tejasvi, et al.
Published: (2025)
Solving Spatial Supersensing Without Spatial Supersensing
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
How to Merge Your Multimodal Models Over Time?
by: Dziadzio, Sebastian, et al.
Published: (2024)
by: Dziadzio, Sebastian, et al.
Published: (2024)
Personalizing Text-to-Image Generation to Individual Taste
by: Maerten, Anne-Sofie, et al.
Published: (2026)
by: Maerten, Anne-Sofie, et al.
Published: (2026)
SceneGraMMi: Scene Graph-boosted Hybrid-fusion for Multi-Modal Misinformation Veracity Prediction
by: Joshi, Swarang, et al.
Published: (2024)
by: Joshi, Swarang, et al.
Published: (2024)
Long-context Non-factoid Question Answering in Indic Languages
by: Mishra, Ritwik, et al.
Published: (2025)
by: Mishra, Ritwik, et al.
Published: (2025)
LLM Vocabulary Compression for Low-Compute Environments
by: Vennam, Sreeram, et al.
Published: (2024)
by: Vennam, Sreeram, et al.
Published: (2024)
Analyzing Patterns and Influence of Advertising in Print Newspapers
by: Vardhan, N Harsha, et al.
Published: (2025)
by: Vardhan, N Harsha, et al.
Published: (2025)
Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations
by: Shankar, Hari, et al.
Published: (2025)
by: Shankar, Hari, et al.
Published: (2025)
VGGSounder: Audio-Visual Evaluations for Foundation Models
by: Zverev, Daniil, et al.
Published: (2025)
by: Zverev, Daniil, et al.
Published: (2025)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
Generalisation of an IMO Geometry Problem
by: Hamzić, Dina Kamber, et al.
Published: (2024)
by: Hamzić, Dina Kamber, et al.
Published: (2024)
Rethinking Thinking Tokens: Understanding Why They Underperform in Practice
by: Vennam, Sreeram, et al.
Published: (2024)
by: Vennam, Sreeram, et al.
Published: (2024)
Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
by: Khatri, Mann, et al.
Published: (2025)
by: Khatri, Mann, et al.
Published: (2025)
Measuring Moral Inconsistencies in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
Multilingual Coreference Resolution in Low-resource South Asian Languages
by: Mishra, Ritwik, et al.
Published: (2024)
by: Mishra, Ritwik, et al.
Published: (2024)
Similar Items
-
Random Representations Outperform Online Continually Learned Representations
by: Prabhu, Ameya, et al.
Published: (2024) -
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
by: Sinha, Shiven, et al.
Published: (2025) -
Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2
by: Chervonyi, Yuri, et al.
Published: (2025) -
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026) -
Corrective Machine Unlearning
by: Goel, Shashwat, et al.
Published: (2024)