miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ospanov, Azim, Farnia, Farzan, Yousefzadeh, Roozbeh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning
von: Ospanov, Azim, et al.
Veröffentlicht: (2025)
von: Ospanov, Azim, et al.
Veröffentlicht: (2025)
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error
von: Farnia, Farzan, et al.
Veröffentlicht: (2026)
von: Farnia, Farzan, et al.
Veröffentlicht: (2026)
A Lean Dataset for International Math Olympiad: Small Steps towards Writing Math Proofs for Hard Problems
von: Yousefzadeh, Roozbeh, et al.
Veröffentlicht: (2024)
von: Yousefzadeh, Roozbeh, et al.
Veröffentlicht: (2024)
Conditional Vendi Score: An Information-Theoretic Approach to Diversity Evaluation of Prompt-based Generative Models
von: Jalali, Mohammad, et al.
Veröffentlicht: (2024)
von: Jalali, Mohammad, et al.
Veröffentlicht: (2024)
Towards a Scalable Reference-Free Evaluation of Generative Models
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
von: Ospanov, Azim, et al.
Veröffentlicht: (2025)
von: Ospanov, Azim, et al.
Veröffentlicht: (2025)
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
Certified Adversarial Robustness via Partition-based Randomized Smoothing
von: Goli, Hossein, et al.
Veröffentlicht: (2024)
von: Goli, Hossein, et al.
Veröffentlicht: (2024)
Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs
von: Yousefzadeh, Roozbeh, et al.
Veröffentlicht: (2025)
von: Yousefzadeh, Roozbeh, et al.
Veröffentlicht: (2025)
When Exploration Comes for Free with Mixture-Greedy: Do we need UCB in Diversity-Aware Multi-Armed Bandits?
von: Nia, Bahar Dibaei, et al.
Veröffentlicht: (2026)
von: Nia, Bahar Dibaei, et al.
Veröffentlicht: (2026)
On the Inductive Biases of Demographic Parity-based Fair Learning Algorithms
von: Lei, Haoyu, et al.
Veröffentlicht: (2024)
von: Lei, Haoyu, et al.
Veröffentlicht: (2024)
Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
von: Oskouie, Haniyeh Ehsani, et al.
Veröffentlicht: (2022)
von: Oskouie, Haniyeh Ehsani, et al.
Veröffentlicht: (2022)
PromptSplit: Revealing Prompt-Level Disagreement in Generative Models
von: Lotfian, Mehdi, et al.
Veröffentlicht: (2026)
von: Lotfian, Mehdi, et al.
Veröffentlicht: (2026)
PromptWise: Online Learning for Cost-Aware Prompt Assignment in Generative Models
von: Hu, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoyan, et al.
Veröffentlicht: (2025)
Towards an Explainable Comparison and Alignment of Feature Embeddings
von: Jalali, Mohammad, et al.
Veröffentlicht: (2025)
von: Jalali, Mohammad, et al.
Veröffentlicht: (2025)
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE Score
von: Jalali, Mohammad, et al.
Veröffentlicht: (2025)
von: Jalali, Mohammad, et al.
Veröffentlicht: (2025)
Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time Adaptation
von: Lei, Haoyu, et al.
Veröffentlicht: (2025)
von: Lei, Haoyu, et al.
Veröffentlicht: (2025)
On the Fragility of AI-Based Channel Decoders under Small Channel Perturbations
von: Lei, Haoyu, et al.
Veröffentlicht: (2026)
von: Lei, Haoyu, et al.
Veröffentlicht: (2026)
Syndrome-Flow Consistency Model Achieves One-step Denoising Error Correction Codes
von: Lei, Haoyu, et al.
Veröffentlicht: (2025)
von: Lei, Haoyu, et al.
Veröffentlicht: (2025)
Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance
von: Sani, Matina Mahdizadeh, et al.
Veröffentlicht: (2026)
von: Sani, Matina Mahdizadeh, et al.
Veröffentlicht: (2026)
PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
von: Zou, Lancheng, et al.
Veröffentlicht: (2025)
von: Zou, Lancheng, et al.
Veröffentlicht: (2025)
From RISC-V Cores to Neuromorphic Arrays: A Tutorial on Building Scalable Digital Neuromorphic Processors
von: Yousefzadeh, Amirreza
Veröffentlicht: (2025)
von: Yousefzadeh, Amirreza
Veröffentlicht: (2025)
ChatPattern: Layout Pattern Customization via Natural Language
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
Scaling Item-to-Standard Alignment with Large Language Models: Accuracy, Limits, and Solutions
von: Karimi-Malekabadi, Farzan, et al.
Veröffentlicht: (2025)
von: Karimi-Malekabadi, Farzan, et al.
Veröffentlicht: (2025)
Text2Graph VPR: A Text-to-Graph Expert System for Explainable Place Recognition in Changing Environments
von: Yousefzadeh, Saeideh, et al.
Veröffentlicht: (2025)
von: Yousefzadeh, Saeideh, et al.
Veröffentlicht: (2025)
Exploring the Limitations of Layer Synchronization in Spiking Neural Networks
von: Koopman, Roel, et al.
Veröffentlicht: (2024)
von: Koopman, Roel, et al.
Veröffentlicht: (2024)
Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward
von: Patro, Gourab K, et al.
Veröffentlicht: (2026)
von: Patro, Gourab K, et al.
Veröffentlicht: (2026)
LeanGeo: Formalizing Competitional Geometry problems in Lean
von: Song, Chendong, et al.
Veröffentlicht: (2025)
von: Song, Chendong, et al.
Veröffentlicht: (2025)
ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
von: Li, Xinhang, et al.
Veröffentlicht: (2025)
von: Li, Xinhang, et al.
Veröffentlicht: (2025)
SynChart: Synthesizing Charts from Language Models
von: Liu, Mengchen, et al.
Veröffentlicht: (2024)
von: Liu, Mengchen, et al.
Veröffentlicht: (2024)
ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2025)
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2025)
Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
von: Shi, Yaorui, et al.
Veröffentlicht: (2025)
von: Shi, Yaorui, et al.
Veröffentlicht: (2025)
ChartDiff: A Large-Scale Benchmark for Comprehending Pairs of Charts
von: Ye, Rongtian
Veröffentlicht: (2026)
von: Ye, Rongtian
Veröffentlicht: (2026)
LeanSearch v2: Global Premise Retrieval for Lean 4 Theorem Proving
von: Gao, Guoxiong, et al.
Veröffentlicht: (2026)
von: Gao, Guoxiong, et al.
Veröffentlicht: (2026)
Text2Chart31: Instruction Tuning for Chart Generation with Automatic Feedback
von: Zadeh, Fatemeh Pesaran, et al.
Veröffentlicht: (2024)
von: Zadeh, Fatemeh Pesaran, et al.
Veröffentlicht: (2024)
ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
Sociodemographic Bias in Language Models: A Survey and Forward Path
von: Gupta, Vipul, et al.
Veröffentlicht: (2023)
von: Gupta, Vipul, et al.
Veröffentlicht: (2023)
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
von: Yao, Wei, et al.
Veröffentlicht: (2025)
von: Yao, Wei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning
von: Ospanov, Azim, et al.
Veröffentlicht: (2025) -
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
von: Ospanov, Azim, et al.
Veröffentlicht: (2024) -
Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error
von: Farnia, Farzan, et al.
Veröffentlicht: (2026) -
A Lean Dataset for International Math Olympiad: Small Steps towards Writing Math Proofs for Hard Problems
von: Yousefzadeh, Roozbeh, et al.
Veröffentlicht: (2024) -
Conditional Vendi Score: An Information-Theoretic Approach to Diversity Evaluation of Prompt-based Generative Models
von: Jalali, Mohammad, et al.
Veröffentlicht: (2024)