Putnam-like dataset summary: LLMs as mathematical competition contestants
Fuente:
arXiv
Guardado en:
| Autores principales: | Bieganowski, Bartosz, Strzelecki, Daniel, Skiba, Robert, Topolewski, Mateusz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Nonlinear scalar field equations with a critical Hardy potential
por: Bieganowski, Bartosz, et al.
Publicado: (2025)
por: Bieganowski, Bartosz, et al.
Publicado: (2025)
Note on the multiplicity of solutions for nonlinear scalar field equations with a critical inverse-square potential
por: Bieganowski, Bartosz, et al.
Publicado: (2026)
por: Bieganowski, Bartosz, et al.
Publicado: (2026)
Supervised Autoencoder MLP for Financial Time Series Forecasting
por: Bieganowski, Bartosz, et al.
Publicado: (2024)
por: Bieganowski, Bartosz, et al.
Publicado: (2024)
Supervised Autoencoders with Fractionally Differentiated Features and Triple Barrier Labelling Enhance Predictions on Noisy Data
por: Bieganowski, Bartosz, et al.
Publicado: (2024)
por: Bieganowski, Bartosz, et al.
Publicado: (2024)
Multiplicity of critical orbits to nonlinear, strongly indefinite functionals with sign-changing nonlinear part
por: Bernini, Federico, et al.
Publicado: (2024)
por: Bernini, Federico, et al.
Publicado: (2024)
Note on homoclinic solutions to nonautonomous Hamiltonian systems with sign-changing nonlinear part
por: Bernini, Federico, et al.
Publicado: (2024)
por: Bernini, Federico, et al.
Publicado: (2024)
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
por: Tsoukalas, George, et al.
Publicado: (2024)
por: Tsoukalas, George, et al.
Publicado: (2024)
Putnam's Critical and Explanatory Tendencies Interpreted from a Machine Learning Perspective
por: Soudin, Sheldon Z.
Publicado: (2025)
por: Soudin, Sheldon Z.
Publicado: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
por: Cywiński, Bartosz, et al.
Publicado: (2025)
por: Cywiński, Bartosz, et al.
Publicado: (2025)
Self-rewarding correction for mathematical reasoning
por: Xiong, Wei, et al.
Publicado: (2025)
por: Xiong, Wei, et al.
Publicado: (2025)
LucidPPN: Unambiguous Prototypical Parts Network for User-centric Interpretable Computer Vision
por: Pach, Mateusz, et al.
Publicado: (2024)
por: Pach, Mateusz, et al.
Publicado: (2024)
TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration
por: Olszewski, Jan, et al.
Publicado: (2023)
por: Olszewski, Jan, et al.
Publicado: (2023)
Pre-Hoc Predictions in AutoML: Leveraging LLMs to Enhance Model Selection and Benchmarking for Tabular datasets
por: Belkhiter, Yannis, et al.
Publicado: (2025)
por: Belkhiter, Yannis, et al.
Publicado: (2025)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
por: Casademunt, Helena, et al.
Publicado: (2026)
por: Casademunt, Helena, et al.
Publicado: (2026)
Neuroplasticity-inspired dynamic ANNs for multi-task demand forecasting
por: Żarski, Mateusz, et al.
Publicado: (2025)
por: Żarski, Mateusz, et al.
Publicado: (2025)
GraphXAIN: Narratives to Explain Graph Neural Networks
por: Cedro, Mateusz, et al.
Publicado: (2024)
por: Cedro, Mateusz, et al.
Publicado: (2024)
HAL: Inducing Human-likeness in LLMs with Alignment
por: Hasan, Masum, et al.
Publicado: (2026)
por: Hasan, Masum, et al.
Publicado: (2026)
A mathematical theory of balancing relational generalization and memorization
por: Cheng, Luke, et al.
Publicado: (2026)
por: Cheng, Luke, et al.
Publicado: (2026)
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
por: Chen, Yimeng, et al.
Publicado: (2025)
por: Chen, Yimeng, et al.
Publicado: (2025)
Math Takes Two: A test for emergent mathematical reasoning in communication
por: Cooper, Michael, et al.
Publicado: (2026)
por: Cooper, Michael, et al.
Publicado: (2026)
Red-teaming Activation Probes using Prompted LLMs
por: Blandfort, Phil, et al.
Publicado: (2025)
por: Blandfort, Phil, et al.
Publicado: (2025)
Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning
por: Praski, Mateusz, et al.
Publicado: (2025)
por: Praski, Mateusz, et al.
Publicado: (2025)
A multi-criteria approach for selecting an explanation from the set of counterfactuals produced by an ensemble of explainers
por: Stępka, Ignacy, et al.
Publicado: (2024)
por: Stępka, Ignacy, et al.
Publicado: (2024)
Counterfactual Explanations with Probabilistic Guarantees on their Robustness to Model Change
por: Stępka, Ignacy, et al.
Publicado: (2024)
por: Stępka, Ignacy, et al.
Publicado: (2024)
The ecosystem of machine learning competitions: Platforms, participants, and their impact on AI development
por: Nasios, Ioannis
Publicado: (2026)
por: Nasios, Ioannis
Publicado: (2026)
Human-like Working Memory Interference in Large Language Models
por: Xiong, Hua-Dong, et al.
Publicado: (2026)
por: Xiong, Hua-Dong, et al.
Publicado: (2026)
SONG: Self-Organizing Neural Graphs
por: Struski, Łukasz, et al.
Publicado: (2021)
por: Struski, Łukasz, et al.
Publicado: (2021)
ClickAgent: Enhancing UI Location Capabilities of Autonomous Agents
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
por: Hoscilowicz, Jakub, et al.
Publicado: (2024)
Ratio law: mathematical descriptions for a universal relationship between AI performance and input samples
por: Kang, Boming, et al.
Publicado: (2024)
por: Kang, Boming, et al.
Publicado: (2024)
Federated Fine-Tuning of LLMs: Framework Comparison and Research Directions
por: Yan, Na, et al.
Publicado: (2025)
por: Yan, Na, et al.
Publicado: (2025)
Is Temporal Difference Learning the Gold Standard for Stitching in RL?
por: Bortkiewicz, Michał, et al.
Publicado: (2025)
por: Bortkiewicz, Michał, et al.
Publicado: (2025)
Scaling FP8 training to trillion-token LLMs
por: Fishman, Maxim, et al.
Publicado: (2024)
por: Fishman, Maxim, et al.
Publicado: (2024)
FP4 All the Way: Fully Quantized Training of LLMs
por: Chmiel, Brian, et al.
Publicado: (2025)
por: Chmiel, Brian, et al.
Publicado: (2025)
ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models
por: Sienkiewicz, Bruno, et al.
Publicado: (2026)
por: Sienkiewicz, Bruno, et al.
Publicado: (2026)
ASTE Transformer Modelling Dependencies in Aspect-Sentiment Triplet Extraction
por: Naglik, Iwo, et al.
Publicado: (2024)
por: Naglik, Iwo, et al.
Publicado: (2024)
Enhanced N-BEATS for Mid-Term Electricity Demand Forecasting
por: Kasprzyk, Mateusz, et al.
Publicado: (2024)
por: Kasprzyk, Mateusz, et al.
Publicado: (2024)
Int2Int: a framework for mathematics with transformers
por: Charton, François
Publicado: (2025)
por: Charton, François
Publicado: (2025)
Cueless EEG imagined speech for subject identification: dataset and benchmarks
por: Derakhshesh, Ali, et al.
Publicado: (2025)
por: Derakhshesh, Ali, et al.
Publicado: (2025)
Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs
por: Asante, Daniel Agyei, et al.
Publicado: (2025)
por: Asante, Daniel Agyei, et al.
Publicado: (2025)
The Deep-Match Framework for Event-Related Potential Detection in EEG
por: Zylinski, Marek, et al.
Publicado: (2026)
por: Zylinski, Marek, et al.
Publicado: (2026)
Ejemplares similares
-
Nonlinear scalar field equations with a critical Hardy potential
por: Bieganowski, Bartosz, et al.
Publicado: (2025) -
Note on the multiplicity of solutions for nonlinear scalar field equations with a critical inverse-square potential
por: Bieganowski, Bartosz, et al.
Publicado: (2026) -
Supervised Autoencoder MLP for Financial Time Series Forecasting
por: Bieganowski, Bartosz, et al.
Publicado: (2024) -
Supervised Autoencoders with Fractionally Differentiated Features and Triple Barrier Labelling Enhance Predictions on Noisy Data
por: Bieganowski, Bartosz, et al.
Publicado: (2024) -
Multiplicity of critical orbits to nonlinear, strongly indefinite functionals with sign-changing nonlinear part
por: Bernini, Federico, et al.
Publicado: (2024)