Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Maiti, Shalini, Budhiraja, Amar, Gauri, Bhavul, Chaurasia, Gaurav, Protopopov, Anton, Audran-Reiss, Alexis, Slater, Michael, Magka, Despoina, Shavrina, Tatiana, Raileanu, Roberta, Bachrach, Yoram |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
by: Nathani, Deepak, et al.
Published: (2025)
by: Nathani, Deepak, et al.
Published: (2025)
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
by: Audran-Reiss, Alexis, et al.
Published: (2025)
by: Audran-Reiss, Alexis, et al.
Published: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
by: Pepe, Alberto, et al.
Published: (2026)
by: Pepe, Alberto, et al.
Published: (2026)
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
by: Lupidi, Alisia, et al.
Published: (2026)
by: Lupidi, Alisia, et al.
Published: (2026)
APRES: An Agentic Paper Revision and Evaluation System
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025)
by: Toledo, Edan, et al.
Published: (2025)
AIRA_2: Overcoming Bottlenecks in AI Research Agents
by: Hambardzumyan, Karen, et al.
Published: (2026)
by: Hambardzumyan, Karen, et al.
Published: (2026)
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
Bootstrapping Task Spaces for Self-Improvement
by: Jiang, Minqi, et al.
Published: (2025)
by: Jiang, Minqi, et al.
Published: (2025)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
by: Kornilov, Albert, et al.
Published: (2024)
by: Kornilov, Albert, et al.
Published: (2024)
Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations
by: Thöni, Anna C. M., et al.
Published: (2025)
by: Thöni, Anna C. M., et al.
Published: (2025)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
by: Alexey, Protopopov
Published: (2026)
by: Alexey, Protopopov
Published: (2026)
A Robust Remote Photoplethysmography Method
by: Protopopov, Alexey
Published: (2025)
by: Protopopov, Alexey
Published: (2025)
A Robust Camera-based Method for Breath Rate Measurement
by: Protopopov, Alexey
Published: (2025)
by: Protopopov, Alexey
Published: (2025)
Scaling Small Agents Through Strategy Auctions
by: Alazraki, Lisa, et al.
Published: (2026)
by: Alazraki, Lisa, et al.
Published: (2026)
LLM-First Search: Self-Guided Exploration of the Solution Space
by: Herr, Nathan, et al.
Published: (2025)
by: Herr, Nathan, et al.
Published: (2025)
PrediPrune: Reducing Verification Overhead in Souper with Machine Learning Driven Pruning
by: Ishimwe, Ange-Thierry, et al.
Published: (2025)
by: Ishimwe, Ange-Thierry, et al.
Published: (2025)
Epistemic Dissonance and Modal Boundaries
by: Raileanu, Dragos
Published: (2025)
by: Raileanu, Dragos
Published: (2025)
Crowd IQ -- Aggregating Opinions to Boost Performance
by: Kosinski, Michal, et al.
Published: (2024)
by: Kosinski, Michal, et al.
Published: (2024)
On Some Extensions of the Boué-Dupuis Variational Formula
by: Budhiraja, A.
Published: (2024)
by: Budhiraja, A.
Published: (2024)
The Generalization Gap in Offline Reinforcement Learning
by: Mediratta, Ishita, et al.
Published: (2023)
by: Mediratta, Ishita, et al.
Published: (2023)
Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
by: Maiti, Shalini, et al.
Published: (2025)
by: Maiti, Shalini, et al.
Published: (2025)
Unsupervised 2D-3D lifting of non-rigid objects using local constraints
by: Maiti, Shalini, et al.
Published: (2025)
by: Maiti, Shalini, et al.
Published: (2025)
Feature Likelihood Divergence: Evaluating the Generalization of Generative Models Using Samples
by: Jiralerspong, Marco, et al.
Published: (2023)
by: Jiralerspong, Marco, et al.
Published: (2023)
Modelling Chemical Reaction Networks using Neural Ordinary Differential Equations
by: Thöni, Anna C. M., et al.
Published: (2025)
by: Thöni, Anna C. M., et al.
Published: (2025)
DUAS FACES DO PODER
by: Peter Bachrach
Published: (2011)
by: Peter Bachrach
Published: (2011)
Asking the Right Questions: Improving Reasoning with Generated Stepping Stones
by: Hu, Hengyuan, et al.
Published: (2026)
by: Hu, Hengyuan, et al.
Published: (2026)
DreamCraft: Text-Guided Generation of Functional 3D Environments in Minecraft
by: Earle, Sam, et al.
Published: (2024)
by: Earle, Sam, et al.
Published: (2024)
GazeProphetV2: Head-Movement-Based Gaze Prediction Enabling Efficient Foveated Rendering on Mobile VR
by: Ebadulla, Farhaan, et al.
Published: (2025)
by: Ebadulla, Farhaan, et al.
Published: (2025)
Copyright Aspects of CATV as Utilized in Information Networking.
by: Bachrach, Morton W.
Published: (1970)
by: Bachrach, Morton W.
Published: (1970)
On free boundary problems for the Atlas model
by: Atar, Rami, et al.
Published: (2025)
by: Atar, Rami, et al.
Published: (2025)
Jump Processes with Self-Interactions: Large Deviation Asymptotics
by: Budhiraja, Amarjit, et al.
Published: (2025)
by: Budhiraja, Amarjit, et al.
Published: (2025)
Extremal Invariant Distributions of Infinite Brownian Particle Systems with Rank Dependent Drifts
by: Banerjee, Sayan, et al.
Published: (2022)
by: Banerjee, Sayan, et al.
Published: (2022)
Large Deviation Asymptotics for the Supermarket Model with Growing Choices
by: Budhiraja, Amarjit, et al.
Published: (2025)
by: Budhiraja, Amarjit, et al.
Published: (2025)
Diffusion limits in the quarter plane and non-semimartingale reflected Brownian motion
by: Atar, Rami, et al.
Published: (2024)
by: Atar, Rami, et al.
Published: (2024)
Are Large Language Models Strategic Decision Makers? A Study of Performance and Bias in Two-Player Non-Zero-Sum Games
by: Herr, Nathan, et al.
Published: (2024)
by: Herr, Nathan, et al.
Published: (2024)
Enceladus's Tidal Heating
by: Lithwick, Yoram
Published: (2025)
by: Lithwick, Yoram
Published: (2025)
Secret leviathan: Secrecy and state capacity under Soviet Communism.MarkHarrison, (Stanford, CA: Stanford University Press, 2023. pp. 372. 9 figs. 23 tabs. ISBN: 9781503628892 $65)
by: Yoram Gorlizki
Published: (2024)
by: Yoram Gorlizki
Published: (2024)
Los derechos humanos internacionales de los judíos soviéticos
by: Dinstein, Yoram
Published: (1974)
by: Dinstein, Yoram
Published: (1974)
Similar Items
-
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
by: Nathani, Deepak, et al.
Published: (2025) -
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
by: Audran-Reiss, Alexis, et al.
Published: (2025) -
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
by: Pepe, Alberto, et al.
Published: (2026) -
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
by: Lupidi, Alisia, et al.
Published: (2026) -
APRES: An Agentic Paper Revision and Evaluation System
by: Zhao, Bingchen, et al.
Published: (2026)