Salvato in:
| Autori principali: | Nathani, Deepak, Madaan, Lovish, Roberts, Nicholas, Bashlykov, Nikolay, Menon, Ajay, Moens, Vincent, Budhiraja, Amar, Magka, Despoina, Vorotilov, Vladislav, Chaurasia, Gaurav, Hupkes, Dieuwke, Cabral, Ricardo Silveira, Shavrina, Tatiana, Foerster, Jakob, Bachrach, Yoram, Wang, William Yang, Raileanu, Roberta |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.14499 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
di: Maiti, Shalini, et al.
Pubblicazione: (2025)
di: Maiti, Shalini, et al.
Pubblicazione: (2025)
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
di: Madaan, Lovish, et al.
Pubblicazione: (2024)
di: Madaan, Lovish, et al.
Pubblicazione: (2024)
APRES: An Agentic Paper Revision and Evaluation System
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
Quantifying Variance in Evaluation Benchmarks
di: Madaan, Lovish, et al.
Pubblicazione: (2024)
di: Madaan, Lovish, et al.
Pubblicazione: (2024)
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
di: Hupkes, Dieuwke, et al.
Pubblicazione: (2025)
di: Hupkes, Dieuwke, et al.
Pubblicazione: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
di: Pepe, Alberto, et al.
Pubblicazione: (2026)
di: Pepe, Alberto, et al.
Pubblicazione: (2026)
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
di: Ohmer, Xenia, et al.
Pubblicazione: (2024)
di: Ohmer, Xenia, et al.
Pubblicazione: (2024)
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
di: Audran-Reiss, Alexis, et al.
Pubblicazione: (2025)
di: Audran-Reiss, Alexis, et al.
Pubblicazione: (2025)
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
Interpretability of Language Models via Task Spaces
di: Weber, Lucas, et al.
Pubblicazione: (2024)
di: Weber, Lucas, et al.
Pubblicazione: (2024)
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
di: Lupidi, Alisia, et al.
Pubblicazione: (2026)
di: Lupidi, Alisia, et al.
Pubblicazione: (2026)
LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
di: Mondorf, Philipp, et al.
Pubblicazione: (2026)
di: Mondorf, Philipp, et al.
Pubblicazione: (2026)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
di: Toledo, Edan, et al.
Pubblicazione: (2025)
di: Toledo, Edan, et al.
Pubblicazione: (2025)
Bootstrapping Task Spaces for Self-Improvement
di: Jiang, Minqi, et al.
Pubblicazione: (2025)
di: Jiang, Minqi, et al.
Pubblicazione: (2025)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
di: Tang, Yunhao, et al.
Pubblicazione: (2025)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
di: Roberts, Nicholas, et al.
Pubblicazione: (2025)
di: Roberts, Nicholas, et al.
Pubblicazione: (2025)
Asking the Right Questions: Improving Reasoning with Generated Stepping Stones
di: Hu, Hengyuan, et al.
Pubblicazione: (2026)
di: Hu, Hengyuan, et al.
Pubblicazione: (2026)
Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations
di: Thöni, Anna C. M., et al.
Pubblicazione: (2025)
di: Thöni, Anna C. M., et al.
Pubblicazione: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
Scaling Small Agents Through Strategy Auctions
di: Alazraki, Lisa, et al.
Pubblicazione: (2026)
di: Alazraki, Lisa, et al.
Pubblicazione: (2026)
AIRA_2: Overcoming Bottlenecks in AI Research Agents
di: Hambardzumyan, Karen, et al.
Pubblicazione: (2026)
di: Hambardzumyan, Karen, et al.
Pubblicazione: (2026)
HARP: A challenging human-annotated math reasoning benchmark
di: Yue, Albert S., et al.
Pubblicazione: (2024)
di: Yue, Albert S., et al.
Pubblicazione: (2024)
Adversarial Training for Process Reward Models
di: Juneja, Gurusha, et al.
Pubblicazione: (2025)
di: Juneja, Gurusha, et al.
Pubblicazione: (2025)
Crowd IQ -- Aggregating Opinions to Boost Performance
di: Kosinski, Michal, et al.
Pubblicazione: (2024)
di: Kosinski, Michal, et al.
Pubblicazione: (2024)
Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
A detailed study of the variations found in the chrysalises of Aglais caschmirensis Kollar, 1844 (Lepidoptera: Papilionoidea, Nymphalidae)
di: Lovish Garlani
Pubblicazione: (2023)
di: Lovish Garlani
Pubblicazione: (2023)
Annotated Checklist of Rhopalocera of Himachal Pradesh, India (Insecta: Lepidoptera)
di: Lovish Garlan
Pubblicazione: (2024)
di: Lovish Garlan
Pubblicazione: (2024)
First record of Celaenorrhinus ratna daphne Evans, 1949 from Himachal Pradesh and its first photographic record from the Western Himalayas (Lepidoptera: Hesperiidae, Pyrginae)
di: Lovish Garlani
Pubblicazione: (2022)
di: Lovish Garlani
Pubblicazione: (2022)
Unveiling the Hidden Gem: An Observational Report, Taxonomic Insights and First Photographic Evidence of Pseudochazara baldiva Moore, 1865, from India (Lepidoptera: Nymphalidae)
di: Lovish Garlani
Pubblicazione: (2024)
di: Lovish Garlani
Pubblicazione: (2024)
Epistemic Dissonance and Modal Boundaries
di: Raileanu, Dragos
Pubblicazione: (2025)
di: Raileanu, Dragos
Pubblicazione: (2025)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
di: Kornilov, Albert, et al.
Pubblicazione: (2024)
di: Kornilov, Albert, et al.
Pubblicazione: (2024)
A Comparative Study of Transfer Learning for Emotion Recognition using CNN and Modified VGG16 Models
di: Nathani, Samay
Pubblicazione: (2024)
di: Nathani, Samay
Pubblicazione: (2024)
Feature Likelihood Divergence: Evaluating the Generalization of Generative Models Using Samples
di: Jiralerspong, Marco, et al.
Pubblicazione: (2023)
di: Jiralerspong, Marco, et al.
Pubblicazione: (2023)
Modelling Chemical Reaction Networks using Neural Ordinary Differential Equations
di: Thöni, Anna C. M., et al.
Pubblicazione: (2025)
di: Thöni, Anna C. M., et al.
Pubblicazione: (2025)
Understanding the Effects of Domain Finetuning on LLMs
di: Tanwar, Eshaan, et al.
Pubblicazione: (2025)
di: Tanwar, Eshaan, et al.
Pubblicazione: (2025)
On Some Extensions of the Boué-Dupuis Variational Formula
di: Budhiraja, A.
Pubblicazione: (2024)
di: Budhiraja, A.
Pubblicazione: (2024)
Hyperagents
di: Zhang, Jenny, et al.
Pubblicazione: (2026)
di: Zhang, Jenny, et al.
Pubblicazione: (2026)
DUAS FACES DO PODER
di: Peter Bachrach
Pubblicazione: (2011)
di: Peter Bachrach
Pubblicazione: (2011)
Rethinking Thinking Tokens: LLMs as Improvement Operators
di: Madaan, Lovish, et al.
Pubblicazione: (2025)
di: Madaan, Lovish, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
di: Maiti, Shalini, et al.
Pubblicazione: (2025) -
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
di: Madaan, Lovish, et al.
Pubblicazione: (2024) -
APRES: An Agentic Paper Revision and Evaluation System
di: Zhao, Bingchen, et al.
Pubblicazione: (2026) -
Quantifying Variance in Evaluation Benchmarks
di: Madaan, Lovish, et al.
Pubblicazione: (2024) -
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
di: Hupkes, Dieuwke, et al.
Pubblicazione: (2025)