Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Etzine, Bryan, Hashemi, Masoud, Madhusudhan, Nishanth, Davasam, Sagar, Sharma, Roshnee, Madhusudhan, Sathwik Tejaswi, Yadav, Vikas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
by: Madhusudhan, Nishanth, et al.
Published: (2024)
by: Madhusudhan, Nishanth, et al.
Published: (2024)
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
by: Hashemi, Masoud, et al.
Published: (2025)
by: Hashemi, Masoud, et al.
Published: (2025)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
by: Tiwari, Aman, et al.
Published: (2024)
by: Tiwari, Aman, et al.
Published: (2024)
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
by: Maheshwary, Rishabh, et al.
Published: (2024)
by: Maheshwary, Rishabh, et al.
Published: (2024)
DeepSRGM -- Sequence Classification and Ranking in Indian Classical Music with Deep Learning
by: Madhusudhan, Sathwik Tejaswi, et al.
Published: (2024)
by: Madhusudhan, Sathwik Tejaswi, et al.
Published: (2024)
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
by: Madhusudhan, Nishanth, et al.
Published: (2026)
by: Madhusudhan, Nishanth, et al.
Published: (2026)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
by: Pattnaik, Pulkit, et al.
Published: (2024)
by: Pattnaik, Pulkit, et al.
Published: (2024)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
by: Maheshwary, Rishabh, et al.
Published: (2025)
by: Maheshwary, Rishabh, et al.
Published: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
Grammar Search for Multi-Agent Systems
by: Singh, Mayank, et al.
Published: (2025)
by: Singh, Mayank, et al.
Published: (2025)
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
Size and shape of terrestrial animals
by: Sharma, Neelima, et al.
Published: (2026)
by: Sharma, Neelima, et al.
Published: (2026)
The Hycean Paradigm in the Search for Life Elsewhere
by: Madhusudhan, Nikku
Published: (2024)
by: Madhusudhan, Nikku
Published: (2024)
Habitability and Biosignatures
by: Madhusudhan, Nikku
Published: (2025)
by: Madhusudhan, Nikku
Published: (2025)
RFID Technology Implementation in Two Libraries in New Delhi
by: Madhusudhan, Margam
Published: (2010)
by: Madhusudhan, Margam
Published: (2010)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
General Circulation Models of Hycean Worlds
by: Barrier, Edouard, et al.
Published: (2025)
by: Barrier, Edouard, et al.
Published: (2025)
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
by: Nguyen, Hoang, et al.
Published: (2025)
by: Nguyen, Hoang, et al.
Published: (2025)
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
by: Bogavelli, Tara, et al.
Published: (2026)
by: Bogavelli, Tara, et al.
Published: (2026)
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
by: Bogavelli, Tara, et al.
Published: (2025)
by: Bogavelli, Tara, et al.
Published: (2025)
Apriel-1.5-15b-Thinker
by: Radhakrishna, Shruthan, et al.
Published: (2025)
by: Radhakrishna, Shruthan, et al.
Published: (2025)
Considerations for Photochemical Modeling of Possible Hycean Worlds
by: Cooke, Gregory J., et al.
Published: (2024)
by: Cooke, Gregory J., et al.
Published: (2024)
Characterising M dwarf host stars of two candidate Hycean worlds
by: Sairam, Lalitha, et al.
Published: (2025)
by: Sairam, Lalitha, et al.
Published: (2025)
Trace Relations in Deformed Gauge Theories
by: Raman, Madhusudhan, et al.
Published: (2024)
by: Raman, Madhusudhan, et al.
Published: (2024)
Feasibility of High-Resolution Transmission Spectroscopy for Low-Velocity Exoplanets
by: Cheverall, Connor, et al.
Published: (2024)
by: Cheverall, Connor, et al.
Published: (2024)
Possible Hycean conditions in the sub-Neptune TOI-270 d
by: Holmberg, Måns, et al.
Published: (2024)
by: Holmberg, Måns, et al.
Published: (2024)
VIRA: An Exoplanet Atmospheric Retrieval Framework for JWST Transmission Spectroscopy
by: Constantinou, Savvas, et al.
Published: (2024)
by: Constantinou, Savvas, et al.
Published: (2024)
Web-Based Online Public Access Catalogues of IIT Libraries in India: An Evaluative Study
by: Madhusudhan, Margam, et al.
Published: (2011)
by: Madhusudhan, Margam, et al.
Published: (2011)
An Algorithmic Upper Bound for Permanents via a Permanental Schur Inequality
by: Laddha, Aditi, et al.
Published: (2025)
by: Laddha, Aditi, et al.
Published: (2025)
On the Ocean Conditions of Hycean Worlds
by: Rigby, Frances E., et al.
Published: (2024)
by: Rigby, Frances E., et al.
Published: (2024)
Prospects for biological evolution on Hycean worlds
by: Mitchell, Emily G., et al.
Published: (2025)
by: Mitchell, Emily G., et al.
Published: (2025)
The Surface and Interior Conditions of Temperate Sub-Neptune TOI-270 d
by: Rigby, Frances E., et al.
Published: (2025)
by: Rigby, Frances E., et al.
Published: (2025)
ClaimHack: A Benchmark Dataset of Check-worthy Factual Claims
by: Rayapet Madhusudhan, Akshay Kumar, et al.
Published: (2025)
by: Rayapet Madhusudhan, Akshay Kumar, et al.
Published: (2025)
Approximation Algorithms for the Weighted Nash Social Welfare via Convex and Non-Convex Programs
by: Brown, Adam, et al.
Published: (2024)
by: Brown, Adam, et al.
Published: (2024)
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
by: Rajeev, Meghana, et al.
Published: (2025)
by: Rajeev, Meghana, et al.
Published: (2025)
A new convection scheme for GCMs of temperate sub-Neptunes
by: Barrier, Edouard F. L., et al.
Published: (2025)
by: Barrier, Edouard F. L., et al.
Published: (2025)
The atmospheric composition of TOI-270 d
by: Constantinou, Savvas, et al.
Published: (2025)
by: Constantinou, Savvas, et al.
Published: (2025)
Generalised Symmetries and Manifest Duality I: Flat Spacetime
by: Chakrabarti, Subhroneel, et al.
Published: (2025)
by: Chakrabarti, Subhroneel, et al.
Published: (2025)
Monopoles, Clarified
by: Aviral Aggarwal, et al.
Published: (2026)
by: Aviral Aggarwal, et al.
Published: (2026)
Monopoles, Clarified
by: Aggarwal, Aviral, et al.
Published: (2025)
by: Aggarwal, Aviral, et al.
Published: (2025)
Similar Items
-
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
by: Madhusudhan, Nishanth, et al.
Published: (2024) -
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
by: Hashemi, Masoud, et al.
Published: (2025) -
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
by: Tiwari, Aman, et al.
Published: (2024) -
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
by: Maheshwary, Rishabh, et al.
Published: (2024) -
DeepSRGM -- Sequence Classification and Ranking in Indian Classical Music with Deep Learning
by: Madhusudhan, Sathwik Tejaswi, et al.
Published: (2024)