Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
Fuente:
arXiv
Saved in:
| Main Authors: | Pant, Aakash, Shah, Kavya, Agnihotri, Apoorv, Nikam, Sneha, Balraj, Prasaanth, Jain, Nakul |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
by: Dong, Jia-Kai, et al.
Published: (2025)
by: Dong, Jia-Kai, et al.
Published: (2025)
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
by: Hooshyar, Danial, et al.
Published: (2023)
by: Hooshyar, Danial, et al.
Published: (2023)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
by: Wu, Dekun, et al.
Published: (2023)
by: Wu, Dekun, et al.
Published: (2023)
Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
by: Pareschi, Remo
Published: (2025)
by: Pareschi, Remo
Published: (2025)
Development of Hybrid Artificial Intelligence Training on Real and Synthetic Data: Benchmark on Two Mixed Training Strategies
by: Wachter, Paul, et al.
Published: (2025)
by: Wachter, Paul, et al.
Published: (2025)
Empowering Decision Trees via Shape Function Branching
by: Upadhya, Nakul, et al.
Published: (2025)
by: Upadhya, Nakul, et al.
Published: (2025)
Boundary Regression for Leitmotif Detection in Music Audio
by: Lee, Sihun, et al.
Published: (2025)
by: Lee, Sihun, et al.
Published: (2025)
A Hierarchical conv-LSTM and LLM Integrated Model for Holistic Stock Forecasting
by: Chakraborty, Arya, et al.
Published: (2024)
by: Chakraborty, Arya, et al.
Published: (2024)
Combination of Weak Learners eXplanations to Improve Random Forest eXplicability Robustness
by: Pala, Riccardo, et al.
Published: (2024)
by: Pala, Riccardo, et al.
Published: (2024)
Evaluating Large Language Models for Causal Modeling
by: Razouk, Houssam, et al.
Published: (2024)
by: Razouk, Houssam, et al.
Published: (2024)
Exploring the Technology Landscape through Topic Modeling, Expert Involvement, and Reinforcement Learning
by: Nazari, Ali, et al.
Published: (2025)
by: Nazari, Ali, et al.
Published: (2025)
Making High-Level AI Design Decisions Explicit Using a Binary Stream System-Designation Approach
by: Mossbridge, Julia
Published: (2024)
by: Mossbridge, Julia
Published: (2024)
KernelOracle: Predicting the Linux Scheduler's Next Move with Deep Learning
by: Kahu, Sampanna Yashwant
Published: (2025)
by: Kahu, Sampanna Yashwant
Published: (2025)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2025)
by: Nfissi, Alaa, et al.
Published: (2025)
A survey of air combat behavior modeling using machine learning
by: Gorton, Patrick Ribu, et al.
Published: (2024)
by: Gorton, Patrick Ribu, et al.
Published: (2024)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
by: Kim, Sejin, et al.
Published: (2025)
by: Kim, Sejin, et al.
Published: (2025)
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
by: Deng, Minghang, et al.
Published: (2025)
by: Deng, Minghang, et al.
Published: (2025)
Right-to-Act: A Pre-Execution Non-Compensatory Decision Protocol for AI Systems
by: Lavi, Gadi
Published: (2026)
by: Lavi, Gadi
Published: (2026)
Large Language Models (LLMs) for Requirements Engineering (RE): A Systematic Literature Review
by: Zadenoori, Mohammad Amin, et al.
Published: (2025)
by: Zadenoori, Mohammad Amin, et al.
Published: (2025)
Gyan: An Explainable Neuro-Symbolic Language Model
by: Srinivasan, Venkat, et al.
Published: (2026)
by: Srinivasan, Venkat, et al.
Published: (2026)
Prescriptive Artificial Intelligence: A Formal Paradigm for Auditing Human Decisions Under Uncertainty
by: Passos, Pedro, et al.
Published: (2025)
by: Passos, Pedro, et al.
Published: (2025)
Discovering Differences in Strategic Behavior Between Humans and LLMs
by: Wang, Caroline, et al.
Published: (2026)
by: Wang, Caroline, et al.
Published: (2026)
Reasoning Beyond the Obvious: Evaluating Divergent and Convergent Thinking in LLMs for Financial Scenarios
by: Bok, Zhuang Qiang, et al.
Published: (2025)
by: Bok, Zhuang Qiang, et al.
Published: (2025)
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
by: Baranwal, Aaditya, et al.
Published: (2026)
by: Baranwal, Aaditya, et al.
Published: (2026)
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
by: Yun, Bhada, et al.
Published: (2026)
by: Yun, Bhada, et al.
Published: (2026)
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
by: Benfeghoul, Martin, et al.
Published: (2024)
by: Benfeghoul, Martin, et al.
Published: (2024)
Unlocking the Potential of Metaverse in Innovative and Immersive Digital Health
by: Ebrahimzadeh, Fatemeh, et al.
Published: (2024)
by: Ebrahimzadeh, Fatemeh, et al.
Published: (2024)
Kolmogorov Arnold Networks and Multi-Layer Perceptrons: A Paradigm Shift in Neural Modelling
by: Gaonkar, Aradhya, et al.
Published: (2026)
by: Gaonkar, Aradhya, et al.
Published: (2026)
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
by: Teixeira, Tiago, et al.
Published: (2026)
by: Teixeira, Tiago, et al.
Published: (2026)
Study Design and Demystification of Physics Informed Neural Networks for Power Flow Simulation
by: Leyli-abadi, Milad, et al.
Published: (2025)
by: Leyli-abadi, Milad, et al.
Published: (2025)
Carefully Structured Compression: Efficiently Managing StarCraft II Data
by: Ferenczi, Bryce, et al.
Published: (2024)
by: Ferenczi, Bryce, et al.
Published: (2024)
How much can change in a year? Revisiting Evaluation in Multi-Agent Reinforcement Learning
by: Singh, Siddarth, et al.
Published: (2023)
by: Singh, Siddarth, et al.
Published: (2023)
Efficiently Quantifying Individual Agent Importance in Cooperative MARL
by: Mahjoub, Omayma, et al.
Published: (2023)
by: Mahjoub, Omayma, et al.
Published: (2023)
Efficiently Scanning and Resampling Spatio-Temporal Tasks with Irregular Observations
by: Ferenczi, Bryce, et al.
Published: (2024)
by: Ferenczi, Bryce, et al.
Published: (2024)
PCA- and SVM-Grad-CAM for Convolutional Neural Networks: Closed-form Jacobian Expression
by: Omae, Yuto
Published: (2025)
by: Omae, Yuto
Published: (2025)
The Station: An Open-World Environment for AI-Driven Discovery
by: Chung, Stephen, et al.
Published: (2025)
by: Chung, Stephen, et al.
Published: (2025)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
Automated Circuit Interpretation via Probe Prompting
by: Birardi, Giuseppe
Published: (2025)
by: Birardi, Giuseppe
Published: (2025)
EduQate: Generating Adaptive Curricula through RMABs in Education Settings
by: Tio, Sidney, et al.
Published: (2024)
by: Tio, Sidney, et al.
Published: (2024)
Similar Items
-
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
by: Dong, Jia-Kai, et al.
Published: (2025) -
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
by: Hooshyar, Danial, et al.
Published: (2023) -
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
by: Wu, Dekun, et al.
Published: (2023) -
Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
by: Pareschi, Remo
Published: (2025) -
Development of Hybrid Artificial Intelligence Training on Real and Synthetic Data: Benchmark on Two Mixed Training Strategies
by: Wachter, Paul, et al.
Published: (2025)