Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pant, Aakash, Shah, Kavya, Agnihotri, Apoorv, Nikam, Sneha, Balraj, Prasaanth, Jain, Nakul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
von: Dong, Jia-Kai, et al.
Veröffentlicht: (2025)
von: Dong, Jia-Kai, et al.
Veröffentlicht: (2025)
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
von: Hooshyar, Danial, et al.
Veröffentlicht: (2023)
von: Hooshyar, Danial, et al.
Veröffentlicht: (2023)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
von: Pareschi, Remo
Veröffentlicht: (2025)
von: Pareschi, Remo
Veröffentlicht: (2025)
Development of Hybrid Artificial Intelligence Training on Real and Synthetic Data: Benchmark on Two Mixed Training Strategies
von: Wachter, Paul, et al.
Veröffentlicht: (2025)
von: Wachter, Paul, et al.
Veröffentlicht: (2025)
Empowering Decision Trees via Shape Function Branching
von: Upadhya, Nakul, et al.
Veröffentlicht: (2025)
von: Upadhya, Nakul, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models for Causal Modeling
von: Razouk, Houssam, et al.
Veröffentlicht: (2024)
von: Razouk, Houssam, et al.
Veröffentlicht: (2024)
Boundary Regression for Leitmotif Detection in Music Audio
von: Lee, Sihun, et al.
Veröffentlicht: (2025)
von: Lee, Sihun, et al.
Veröffentlicht: (2025)
A Hierarchical conv-LSTM and LLM Integrated Model for Holistic Stock Forecasting
von: Chakraborty, Arya, et al.
Veröffentlicht: (2024)
von: Chakraborty, Arya, et al.
Veröffentlicht: (2024)
Combination of Weak Learners eXplanations to Improve Random Forest eXplicability Robustness
von: Pala, Riccardo, et al.
Veröffentlicht: (2024)
von: Pala, Riccardo, et al.
Veröffentlicht: (2024)
Exploring the Technology Landscape through Topic Modeling, Expert Involvement, and Reinforcement Learning
von: Nazari, Ali, et al.
Veröffentlicht: (2025)
von: Nazari, Ali, et al.
Veröffentlicht: (2025)
Making High-Level AI Design Decisions Explicit Using a Binary Stream System-Designation Approach
von: Mossbridge, Julia
Veröffentlicht: (2024)
von: Mossbridge, Julia
Veröffentlicht: (2024)
KernelOracle: Predicting the Linux Scheduler's Next Move with Deep Learning
von: Kahu, Sampanna Yashwant
Veröffentlicht: (2025)
von: Kahu, Sampanna Yashwant
Veröffentlicht: (2025)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2025)
von: Nfissi, Alaa, et al.
Veröffentlicht: (2025)
A survey of air combat behavior modeling using machine learning
von: Gorton, Patrick Ribu, et al.
Veröffentlicht: (2024)
von: Gorton, Patrick Ribu, et al.
Veröffentlicht: (2024)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
von: Du, Bangde, et al.
Veröffentlicht: (2025)
von: Du, Bangde, et al.
Veröffentlicht: (2025)
ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
von: Kim, Sejin, et al.
Veröffentlicht: (2025)
von: Kim, Sejin, et al.
Veröffentlicht: (2025)
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
von: Deng, Minghang, et al.
Veröffentlicht: (2025)
von: Deng, Minghang, et al.
Veröffentlicht: (2025)
Right-to-Act: A Pre-Execution Non-Compensatory Decision Protocol for AI Systems
von: Lavi, Gadi
Veröffentlicht: (2026)
von: Lavi, Gadi
Veröffentlicht: (2026)
Gyan: An Explainable Neuro-Symbolic Language Model
von: Srinivasan, Venkat, et al.
Veröffentlicht: (2026)
von: Srinivasan, Venkat, et al.
Veröffentlicht: (2026)
Prescriptive Artificial Intelligence: A Formal Paradigm for Auditing Human Decisions Under Uncertainty
von: Passos, Pedro, et al.
Veröffentlicht: (2025)
von: Passos, Pedro, et al.
Veröffentlicht: (2025)
Discovering Differences in Strategic Behavior Between Humans and LLMs
von: Wang, Caroline, et al.
Veröffentlicht: (2026)
von: Wang, Caroline, et al.
Veröffentlicht: (2026)
Large Language Models (LLMs) for Requirements Engineering (RE): A Systematic Literature Review
von: Zadenoori, Mohammad Amin, et al.
Veröffentlicht: (2025)
von: Zadenoori, Mohammad Amin, et al.
Veröffentlicht: (2025)
Reasoning Beyond the Obvious: Evaluating Divergent and Convergent Thinking in LLMs for Financial Scenarios
von: Bok, Zhuang Qiang, et al.
Veröffentlicht: (2025)
von: Bok, Zhuang Qiang, et al.
Veröffentlicht: (2025)
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
von: Baranwal, Aaditya, et al.
Veröffentlicht: (2026)
von: Baranwal, Aaditya, et al.
Veröffentlicht: (2026)
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
von: Yun, Bhada, et al.
Veröffentlicht: (2026)
von: Yun, Bhada, et al.
Veröffentlicht: (2026)
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2024)
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2024)
Unlocking the Potential of Metaverse in Innovative and Immersive Digital Health
von: Ebrahimzadeh, Fatemeh, et al.
Veröffentlicht: (2024)
von: Ebrahimzadeh, Fatemeh, et al.
Veröffentlicht: (2024)
Kolmogorov Arnold Networks and Multi-Layer Perceptrons: A Paradigm Shift in Neural Modelling
von: Gaonkar, Aradhya, et al.
Veröffentlicht: (2026)
von: Gaonkar, Aradhya, et al.
Veröffentlicht: (2026)
Study Design and Demystification of Physics Informed Neural Networks for Power Flow Simulation
von: Leyli-abadi, Milad, et al.
Veröffentlicht: (2025)
von: Leyli-abadi, Milad, et al.
Veröffentlicht: (2025)
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
von: Teixeira, Tiago, et al.
Veröffentlicht: (2026)
von: Teixeira, Tiago, et al.
Veröffentlicht: (2026)
Carefully Structured Compression: Efficiently Managing StarCraft II Data
von: Ferenczi, Bryce, et al.
Veröffentlicht: (2024)
von: Ferenczi, Bryce, et al.
Veröffentlicht: (2024)
How much can change in a year? Revisiting Evaluation in Multi-Agent Reinforcement Learning
von: Singh, Siddarth, et al.
Veröffentlicht: (2023)
von: Singh, Siddarth, et al.
Veröffentlicht: (2023)
Efficiently Quantifying Individual Agent Importance in Cooperative MARL
von: Mahjoub, Omayma, et al.
Veröffentlicht: (2023)
von: Mahjoub, Omayma, et al.
Veröffentlicht: (2023)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2024)
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2024)
Automated Circuit Interpretation via Probe Prompting
von: Birardi, Giuseppe
Veröffentlicht: (2025)
von: Birardi, Giuseppe
Veröffentlicht: (2025)
Efficiently Scanning and Resampling Spatio-Temporal Tasks with Irregular Observations
von: Ferenczi, Bryce, et al.
Veröffentlicht: (2024)
von: Ferenczi, Bryce, et al.
Veröffentlicht: (2024)
PCA- and SVM-Grad-CAM for Convolutional Neural Networks: Closed-form Jacobian Expression
von: Omae, Yuto
Veröffentlicht: (2025)
von: Omae, Yuto
Veröffentlicht: (2025)
The Station: An Open-World Environment for AI-Driven Discovery
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
EduQate: Generating Adaptive Curricula through RMABs in Education Settings
von: Tio, Sidney, et al.
Veröffentlicht: (2024)
von: Tio, Sidney, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
von: Dong, Jia-Kai, et al.
Veröffentlicht: (2025) -
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
von: Hooshyar, Danial, et al.
Veröffentlicht: (2023) -
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023) -
Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
von: Pareschi, Remo
Veröffentlicht: (2025) -
Development of Hybrid Artificial Intelligence Training on Real and Synthetic Data: Benchmark on Two Mixed Training Strategies
von: Wachter, Paul, et al.
Veröffentlicht: (2025)