Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
Fuente:
arXiv
Saved in:
| Main Authors: | Jallad, Khloud AL, Ghneim, Nada, Rebdawi, Ghaida |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
ArEEG_Chars: Dataset for Envisioned Speech Recognition using EEG for Arabic Characters
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
Arabic Little STT: Arabic Children Speech Recognition Dataset
by: Alkadri, Mouhand, et al.
Published: (2025)
by: Alkadri, Mouhand, et al.
Published: (2025)
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
by: Khalil, Mohammad Amer, et al.
Published: (2026)
by: Khalil, Mohammad Amer, et al.
Published: (2026)
Benchmarking Gender and Political Bias in Large Language Models
by: Yang, Jinrui, et al.
Published: (2025)
by: Yang, Jinrui, et al.
Published: (2025)
Voting-based Multimodal Automatic Deception Detection
by: Touma, Lana, et al.
Published: (2023)
by: Touma, Lana, et al.
Published: (2023)
Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI
by: Saha, Agnik, et al.
Published: (2025)
by: Saha, Agnik, et al.
Published: (2025)
Benchmarking Agentic Workflow Generation
by: Qiao, Shuofei, et al.
Published: (2024)
by: Qiao, Shuofei, et al.
Published: (2024)
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
by: Garbacea, Cristina, et al.
Published: (2026)
by: Garbacea, Cristina, et al.
Published: (2026)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
by: Xu, Jia, et al.
Published: (2025)
by: Xu, Jia, et al.
Published: (2025)
Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
by: Zhang, Ziyuan, et al.
Published: (2025)
by: Zhang, Ziyuan, et al.
Published: (2025)
Survey of User Interface Design and Interaction Techniques in Generative AI Applications
by: Luera, Reuben, et al.
Published: (2024)
by: Luera, Reuben, et al.
Published: (2024)
Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
by: Xu, Haoming, et al.
Published: (2026)
by: Xu, Haoming, et al.
Published: (2026)
Benchmark It Yourself (BIY): Preparing a Dataset and Benchmarking AI Models for Scatterplot-Related Tasks
by: Palmeiro, João, et al.
Published: (2025)
by: Palmeiro, João, et al.
Published: (2025)
Deterministic AI Agent Personality Expression through Standard Psychological Diagnostics
by: Kruijssen, J. M. Diederik, et al.
Published: (2025)
by: Kruijssen, J. M. Diederik, et al.
Published: (2025)
A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
by: Machlovi, Naseem, et al.
Published: (2025)
by: Machlovi, Naseem, et al.
Published: (2025)
Benchmarking Mobile Device Control Agents across Diverse Configurations
by: Lee, Juyong, et al.
Published: (2024)
by: Lee, Juyong, et al.
Published: (2024)
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
RuleAlign: Making Large Language Models Better Physicians with Diagnostic Rule Alignment
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
by: Zou, Henry Peng, et al.
Published: (2025)
by: Zou, Henry Peng, et al.
Published: (2025)
From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making
by: Lee, Min Hun
Published: (2026)
by: Lee, Min Hun
Published: (2026)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
by: Sinacola, Enzo, et al.
Published: (2025)
by: Sinacola, Enzo, et al.
Published: (2025)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
IDRBench: Interactive Deep Research Benchmark
by: Feng, Yingchaojie, et al.
Published: (2026)
by: Feng, Yingchaojie, et al.
Published: (2026)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Benchmarking System Dynamics AI Assistants: Cloud Versus Local LLMs on CLD Extraction and Discussion
by: Leitch, Terry
Published: (2026)
by: Leitch, Terry
Published: (2026)
Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
by: Silva, Inês Oliveira e, et al.
Published: (2026)
by: Silva, Inês Oliveira e, et al.
Published: (2026)
Cognitively-Inspired Episodic Memory Architectures for Accurate and Efficient Character AI
by: Gonzalez, Rafael Arias, et al.
Published: (2025)
by: Gonzalez, Rafael Arias, et al.
Published: (2025)
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
Agent Laboratory: Using LLM Agents as Research Assistants
by: Schmidgall, Samuel, et al.
Published: (2025)
by: Schmidgall, Samuel, et al.
Published: (2025)
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
by: Ayonrinde, Kola, et al.
Published: (2025)
by: Ayonrinde, Kola, et al.
Published: (2025)
Value Profiles for Encoding Human Variation
by: Sorensen, Taylor, et al.
Published: (2025)
by: Sorensen, Taylor, et al.
Published: (2025)
OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
by: Verma, Sahil, et al.
Published: (2025)
by: Verma, Sahil, et al.
Published: (2025)
Feedback-Aware Monte Carlo Tree Search for Efficient Information Seeking in Goal-Oriented Conversations
by: Chopra, Harshita, et al.
Published: (2025)
by: Chopra, Harshita, et al.
Published: (2025)
Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
by: Liu, Jonathan, et al.
Published: (2025)
by: Liu, Jonathan, et al.
Published: (2025)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
by: Soni, Nikita, et al.
Published: (2025)
by: Soni, Nikita, et al.
Published: (2025)
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
by: He, Zeyu, et al.
Published: (2025)
by: He, Zeyu, et al.
Published: (2025)
Similar Items
-
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
by: Darwish, Hazem, et al.
Published: (2024) -
ArEEG_Chars: Dataset for Envisioned Speech Recognition using EEG for Arabic Characters
by: Darwish, Hazem, et al.
Published: (2024) -
Arabic Little STT: Arabic Children Speech Recognition Dataset
by: Alkadri, Mouhand, et al.
Published: (2025) -
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts
by: Mohammadi, Seyedali, et al.
Published: (2025) -
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
by: Khalil, Mohammad Amer, et al.
Published: (2026)