Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
Fuente:
arXiv
Saved in:
| Main Authors: | Sühr, Tom, Dorner, Florian E., Salaudeen, Olawale, Kelava, Augustin, Samadi, Samira |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Challenging the Validity of Personality Tests for Large Language Models
by: Sühr, Tom, et al.
Published: (2023)
by: Sühr, Tom, et al.
Published: (2023)
The Transparency Paradox in Explainable AI: A Theory of Autonomy Depletion Through Cognitive Load
by: Margondai, Ancuta, et al.
Published: (2026)
by: Margondai, Ancuta, et al.
Published: (2026)
AI Psychometrics: Evaluating the Psychological Reasoning of Large Language Models with Psychometric Validities
by: Li, Yibai, et al.
Published: (2026)
by: Li, Yibai, et al.
Published: (2026)
Multimodal AI-based visualization of strategic leaders' emotional dynamics: a deep behavioral analysis of Trump's trade war discourse
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Function Alignment: A New Theory of Mind and Intelligence, Part I: Foundations
by: Xia, Gus G.
Published: (2025)
by: Xia, Gus G.
Published: (2025)
A Dynamic Model of Performative Human-ML Collaboration: Theory and Empirical Evidence
by: Sühr, Tom, et al.
Published: (2024)
by: Sühr, Tom, et al.
Published: (2024)
Generative AI on Wall Street -- Opportunities and Risk Controls
by: Shen, Jackie
Published: (2025)
by: Shen, Jackie
Published: (2025)
The Trap of Presumed Equivalence: Artificial General Intelligence Should Not Be Assessed on the Scale of Human Intelligence
by: Dolgikh, Serge
Published: (2024)
by: Dolgikh, Serge
Published: (2024)
Weighted simple games and the topology of simplicial complexes
by: Brooks, Anastasia, et al.
Published: (2022)
by: Brooks, Anastasia, et al.
Published: (2022)
Partial Information in a Mean-Variance Portfolio Selection Game
by: Huang, Yu-Jui, et al.
Published: (2023)
by: Huang, Yu-Jui, et al.
Published: (2023)
The Association of Transformer-based Sentiment Analysis with Symptom Distress and Deterioration in Routine Psychotherapy Care
by: Faust, Douglas K., et al.
Published: (2026)
by: Faust, Douglas K., et al.
Published: (2026)
Nonverbal Immediacy Analysis in Education: A Multimodal Computational Model
by: Petković, Uroš, et al.
Published: (2024)
by: Petković, Uroš, et al.
Published: (2024)
Spatiotemporal Heterogeneity of AI-Driven Traffic Flow Patterns and Land Use Interaction: A GeoAI-Based Analysis of Multimodal Urban Mobility
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
Unlocking NACE Classification Embeddings with OpenAI for Enhanced Analysis and Processing
by: Vidali, Andrea, et al.
Published: (2024)
by: Vidali, Andrea, et al.
Published: (2024)
Inclusive AI for Group Interactions: Predicting Gaze-Direction Behaviors in People with Intellectual and Developmental Disabilities
by: Huang, Giulia, et al.
Published: (2026)
by: Huang, Giulia, et al.
Published: (2026)
Can Nash inform capital requirements? Allocating systemic risk measures
by: Ararat, Çağın, et al.
Published: (2025)
by: Ararat, Çağın, et al.
Published: (2025)
Detecting Corporate AI-Washing via Cross-Modal Semantic Inconsistency Learning
by: Wen, Zhanjie, et al.
Published: (2026)
by: Wen, Zhanjie, et al.
Published: (2026)
On the Separability of Vector-Valued Risk Measures
by: Ararat, Çağın, et al.
Published: (2024)
by: Ararat, Çağın, et al.
Published: (2024)
Systemic values-at-risk and their sample-average approximations
by: AlAli, Wissam, et al.
Published: (2024)
by: AlAli, Wissam, et al.
Published: (2024)
Embeddability of joinpowers, and minimal rank of partial matrices
by: Skopenkov, A., et al.
Published: (2023)
by: Skopenkov, A., et al.
Published: (2023)
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
by: Salehi, Pegah, et al.
Published: (2025)
by: Salehi, Pegah, et al.
Published: (2025)
Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test
by: Adl, Rommin, et al.
Published: (2026)
by: Adl, Rommin, et al.
Published: (2026)
MEMOA: Massive Mixtures of Online Agents via Mean-Field Decentralized Nash Equilibria
by: Yang, Xuwei, et al.
Published: (2026)
by: Yang, Xuwei, et al.
Published: (2026)
Simulation of Non-Ordinary Consciousness
by: Saqr, Khalid M.
Published: (2025)
by: Saqr, Khalid M.
Published: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
by: Sáez, Arnau Igualde, et al.
Published: (2025)
by: Sáez, Arnau Igualde, et al.
Published: (2025)
Fourier Neural Network Approximation of Transition Densities in Finance
by: Du, Rong, et al.
Published: (2023)
by: Du, Rong, et al.
Published: (2023)
Dynamic Weight Optimization for Double Linear Policy: A Stochastic Model Predictive Control Approach
by: Hong, Tan Chin, et al.
Published: (2026)
by: Hong, Tan Chin, et al.
Published: (2026)
Detecting Cognitive Signatures in Typing Behavior for Non-Intrusive Authorship Verification
by: Condrey, David
Published: (2026)
by: Condrey, David
Published: (2026)
A Theory of Multilevel Interactive Equilibrium in NeuroAI
by: Chen, Zhe Sage, et al.
Published: (2026)
by: Chen, Zhe Sage, et al.
Published: (2026)
Predicting Femicide in Veracruz: A Fuzzy Logic Approach with the Expanded MFM-FEM-VER-CP-2024 Model
by: Medel-Ramírez, Carlos, et al.
Published: (2024)
by: Medel-Ramírez, Carlos, et al.
Published: (2024)
Uncertain Regulations, Definite Impacts: The Impact of the US Securities and Exchange Commission's Regulatory Interventions on Crypto Assets
by: Saggu, Aman, et al.
Published: (2024)
by: Saggu, Aman, et al.
Published: (2024)
$C_{pq}$-Injective Diagrams and a Combination Theorem for Minimal Models
by: Thandar, Soumyadip
Published: (2025)
by: Thandar, Soumyadip
Published: (2025)
Differential Beliefs in Financial Markets Under Information Constraints: A Modeling Perspective
by: Grigorian, Karen, et al.
Published: (2025)
by: Grigorian, Karen, et al.
Published: (2025)
Equivariant Intrinsic Formality
by: Santhanam, Rekha, et al.
Published: (2023)
by: Santhanam, Rekha, et al.
Published: (2023)
A ring structure on Tor
by: Carlson, Jeffrey D.
Published: (2023)
by: Carlson, Jeffrey D.
Published: (2023)
Kernel Based Maximum Entropy Inverse Reinforcement Learning for Mean-Field Games
by: Anahtarci, Berkay, et al.
Published: (2025)
by: Anahtarci, Berkay, et al.
Published: (2025)
Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
by: Skorski, Maciej
Published: (2026)
by: Skorski, Maciej
Published: (2026)
Separating Maximality Principles
by: Gappo, Takehiko, et al.
Published: (2025)
by: Gappo, Takehiko, et al.
Published: (2025)
Nonlocal Stochastic Optimal Control for Diffusion Processes: Existence, Maximum Principle and Financial Applications
by: Anita, Stefana-Lucia, et al.
Published: (2025)
by: Anita, Stefana-Lucia, et al.
Published: (2025)
Revenue-Sharing as Infrastructure: A Distributed Business Model for Generative AI Platforms
by: Mondjo, Ghislain Dorian Tchuente
Published: (2026)
by: Mondjo, Ghislain Dorian Tchuente
Published: (2026)
Similar Items
-
Challenging the Validity of Personality Tests for Large Language Models
by: Sühr, Tom, et al.
Published: (2023) -
The Transparency Paradox in Explainable AI: A Theory of Autonomy Depletion Through Cognitive Load
by: Margondai, Ancuta, et al.
Published: (2026) -
AI Psychometrics: Evaluating the Psychological Reasoning of Large Language Models with Psychometric Validities
by: Li, Yibai, et al.
Published: (2026) -
Multimodal AI-based visualization of strategic leaders' emotional dynamics: a deep behavioral analysis of Trump's trade war discourse
by: Meng, Wei
Published: (2025) -
Function Alignment: A New Theory of Mind and Intelligence, Part I: Foundations
by: Xia, Gus G.
Published: (2025)