AI Evaluation Should Require Standardized Item-Level Data Releases
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Han, Zhang, Susu, Zhu, Dongyao, Bai, Yuzhuo, Truong, Sang T., Yi, Xiaoyuan, Koyejo, Sanmi, Xie, Xing, Xiao, Ziang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
von: Vo, Truong, et al.
Veröffentlicht: (2025)
von: Vo, Truong, et al.
Veröffentlicht: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
von: Jiang, Han, et al.
Veröffentlicht: (2025)
von: Jiang, Han, et al.
Veröffentlicht: (2025)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
von: Jiang, Han, et al.
Veröffentlicht: (2025)
von: Jiang, Han, et al.
Veröffentlicht: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
von: Tang, Zeyu, et al.
Veröffentlicht: (2026)
von: Tang, Zeyu, et al.
Veröffentlicht: (2026)
Reliable and Efficient Amortized Model-based Evaluation
von: Truong, Sang, et al.
Veröffentlicht: (2025)
von: Truong, Sang, et al.
Veröffentlicht: (2025)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
von: Robertson, Zachary, et al.
Veröffentlicht: (2025)
von: Robertson, Zachary, et al.
Veröffentlicht: (2025)
Why Do Safety Guardrails Degrade Across Languages?
von: Zhang, Max, et al.
Veröffentlicht: (2026)
von: Zhang, Max, et al.
Veröffentlicht: (2026)
Exploring Distance Query Processing in Edge Computing Environments
von: Zhang, Xiubo, et al.
Veröffentlicht: (2024)
von: Zhang, Xiubo, et al.
Veröffentlicht: (2024)
Should I Hide My Duck in the Lake?
von: Dann, Jonas, et al.
Veröffentlicht: (2026)
von: Dann, Jonas, et al.
Veröffentlicht: (2026)
Crypto-Assisted Graph Degree Sequence Release under Local Differential Privacy
von: Zhang, Xiaojian, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaojian, et al.
Veröffentlicht: (2025)
UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data
von: Weng, Han, et al.
Veröffentlicht: (2025)
von: Weng, Han, et al.
Veröffentlicht: (2025)
When Focus Enhances Utility: Target Range LDP Frequency Estimation and Unknown Item Discovery
von: Jiang, Bo, et al.
Veröffentlicht: (2024)
von: Jiang, Bo, et al.
Veröffentlicht: (2024)
How to Mine Potentially Popular Items? A Reverse MIPS-based Approach
von: Amagata, Daichi, et al.
Veröffentlicht: (2025)
von: Amagata, Daichi, et al.
Veröffentlicht: (2025)
ThreatKG: An AI-Powered System for Automated Open-Source Cyber Threat Intelligence Gathering and Management
von: Gao, Peng, et al.
Veröffentlicht: (2022)
von: Gao, Peng, et al.
Veröffentlicht: (2022)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
von: Chen, Edward, et al.
Veröffentlicht: (2025)
von: Chen, Edward, et al.
Veröffentlicht: (2025)
Hidden Sketch: A Space-Efficient Reversible Sketch for Tracking Frequent Items in Data Streams
von: Xu, Zicang, et al.
Veröffentlicht: (2025)
von: Xu, Zicang, et al.
Veröffentlicht: (2025)
It's Time to Standardize RDF Messages
von: Colpaert, Pieter, et al.
Veröffentlicht: (2026)
von: Colpaert, Pieter, et al.
Veröffentlicht: (2026)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
von: Patel, Fagun, et al.
Veröffentlicht: (2025)
von: Patel, Fagun, et al.
Veröffentlicht: (2025)
MorphingDB: A Task-Centric AI-Native DBMS for Model Management and Inference
von: Sai, Wu, et al.
Veröffentlicht: (2025)
von: Sai, Wu, et al.
Veröffentlicht: (2025)
My Ontologist: Evaluating BFO-Based AI for Definition Support
von: Benson, Carter, et al.
Veröffentlicht: (2024)
von: Benson, Carter, et al.
Veröffentlicht: (2024)
Towards a Standard for JSON Document Databases
von: Botoeva, Elena, et al.
Veröffentlicht: (2025)
von: Botoeva, Elena, et al.
Veröffentlicht: (2025)
Towards Effective Orchestration of AI x DB Workloads
von: Xing, Naili, et al.
Veröffentlicht: (2026)
von: Xing, Naili, et al.
Veröffentlicht: (2026)
FuncEvalGMN: Evaluating Functional Correctness of SQL via Graph Matching Network
von: Zhan, Yi, et al.
Veröffentlicht: (2024)
von: Zhan, Yi, et al.
Veröffentlicht: (2024)
Reservoir Sampling over Joins
von: Dai, Binyang, et al.
Veröffentlicht: (2024)
von: Dai, Binyang, et al.
Veröffentlicht: (2024)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
von: Jiang, Han, et al.
Veröffentlicht: (2024)
von: Jiang, Han, et al.
Veröffentlicht: (2024)
Privacy and Accuracy-Aware AI/ML Model Deduplication
von: Guan, Hong, et al.
Veröffentlicht: (2025)
von: Guan, Hong, et al.
Veröffentlicht: (2025)
Real Life Is Uncertain. Consensus Should Be Too!
von: Frank, Reginald, et al.
Veröffentlicht: (2026)
von: Frank, Reginald, et al.
Veröffentlicht: (2026)
SCOPE: Cost-Efficient Model Selection for Compound AI Systems under Quality Constraints
von: Huang, Yiqian, et al.
Veröffentlicht: (2026)
von: Huang, Yiqian, et al.
Veröffentlicht: (2026)
Machine Learning Practitioners' Views on Data Quality in Light of EU Regulatory Requirements: A European Online Survey
von: Wang, Yichun, et al.
Veröffentlicht: (2026)
von: Wang, Yichun, et al.
Veröffentlicht: (2026)
Extended Event Log: Towards a Unified Standard for Process Mining
von: Suleiman, Ali, et al.
Veröffentlicht: (2024)
von: Suleiman, Ali, et al.
Veröffentlicht: (2024)
Finding Non-Redundant Simpson's Paradox from Multidimensional Data
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
von: Wu, Junfeng, et al.
Veröffentlicht: (2025)
von: Wu, Junfeng, et al.
Veröffentlicht: (2025)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
von: Truong, Sang T., et al.
Veröffentlicht: (2024)
von: Truong, Sang T., et al.
Veröffentlicht: (2024)
Preventing the Popular Item Embedding Based Attack in Federated Recommendations
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
CACTUSDB: Unlock Co-Optimization Opportunities for SQL and AI/ML Inferences
von: Zhou, Lixi, et al.
Veröffentlicht: (2026)
von: Zhou, Lixi, et al.
Veröffentlicht: (2026)
Efficient Query Rewrite Rule Discovery via Standardized Enumeration and Learning-to-Rank(extend)
von: Zhang, Yuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuan, et al.
Veröffentlicht: (2026)
Beyond Standard Datacubes: Extracting Features from Irregular and Branching Earth System Data
von: Leuridan, Mathilde, et al.
Veröffentlicht: (2026)
von: Leuridan, Mathilde, et al.
Veröffentlicht: (2026)
Building an OceanBase-based Distributed Nearly Real-time Analytical Processing Database System
von: Xu, Quanqing, et al.
Veröffentlicht: (2026)
von: Xu, Quanqing, et al.
Veröffentlicht: (2026)
Requirements for a User-Friendly OPAC.
von: Fokker, Dirk W.
Veröffentlicht: (1989)
von: Fokker, Dirk W.
Veröffentlicht: (1989)
AQETuner: Reliable Query-level Configuration Tuning for Analytical Query Engines
von: Chen, Lixiang, et al.
Veröffentlicht: (2025)
von: Chen, Lixiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
von: Vo, Truong, et al.
Veröffentlicht: (2025) -
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
von: Jiang, Han, et al.
Veröffentlicht: (2025) -
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
von: Jiang, Han, et al.
Veröffentlicht: (2025) -
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
von: Tang, Zeyu, et al.
Veröffentlicht: (2026) -
Reliable and Efficient Amortized Model-based Evaluation
von: Truong, Sang, et al.
Veröffentlicht: (2025)