On Meta-Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongxiao, Wang, Chenxi, Fan, Fanda, Wang, Zihan, Gao, Wanling, Wang, Lei, Zhan, Jianfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Imputation Matters: A Deeper Look into an Overlooked Step in Longitudinal Health and Behavior Sensing Research
by: Choube, Akshat, et al.
Published: (2024)
by: Choube, Akshat, et al.
Published: (2024)
Perils of Label Indeterminacy: A Case Study on Prediction of Neurological Recovery After Cardiac Arrest
by: Schoeffer, Jakob, et al.
Published: (2025)
by: Schoeffer, Jakob, et al.
Published: (2025)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023)
by: Kıcıman, Emre, et al.
Published: (2023)
VisMoDAl: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models
by: Wang, Huanchen, et al.
Published: (2025)
by: Wang, Huanchen, et al.
Published: (2025)
Visual Analysis of Multi-outcome Causal Graphs
by: Fan, Mengjie, et al.
Published: (2024)
by: Fan, Mengjie, et al.
Published: (2024)
Methodology and Real-World Applications of Dynamic Uncertain Causality Graph for Clinical Diagnosis with Explainability and Invariance
by: Zhang, Zhan, et al.
Published: (2024)
by: Zhang, Zhan, et al.
Published: (2024)
Generative AI for Visualization: State of the Art and Future Directions
by: Ye, Yilin, et al.
Published: (2024)
by: Ye, Yilin, et al.
Published: (2024)
Optimizing Feature Extraction for On-device Model Inference with User Behavior Sequences
by: Gong, Chen, et al.
Published: (2026)
by: Gong, Chen, et al.
Published: (2026)
RAICL: Retrieval-Augmented In-Context Learning for Vision-Language-Model Based EEG Seizure Detection
by: Li, Siyang, et al.
Published: (2026)
by: Li, Siyang, et al.
Published: (2026)
Spatial Distillation based Distribution Alignment (SDDA) for Cross-Headset EEG Classification
by: Liu, Dingkun, et al.
Published: (2025)
by: Liu, Dingkun, et al.
Published: (2025)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
by: Zeng, Qiuhai, et al.
Published: (2025)
by: Zeng, Qiuhai, et al.
Published: (2025)
Agent AI: Surveying the Horizons of Multimodal Interaction
by: Durante, Zane, et al.
Published: (2024)
by: Durante, Zane, et al.
Published: (2024)
SPECTRE: Spectral Pre-training Embeddings with Cylindrical Temporal Rotary Position Encoding for Fine-Grained sEMG-Based Movement Decoding
by: Weng, Zihan, et al.
Published: (2025)
by: Weng, Zihan, et al.
Published: (2025)
From Complexity to Parsimony: Integrating Latent Class Analysis to Uncover Multimodal Learning Patterns in Collaborative Learning
by: Yan, Lixiang, et al.
Published: (2024)
by: Yan, Lixiang, et al.
Published: (2024)
Loop Polarity Analysis to Avoid Underspecification in Deep Learning
by: Martin, Jr., Donald, et al.
Published: (2023)
by: Martin, Jr., Donald, et al.
Published: (2023)
MetaExplainer: A Framework to Generate Multi-Type User-Centered Explanations for AI Systems
by: Chari, Shruthi, et al.
Published: (2025)
by: Chari, Shruthi, et al.
Published: (2025)
Evaluating Deep Networks for Detecting User Familiarity with VR from Hand Interactions
by: Li, Mingjun, et al.
Published: (2024)
by: Li, Mingjun, et al.
Published: (2024)
Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human Alignment
by: Zhang, Chen, et al.
Published: (2024)
by: Zhang, Chen, et al.
Published: (2024)
Domain-Adversarial Anatomical Graph Networks for Cross-User Human Activity Recognition
by: Ye, Xiaozhou, et al.
Published: (2025)
by: Ye, Xiaozhou, et al.
Published: (2025)
Reinforcement Learning Driven Generalizable Feature Representation for Cross-User Activity Recognition
by: Ye, Xiaozhou, et al.
Published: (2025)
by: Ye, Xiaozhou, et al.
Published: (2025)
Alignment-Based Adversarial Training (ABAT) for Improving the Robustness and Accuracy of EEG-Based BCIs
by: Chen, Xiaoqing, et al.
Published: (2024)
by: Chen, Xiaoqing, et al.
Published: (2024)
Adversarial Domain Adaptation for Cross-user Activity Recognition Using Diffusion-based Noise-centred Learning
by: Ye, Xiaozhou, et al.
Published: (2024)
by: Ye, Xiaozhou, et al.
Published: (2024)
Why is "Problems" Predictive of Positive Sentiment? A Case Study of Explaining Unintuitive Features in Sentiment Classification
by: Qu, Jiaming, et al.
Published: (2024)
by: Qu, Jiaming, et al.
Published: (2024)
Human-Computer Interaction and Human-AI Collaboration in Advanced Air Mobility: A Comprehensive Review
by: Sagirli, Fatma Yamac, et al.
Published: (2024)
by: Sagirli, Fatma Yamac, et al.
Published: (2024)
KAIROS: Unified Training for Universal Non-Autoregressive Time Series Forecasting
by: Ding, Kuiye, et al.
Published: (2025)
by: Ding, Kuiye, et al.
Published: (2025)
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
by: Cook, Jonathan, et al.
Published: (2024)
by: Cook, Jonathan, et al.
Published: (2024)
Collaborative Intelligence in Sequential Experiments: A Human-in-the-Loop Framework for Drug Discovery
by: He, Jinghai, et al.
Published: (2024)
by: He, Jinghai, et al.
Published: (2024)
How Users Who are Blind or Low Vision Play Mobile Games: Perceptions, Challenges, and Strategies
by: Ran, Zihe, et al.
Published: (2025)
by: Ran, Zihe, et al.
Published: (2025)
Reinforcement Learning Measurement Model
by: Xu, Wenqian, et al.
Published: (2026)
by: Xu, Wenqian, et al.
Published: (2026)
Domain-Grounded Evaluation of LLMs in International Student Knowledge
by: Daitx, Claudinei, et al.
Published: (2025)
by: Daitx, Claudinei, et al.
Published: (2025)
Steering Robots with Inference-Time Interactions
by: Wang, Yanwei
Published: (2025)
by: Wang, Yanwei
Published: (2025)
Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software
by: Khurana, Anjali, et al.
Published: (2025)
by: Khurana, Anjali, et al.
Published: (2025)
Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors
by: Chen, Wenqiang, et al.
Published: (2024)
by: Chen, Wenqiang, et al.
Published: (2024)
Off-Policy Selection for Initiating Human-Centric Experimental Design
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
by: Baidya, Avinash, et al.
Published: (2025)
by: Baidya, Avinash, et al.
Published: (2025)
Copilot-Assisted Second-Thought Framework for Brain-to-Robot Hand Motion Decoding
by: Li, Yizhe, et al.
Published: (2026)
by: Li, Yizhe, et al.
Published: (2026)
A Systematic Review of User-Centred Evaluation of Explainable AI in Healthcare
by: Donoso-Guzmán, Ivania, et al.
Published: (2025)
by: Donoso-Guzmán, Ivania, et al.
Published: (2025)
Evaluation of Human-Understandability of Global Model Explanations using Decision Tree
by: Sivaprasad, Adarsa, et al.
Published: (2023)
by: Sivaprasad, Adarsa, et al.
Published: (2023)
When Generative Artificial Intelligence meets Extended Reality: A Systematic Review
by: Ning, Xinyu, et al.
Published: (2025)
by: Ning, Xinyu, et al.
Published: (2025)
Similar Items
-
Imputation Matters: A Deeper Look into an Overlooked Step in Longitudinal Health and Behavior Sensing Research
by: Choube, Akshat, et al.
Published: (2024) -
Perils of Label Indeterminacy: A Case Study on Prediction of Neurological Recovery After Cardiac Arrest
by: Schoeffer, Jakob, et al.
Published: (2025) -
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023) -
VisMoDAl: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models
by: Wang, Huanchen, et al.
Published: (2025) -
Visual Analysis of Multi-outcome Causal Graphs
by: Fan, Mengjie, et al.
Published: (2024)