Benchmark It Yourself (BIY): Preparing a Dataset and Benchmarking AI Models for Scatterplot-Related Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Palmeiro, João, Duarte, Diogo, Costa, Rita, Bizarro, Pedro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
por: Silva, Inês Oliveira e, et al.
Publicado: (2026)
por: Silva, Inês Oliveira e, et al.
Publicado: (2026)
"Show Me What's Wrong!": Combining Charts and Text to Guide Data Analysis
por: Feliciano, Beatriz, et al.
Publicado: (2024)
por: Feliciano, Beatriz, et al.
Publicado: (2024)
DiConStruct: Causal Concept-based Explanations through Black-Box Distillation
por: Moreira, Ricardo, et al.
Publicado: (2024)
por: Moreira, Ricardo, et al.
Publicado: (2024)
From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making
por: Lee, Min Hun
Publicado: (2026)
por: Lee, Min Hun
Publicado: (2026)
What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
por: Ibrahim, Adham, et al.
Publicado: (2024)
por: Ibrahim, Adham, et al.
Publicado: (2024)
Benchmarking System Dynamics AI Assistants: Cloud Versus Local LLMs on CLD Extraction and Discussion
por: Leitch, Terry
Publicado: (2026)
por: Leitch, Terry
Publicado: (2026)
Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
por: Ding, Shi, et al.
Publicado: (2025)
por: Ding, Shi, et al.
Publicado: (2025)
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
por: Wang, Xin, et al.
Publicado: (2025)
por: Wang, Xin, et al.
Publicado: (2025)
Benchmarking Mobile Device Control Agents across Diverse Configurations
por: Lee, Juyong, et al.
Publicado: (2024)
por: Lee, Juyong, et al.
Publicado: (2024)
Toward Human-AI Complementarity Across Diverse Tasks
por: Xu, Yuzheng, et al.
Publicado: (2026)
por: Xu, Yuzheng, et al.
Publicado: (2026)
Benchmarking Gender and Political Bias in Large Language Models
por: Yang, Jinrui, et al.
Publicado: (2025)
por: Yang, Jinrui, et al.
Publicado: (2025)
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
por: Jallad, Khloud AL, et al.
Publicado: (2025)
por: Jallad, Khloud AL, et al.
Publicado: (2025)
TutoAI: A Cross-domain Framework for AI-assisted Mixed-media Tutorial Creation on Physical Tasks
por: Chen, Yuexi, et al.
Publicado: (2024)
por: Chen, Yuexi, et al.
Publicado: (2024)
Towards Uncertainty Aware Task Delegation and Human-AI Collaborative Decision-Making
por: Lee, Min Hun, et al.
Publicado: (2025)
por: Lee, Min Hun, et al.
Publicado: (2025)
ABScribe: Rapid Exploration & Organization of Multiple Writing Variations in Human-AI Co-Writing Tasks using Large Language Models
por: Reza, Mohi, et al.
Publicado: (2023)
por: Reza, Mohi, et al.
Publicado: (2023)
Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments
por: Roman, Anthony Cintron, et al.
Publicado: (2023)
por: Roman, Anthony Cintron, et al.
Publicado: (2023)
ConvoLearn: A Learning Sciences Grounded Dataset for Fine-Tuning Dialogic AI Tutors
por: Sharma, Mayank, et al.
Publicado: (2026)
por: Sharma, Mayank, et al.
Publicado: (2026)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
por: Xu, Jia, et al.
Publicado: (2025)
por: Xu, Jia, et al.
Publicado: (2025)
Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
por: Yuan, Yifu, et al.
Publicado: (2024)
por: Yuan, Yifu, et al.
Publicado: (2024)
GQVis: A Dataset of Genomics Data Questions and Visualizations for Generative AI
por: Walters, Skylar Sargent, et al.
Publicado: (2025)
por: Walters, Skylar Sargent, et al.
Publicado: (2025)
Noise Correction on Subjective Datasets
por: Jinadu, Uthman, et al.
Publicado: (2023)
por: Jinadu, Uthman, et al.
Publicado: (2023)
CliMB: An AI-enabled Partner for Clinical Predictive Modeling
por: Saveliev, Evgeny, et al.
Publicado: (2024)
por: Saveliev, Evgeny, et al.
Publicado: (2024)
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems
por: Rajagopal, Shreya, et al.
Publicado: (2024)
por: Rajagopal, Shreya, et al.
Publicado: (2024)
The Model Mastery Lifecycle: A Framework for Designing Human-AI Interaction
por: Chignell, Mark, et al.
Publicado: (2024)
por: Chignell, Mark, et al.
Publicado: (2024)
Jigsaw: Supporting Designers to Prototype Multimodal Applications by Chaining AI Foundation Models
por: Lin, David Chuan-En, et al.
Publicado: (2023)
por: Lin, David Chuan-En, et al.
Publicado: (2023)
RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview
por: Lee, Min Hun, et al.
Publicado: (2026)
por: Lee, Min Hun, et al.
Publicado: (2026)
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts
por: Mohammadi, Seyedali, et al.
Publicado: (2025)
por: Mohammadi, Seyedali, et al.
Publicado: (2025)
The case for delegated AI autonomy for Human AI teaming in healthcare
por: Jia, Yan, et al.
Publicado: (2025)
por: Jia, Yan, et al.
Publicado: (2025)
Explaining AI Without Code: A User Study on Explainable AI
por: Abarca, Natalia, et al.
Publicado: (2025)
por: Abarca, Natalia, et al.
Publicado: (2025)
Benchmarking Agentic Workflow Generation
por: Qiao, Shuofei, et al.
Publicado: (2024)
por: Qiao, Shuofei, et al.
Publicado: (2024)
On the Utility of Accounting for Human Beliefs about AI Intention in Human-AI Collaboration
por: Yu, Guanghui, et al.
Publicado: (2024)
por: Yu, Guanghui, et al.
Publicado: (2024)
A Multi-Component AI Framework for Computational Psychology: From Robust Predictive Modeling to Deployed Generative Dialogue
por: Pareek, Anant
Publicado: (2025)
por: Pareek, Anant
Publicado: (2025)
Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation
por: Rondina, Marco, et al.
Publicado: (2025)
por: Rondina, Marco, et al.
Publicado: (2025)
Benchmarking Neural Decoding Backbones towards Enhanced On-edge iBCI Applications
por: Zhou, Zhou, et al.
Publicado: (2024)
por: Zhou, Zhou, et al.
Publicado: (2024)
ZSC-Eval: An Evaluation Toolkit and Benchmark for Multi-agent Zero-shot Coordination
por: Wang, Xihuai, et al.
Publicado: (2023)
por: Wang, Xihuai, et al.
Publicado: (2023)
Improving Health Professionals' Onboarding with AI and XAI for Trustworthy Human-AI Collaborative Decision Making
por: Lee, Min Hun, et al.
Publicado: (2024)
por: Lee, Min Hun, et al.
Publicado: (2024)
Frontend Diffusion: Exploring Intent-Based User Interfaces through Abstract-to-Detailed Task Transitions
por: Zhang, Qinshi, et al.
Publicado: (2024)
por: Zhang, Qinshi, et al.
Publicado: (2024)
EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions
por: Bartkowiak, Patryk, et al.
Publicado: (2025)
por: Bartkowiak, Patryk, et al.
Publicado: (2025)
Human-AI Collaborative Uncertainty Quantification
por: Noorani, Sima, et al.
Publicado: (2025)
por: Noorani, Sima, et al.
Publicado: (2025)
Everyday AR through AI-in-the-Loop
por: Suzuki, Ryo, et al.
Publicado: (2024)
por: Suzuki, Ryo, et al.
Publicado: (2024)
Ejemplares similares
-
Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
por: Silva, Inês Oliveira e, et al.
Publicado: (2026) -
"Show Me What's Wrong!": Combining Charts and Text to Guide Data Analysis
por: Feliciano, Beatriz, et al.
Publicado: (2024) -
DiConStruct: Causal Concept-based Explanations through Black-Box Distillation
por: Moreira, Ricardo, et al.
Publicado: (2024) -
From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making
por: Lee, Min Hun
Publicado: (2026) -
What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
por: Ibrahim, Adham, et al.
Publicado: (2024)