AQuA: A Benchmarking Tool for Label Quality Assessment
Fuente:
arXiv
Saved in:
| Main Authors: | Goswami, Mononito, Sanil, Vedant, Choudhry, Arjun, Srinivasan, Arvind, Udompanyawit, Chalisa, Dubrawski, Artur |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOMENT: A Family of Open Time-series Foundation Models
by: Goswami, Mononito, et al.
Published: (2024)
by: Goswami, Mononito, et al.
Published: (2024)
TimeSeriesExam: A time series understanding exam
by: Cai, Yifu, et al.
Published: (2024)
by: Cai, Yifu, et al.
Published: (2024)
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
by: Gwiazda, Malgorzata, et al.
Published: (2026)
by: Gwiazda, Malgorzata, et al.
Published: (2026)
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
by: Cai, Yifu, et al.
Published: (2025)
by: Cai, Yifu, et al.
Published: (2025)
Exploring Representations and Interventions in Time Series Foundation Models
by: Wiliński, Michał, et al.
Published: (2024)
by: Wiliński, Michał, et al.
Published: (2024)
Towards Long-Context Time Series Foundation Models
by: Żukowska, Nina, et al.
Published: (2024)
by: Żukowska, Nina, et al.
Published: (2024)
AQuA -- Combining Experts' and Non-Experts' Views To Assess Deliberation Quality in Online Discussions Using LLMs
by: Behrendt, Maike, et al.
Published: (2024)
by: Behrendt, Maike, et al.
Published: (2024)
Implicit Reasoning in Deep Time Series Forecasting
by: Potosnak, Willa, et al.
Published: (2024)
by: Potosnak, Willa, et al.
Published: (2024)
MICA: Multivariate Infini Compressive Attention for Time Series Forecasting
by: Potosnak, Willa, et al.
Published: (2026)
by: Potosnak, Willa, et al.
Published: (2026)
STAMP: Spatial-Temporal Adapter with Multi-Head Pooling
by: Shook, Brad, et al.
Published: (2025)
by: Shook, Brad, et al.
Published: (2025)
Investigating Compositional Reasoning in Time Series Foundation Models
by: Potosnak, Willa, et al.
Published: (2025)
by: Potosnak, Willa, et al.
Published: (2025)
Signal Quality Auditing for Time-series Data
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions
by: Jang, Jihyoung, et al.
Published: (2026)
by: Jang, Jihyoung, et al.
Published: (2026)
Exploring Loss Design Techniques For Decision Tree Robustness To Label Noise
by: Sztukiewicz, Lukasz, et al.
Published: (2024)
by: Sztukiewicz, Lukasz, et al.
Published: (2024)
AQuA: Automated Question-Answering in Software Tutorial Videos with Visual Anchors
by: Yang, Saelyne, et al.
Published: (2024)
by: Yang, Saelyne, et al.
Published: (2024)
Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting
by: Garza, Azul, et al.
Published: (2026)
by: Garza, Azul, et al.
Published: (2026)
Non-Stationarity in the Embedding Space of Time Series Foundation Models
by: Choi, Jinmyeong, et al.
Published: (2026)
by: Choi, Jinmyeong, et al.
Published: (2026)
A Rate-Distortion View of Uncertainty Quantification
by: Apostolopoulou, Ifigeneia, et al.
Published: (2024)
by: Apostolopoulou, Ifigeneia, et al.
Published: (2024)
Multimodal Structure Preservation Learning
by: Liu, Chang, et al.
Published: (2024)
by: Liu, Chang, et al.
Published: (2024)
Enhanced Uncertainty Estimation in Ultrasound Image Segmentation with MSU-Net
by: Banerjee, Rohini, et al.
Published: (2024)
by: Banerjee, Rohini, et al.
Published: (2024)
Leveraging Expert Consistency to Improve Algorithmic Decision Support
by: De-Arteaga, Maria, et al.
Published: (2021)
by: De-Arteaga, Maria, et al.
Published: (2021)
Mitigating Persistent Client Dropout in Asynchronous Decentralized Federated Learning
by: Stępka, Ignacy, et al.
Published: (2025)
by: Stępka, Ignacy, et al.
Published: (2025)
SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization
by: Chaudhari, Shravan, et al.
Published: (2025)
by: Chaudhari, Shravan, et al.
Published: (2025)
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
by: Xie, Stephan, et al.
Published: (2026)
by: Xie, Stephan, et al.
Published: (2026)
Global Deep Forecasting with Patient-Specific Pharmacokinetics
by: Potosnak, Willa, et al.
Published: (2023)
by: Potosnak, Willa, et al.
Published: (2023)
A Benchmarking Framework for AI models in Automotive Aerodynamics
by: Tangsali, Kaustubh, et al.
Published: (2025)
by: Tangsali, Kaustubh, et al.
Published: (2025)
Guidelines for the Quality Assessment of Energy-Aware NAS Benchmarks
by: Kocher, Nick, et al.
Published: (2025)
by: Kocher, Nick, et al.
Published: (2025)
Conformal Prediction: A Theoretical Note and Benchmarking Transductive Node Classification in Graphs
by: Maneriker, Pranav, et al.
Published: (2024)
by: Maneriker, Pranav, et al.
Published: (2024)
Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks
by: Palit, Vedant
Published: (2025)
by: Palit, Vedant
Published: (2025)
HierarchicalForecast: A Reference Framework for Hierarchical Forecasting in Python
by: Olivares, Kin G., et al.
Published: (2022)
by: Olivares, Kin G., et al.
Published: (2022)
A Mixture of Experts Gating Network for Enhanced Surrogate Modeling in External Aerodynamics
by: Nabian, Mohammad Amin, et al.
Published: (2025)
by: Nabian, Mohammad Amin, et al.
Published: (2025)
Adaptive Parameter Optimization for Robust Remote Photoplethysmography
by: Morales, Cecilia G., et al.
Published: (2025)
by: Morales, Cecilia G., et al.
Published: (2025)
Brain-Computer Interfaces for Emotional Regulation in Patients with Various Disorders
by: Mehta, Vedant
Published: (2024)
by: Mehta, Vedant
Published: (2024)
Motion Informed Needle Segmentation in Ultrasound Images
by: Goel, Raghavv, et al.
Published: (2023)
by: Goel, Raghavv, et al.
Published: (2023)
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
by: Siddiq, Mohammed Latif, et al.
Published: (2024)
by: Siddiq, Mohammed Latif, et al.
Published: (2024)
Selection of Optimal Number and Location of PMUs for CNN Based Fault Location and Identification
by: Khattak, Khalid Daud, et al.
Published: (2025)
by: Khattak, Khalid Daud, et al.
Published: (2025)
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
by: Goswami, Mandip
Published: (2026)
by: Goswami, Mandip
Published: (2026)
Bifurcation Identification for Ultrasound-driven Robotic Cannulation
by: Morales, Cecilia G., et al.
Published: (2024)
by: Morales, Cecilia G., et al.
Published: (2024)
Does Prompt Design Impact Quality of Data Imputation by LLMs?
by: Srinivasan, Shreenidhi, et al.
Published: (2025)
by: Srinivasan, Shreenidhi, et al.
Published: (2025)
Automated Neuron Labelling Enables Generative Steering and Interpretability in Protein Language Models
by: Banerjee, Arjun, et al.
Published: (2025)
by: Banerjee, Arjun, et al.
Published: (2025)
Similar Items
-
MOMENT: A Family of Open Time-series Foundation Models
by: Goswami, Mononito, et al.
Published: (2024) -
TimeSeriesExam: A time series understanding exam
by: Cai, Yifu, et al.
Published: (2024) -
TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale
by: Gwiazda, Malgorzata, et al.
Published: (2026) -
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
by: Cai, Yifu, et al.
Published: (2025) -
Exploring Representations and Interventions in Time Series Foundation Models
by: Wiliński, Michał, et al.
Published: (2024)