How Many Ratings per Item are Necessary for Reliable Significance Testing?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Homan, Christopher, Korn, Flip, Pandita, Deepak, Welty, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
von: Pandita, Deepak, et al.
Veröffentlicht: (2026)
von: Pandita, Deepak, et al.
Veröffentlicht: (2026)
Forest vs Tree: The $(N, K)$ Trade-off in Reproducible ML Evaluation
von: Pandita, Deepak, et al.
Veröffentlicht: (2025)
von: Pandita, Deepak, et al.
Veröffentlicht: (2025)
Beyond Random Sampling: Instance Quality-Based Data Partitioning via Item Response Theory
von: Cardoso, Lucas, et al.
Veröffentlicht: (2025)
von: Cardoso, Lucas, et al.
Veröffentlicht: (2025)
Super Apriel: One Checkpoint, Many Speeds
von: Labs, SLAM, et al.
Veröffentlicht: (2026)
von: Labs, SLAM, et al.
Veröffentlicht: (2026)
Fusing Rewards and Preferences in Reinforcement Learning
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)
How Reliable and Stable are Explanations of XAI Methods?
von: Ribeiro, José, et al.
Veröffentlicht: (2024)
von: Ribeiro, José, et al.
Veröffentlicht: (2024)
Explanations Based on Item Response Theory (eXirt): A Model-Specific Method to Explain Tree-Ensemble Model in Trust Perspective
von: Ribeiro, José, et al.
Veröffentlicht: (2022)
von: Ribeiro, José, et al.
Veröffentlicht: (2022)
Securing Reliability: A Brief Overview on Enhancing In-Context Learning for Foundation Models
von: Huang, Yunpeng, et al.
Veröffentlicht: (2024)
von: Huang, Yunpeng, et al.
Veröffentlicht: (2024)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
von: McCann, Jordan F.
Veröffentlicht: (2026)
von: McCann, Jordan F.
Veröffentlicht: (2026)
Learning Can Converge Stably to the Wrong Belief under Latent Reliability
von: Zhang, Zhipeng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhipeng, et al.
Veröffentlicht: (2026)
How to Boost Any Loss Function
von: Nock, Richard, et al.
Veröffentlicht: (2024)
von: Nock, Richard, et al.
Veröffentlicht: (2024)
Strengthening the Internal Adversarial Robustness in Lifted Neural Networks
von: Zach, Christopher
Veröffentlicht: (2025)
von: Zach, Christopher
Veröffentlicht: (2025)
Explainable Graph Representation Learning via Graph Pattern Analysis
von: Wang, Xudong, et al.
Veröffentlicht: (2025)
von: Wang, Xudong, et al.
Veröffentlicht: (2025)
Superposition Is Not Necessary: A Mechanistic Interpretability Analysis of Transformer Representations for Time Series Forecasting
von: Yıldırım, Alper
Veröffentlicht: (2026)
von: Yıldırım, Alper
Veröffentlicht: (2026)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024)
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024)
Is ReLU Adversarially Robust?
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
von: Yousaf, Iqra
Veröffentlicht: (2024)
von: Yousaf, Iqra
Veröffentlicht: (2024)
Beyond Normality: Reliable A/B Testing with Non-Gaussian Data
von: Gong, Junpeng, et al.
Veröffentlicht: (2025)
von: Gong, Junpeng, et al.
Veröffentlicht: (2025)
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
von: Mongaras, Gabriel, et al.
Veröffentlicht: (2026)
von: Mongaras, Gabriel, et al.
Veröffentlicht: (2026)
How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures
von: Gupta, Krishnam
Veröffentlicht: (2026)
von: Gupta, Krishnam
Veröffentlicht: (2026)
FluidWorld: Reaction-Diffusion Dynamics as a Predictive Substrate for World Models
von: Polly, Fabien
Veröffentlicht: (2026)
von: Polly, Fabien
Veröffentlicht: (2026)
I-GLIDE: Input Groups for Latent Health Indicators in Degradation Estimation
von: Thil, Lucas, et al.
Veröffentlicht: (2025)
von: Thil, Lucas, et al.
Veröffentlicht: (2025)
Potential-Based Reward Shaping For Intrinsic Motivation
von: Forbes, Grant C., et al.
Veröffentlicht: (2024)
von: Forbes, Grant C., et al.
Veröffentlicht: (2024)
Interpretable Multi-View Clustering
von: Jiang, Mudi, et al.
Veröffentlicht: (2024)
von: Jiang, Mudi, et al.
Veröffentlicht: (2024)
The Bayesian Confidence (BACON) Estimator for Deep Neural Networks
von: Kee, Patrick D., et al.
Veröffentlicht: (2024)
von: Kee, Patrick D., et al.
Veröffentlicht: (2024)
Pre-Ictal Seizure Prediction Using Personalized Deep Learning
von: Jaddu, Shriya, et al.
Veröffentlicht: (2024)
von: Jaddu, Shriya, et al.
Veröffentlicht: (2024)
xLSTM-Mixer: Multivariate Time Series Forecasting by Mixing via Scalar Memories
von: Kraus, Maurice, et al.
Veröffentlicht: (2024)
von: Kraus, Maurice, et al.
Veröffentlicht: (2024)
Representation learning with CGAN for casual inference
von: Weng, Zhaotian, et al.
Veröffentlicht: (2024)
von: Weng, Zhaotian, et al.
Veröffentlicht: (2024)
Boosting gets full Attention for Relational Learning
von: Guillame-Bert, Mathieu, et al.
Veröffentlicht: (2024)
von: Guillame-Bert, Mathieu, et al.
Veröffentlicht: (2024)
Data-Incremental Continual Offline Reinforcement Learning
von: Gai, Sibo, et al.
Veröffentlicht: (2024)
von: Gai, Sibo, et al.
Veröffentlicht: (2024)
Adaptive Epsilon Adversarial Training for Robust Gravitational Wave Parameter Estimation Using Normalizing Flows
von: Yang, Yiqian, et al.
Veröffentlicht: (2024)
von: Yang, Yiqian, et al.
Veröffentlicht: (2024)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
von: Gray, Gavia, et al.
Veröffentlicht: (2024)
von: Gray, Gavia, et al.
Veröffentlicht: (2024)
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
von: Forbes, Grant C., et al.
Veröffentlicht: (2024)
von: Forbes, Grant C., et al.
Veröffentlicht: (2024)
CPT: Competence-progressive Training Strategy for Few-shot Node Classification
von: Yan, Qilong, et al.
Veröffentlicht: (2024)
von: Yan, Qilong, et al.
Veröffentlicht: (2024)
Learning Useful Representations of Recurrent Neural Network Weight Matrices
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
New Paradigm of Adversarial Training: Releasing Accuracy-Robustness Trade-Off via Dummy Class
von: Wang, Yanyun, et al.
Veröffentlicht: (2024)
von: Wang, Yanyun, et al.
Veröffentlicht: (2024)
RobustBlack: Challenging Black-Box Adversarial Attacks on State-of-the-Art Defenses
von: Djilani, Mohamed, et al.
Veröffentlicht: (2024)
von: Djilani, Mohamed, et al.
Veröffentlicht: (2024)
Trusted Multi-view Learning under Noisy Supervision
von: Zhang, Yilin, et al.
Veröffentlicht: (2024)
von: Zhang, Yilin, et al.
Veröffentlicht: (2024)
Versatile Ordering Network: An Attention-based Neural Network for Ordering Across Scales and Quality Metrics
von: Yu, Zehua, et al.
Veröffentlicht: (2024)
von: Yu, Zehua, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
von: Pandita, Deepak, et al.
Veröffentlicht: (2026) -
Forest vs Tree: The $(N, K)$ Trade-off in Reproducible ML Evaluation
von: Pandita, Deepak, et al.
Veröffentlicht: (2025) -
Beyond Random Sampling: Instance Quality-Based Data Partitioning via Item Response Theory
von: Cardoso, Lucas, et al.
Veröffentlicht: (2025) -
Super Apriel: One Checkpoint, Many Speeds
von: Labs, SLAM, et al.
Veröffentlicht: (2026) -
Fusing Rewards and Preferences in Reinforcement Learning
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)