Exploring the Impact of Dataset Statistical Effect Size on Model Performance and Data Sample Size Sufficiency
Fuente:
arXiv
Saved in:
| Main Authors: | Hatamian, Arya, Levine, Lionel, Oskouie, Haniyeh Ehsani, Sarrafzadeh, Majid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Cross-model Neuronal Correlations in the Context of Predicting Model Performance and Generalizability
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2024)
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2024)
MI-to-Mid Distilled Compression (M2M-DC): An Hybrid-Information-Guided-Block Pruning with Progressive Inner Slicing Approach to Model Compression
by: Levine, Lionel, et al.
Published: (2025)
by: Levine, Lionel, et al.
Published: (2025)
Leveraging Large Language Models and Topic Modeling for Toxicity Classification
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2024)
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2024)
Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models
by: Ghiasvand, Sajjad, et al.
Published: (2025)
by: Ghiasvand, Sajjad, et al.
Published: (2025)
Attack on Scene Flow using Point Clouds
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2024)
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2024)
MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
by: Ghiasvand, Sajjad, et al.
Published: (2026)
by: Ghiasvand, Sajjad, et al.
Published: (2026)
Test-Time Defense Against Adversarial Attacks via Stochastic Resonance of Latent Ensembles
by: Lao, Dong, et al.
Published: (2025)
by: Lao, Dong, et al.
Published: (2025)
IT-OSE: Exploring Optimal Sample Size for Industrial Data Augmentation
by: Sun, Mingchun, et al.
Published: (2026)
by: Sun, Mingchun, et al.
Published: (2026)
When to Retrain after Drift: A Data-Only Test of Post-Drift Data Size Sufficiency
by: Fujiwara, Ren, et al.
Published: (2026)
by: Fujiwara, Ren, et al.
Published: (2026)
PRISM: A Transformer-based Language Model of Structured Clinical Event Data
by: Levine, Lionel, et al.
Published: (2025)
by: Levine, Lionel, et al.
Published: (2025)
A Bayesian Model for Online Activity Sample Sizes
by: Richardson, Thomas, et al.
Published: (2021)
by: Richardson, Thomas, et al.
Published: (2021)
Sample Amplification: Increasing Dataset Size even when Learning is Impossible
by: Axelrod, Brian, et al.
Published: (2019)
by: Axelrod, Brian, et al.
Published: (2019)
Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics
by: Wei, Aaron, et al.
Published: (2025)
by: Wei, Aaron, et al.
Published: (2025)
Private Regression via Data-Dependent Sufficient Statistic Perturbation
by: Ferrando, Cecilia, et al.
Published: (2024)
by: Ferrando, Cecilia, et al.
Published: (2024)
Disentangle Sample Size and Initialization Effect on Perfect Generalization for Single-Neuron Target
by: Zhao, Jiajie, et al.
Published: (2024)
by: Zhao, Jiajie, et al.
Published: (2024)
Generalized Data Thinning Using Sufficient Statistics
by: Dharamshi, Ameer, et al.
Published: (2023)
by: Dharamshi, Ameer, et al.
Published: (2023)
Learngene Search Across Multiple Datasets for Building Variable-Sized Models
by: Shi, Boyu, et al.
Published: (2026)
by: Shi, Boyu, et al.
Published: (2026)
The Impact of Model Zoo Size and Composition on Weight Space Learning
by: Falk, Damian, et al.
Published: (2025)
by: Falk, Damian, et al.
Published: (2025)
Unveiling Scaling Behaviors in Molecular Language Models: Effects of Model Size, Data, and Representation
by: Xu, Dong, et al.
Published: (2026)
by: Xu, Dong, et al.
Published: (2026)
Algorithmic Accountability in Small Data: Sample-Size-Induced Bias Within Classification Metrics
by: Briscoe, Jarren, et al.
Published: (2025)
by: Briscoe, Jarren, et al.
Published: (2025)
On Size-Independent Sample Complexity of ReLU Networks
by: Sellke, Mark
Published: (2023)
by: Sellke, Mark
Published: (2023)
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
by: Jayawardhana, Mayuka, et al.
Published: (2025)
by: Jayawardhana, Mayuka, et al.
Published: (2025)
Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
by: Lanzillotta, Giulia, et al.
Published: (2025)
by: Lanzillotta, Giulia, et al.
Published: (2025)
Effective Sample Size and Generalization Bounds for Temporal Networks
by: Gahtan, Barak, et al.
Published: (2025)
by: Gahtan, Barak, et al.
Published: (2025)
Multi-Objective Alignment of Language Models for Personalized Psychotherapy
by: Beikzadeh, Mehrab, et al.
Published: (2026)
by: Beikzadeh, Mehrab, et al.
Published: (2026)
Conditional Generative Models are Sufficient to Sample from Any Causal Effect Estimand
by: Rahman, Md Musfiqur, et al.
Published: (2024)
by: Rahman, Md Musfiqur, et al.
Published: (2024)
Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs
by: Tan, Kevin, et al.
Published: (2024)
by: Tan, Kevin, et al.
Published: (2024)
Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance Constraints
by: Xia, Xiaobo, et al.
Published: (2023)
by: Xia, Xiaobo, et al.
Published: (2023)
Effect of Kernel Size on CNN-Vision-Transformer-Based Gaze Prediction Using Electroencephalography Data
by: Qiu, Chuhui, et al.
Published: (2024)
by: Qiu, Chuhui, et al.
Published: (2024)
A Classical View on Benign Overfitting: The Role of Sample Size
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Fast Rate Bounds for Multi-Task and Meta-Learning with Different Sample Sizes
by: Zakerinia, Hossein, et al.
Published: (2025)
by: Zakerinia, Hossein, et al.
Published: (2025)
Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation
by: Mahapatra, Joy, et al.
Published: (2025)
by: Mahapatra, Joy, et al.
Published: (2025)
A Statistical Theory of Contrastive Learning via Approximate Sufficient Statistics
by: Lin, Licong, et al.
Published: (2025)
by: Lin, Licong, et al.
Published: (2025)
Sample Size Calculations for Developing Clinical Prediction Models: Overview and pmsims R package
by: Shamsutdinova, Diana, et al.
Published: (2026)
by: Shamsutdinova, Diana, et al.
Published: (2026)
Exploring Time-Step Size in Reinforcement Learning for Sepsis Treatment
by: Sun, Yingchuan, et al.
Published: (2025)
by: Sun, Yingchuan, et al.
Published: (2025)
Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models
by: Ehsani, Kimia, et al.
Published: (2025)
by: Ehsani, Kimia, et al.
Published: (2025)
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
by: Wang, Mingze, et al.
Published: (2026)
by: Wang, Mingze, et al.
Published: (2026)
PRISM-Consult: A Panel-of-Experts Architecture for Clinician-Aligned Diagnosis
by: Levine, Lionel, et al.
Published: (2025)
by: Levine, Lionel, et al.
Published: (2025)
Toward Generalizable Machine Learning Models in Speech, Language, and Hearing Sciences: Estimating Sample Size and Reducing Overfitting
by: Ghasemzadeh, Hamzeh, et al.
Published: (2023)
by: Ghasemzadeh, Hamzeh, et al.
Published: (2023)
Similar Items
-
Exploring Cross-model Neuronal Correlations in the Context of Predicting Model Performance and Generalizability
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2024) -
MI-to-Mid Distilled Compression (M2M-DC): An Hybrid-Information-Guided-Block Pruning with Progressive Inner Slicing Approach to Model Compression
by: Levine, Lionel, et al.
Published: (2025) -
Leveraging Large Language Models and Topic Modeling for Toxicity Classification
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2024) -
Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022) -
Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models
by: Ghiasvand, Sajjad, et al.
Published: (2025)