Inherent Trade-Offs between Diversity and Stability in Multi-Task Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Guanhua, Hardt, Moritz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Benchmark Prediction from Fewer Data Misses the Mark
por: Zhang, Guanhua, et al.
Publicado: (2025)
por: Zhang, Guanhua, et al.
Publicado: (2025)
Leaderboard Incentives: Model Rankings under Strategic Post-Training
por: Chen, Yatong, et al.
Publicado: (2026)
por: Chen, Yatong, et al.
Publicado: (2026)
Train-before-Test Harmonizes Language Model Rankings
por: Zhang, Guanhua, et al.
Publicado: (2025)
por: Zhang, Guanhua, et al.
Publicado: (2025)
Good Allocations from Bad Estimates
por: Casacuberta, Sílvia, et al.
Publicado: (2026)
por: Casacuberta, Sílvia, et al.
Publicado: (2026)
Test-Time Training on Nearest Neighbors for Large Language Models
por: Hardt, Moritz, et al.
Publicado: (2023)
por: Hardt, Moritz, et al.
Publicado: (2023)
Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
por: Dorner, Florian E., et al.
Publicado: (2024)
por: Dorner, Florian E., et al.
Publicado: (2024)
Do causal predictors generalize better to new domains?
por: Nastl, Vivian Y., et al.
Publicado: (2024)
por: Nastl, Vivian Y., et al.
Publicado: (2024)
Performative Prediction: Past and Future
por: Hardt, Moritz, et al.
Publicado: (2023)
por: Hardt, Moritz, et al.
Publicado: (2023)
Is your model predicting the past?
por: Hardt, Moritz, et al.
Publicado: (2022)
por: Hardt, Moritz, et al.
Publicado: (2022)
Training on the Test Task Confounds Evaluation and Emergence
por: Dominguez-Olmedo, Ricardo, et al.
Publicado: (2024)
por: Dominguez-Olmedo, Ricardo, et al.
Publicado: (2024)
Trade-Offs of Diagonal Fisher Information Matrix Estimators
por: Soen, Alexander, et al.
Publicado: (2024)
por: Soen, Alexander, et al.
Publicado: (2024)
ImageNot: A contrast with ImageNet preserves model rankings
por: Salaudeen, Olawale, et al.
Publicado: (2024)
por: Salaudeen, Olawale, et al.
Publicado: (2024)
What Makes ImageNet Look Unlike LAION
por: Shirali, Ali, et al.
Publicado: (2023)
por: Shirali, Ali, et al.
Publicado: (2023)
Unprocessing Seven Years of Algorithmic Fairness
por: Cruz, André F., et al.
Publicado: (2023)
por: Cruz, André F., et al.
Publicado: (2023)
Fairness-Accuracy Trade-Offs: A Causal Perspective
por: Plecko, Drago, et al.
Publicado: (2024)
por: Plecko, Drago, et al.
Publicado: (2024)
A Trajectory-Based Bayesian Approach to Multi-Objective Hyperparameter Optimization with Epoch-Aware Trade-Offs
por: Wang, Wenyu, et al.
Publicado: (2024)
por: Wang, Wenyu, et al.
Publicado: (2024)
Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation
por: Cohen, Nadav Z., et al.
Publicado: (2024)
por: Cohen, Nadav Z., et al.
Publicado: (2024)
Computational Arbitrage in AI Model Markets
por: Olmedo, Ricardo, et al.
Publicado: (2026)
por: Olmedo, Ricardo, et al.
Publicado: (2026)
First-See-Then-Design: A Multi-Stakeholder View for Optimal Performance-Fairness Trade-Offs
por: Gupta, Kavya, et al.
Publicado: (2026)
por: Gupta, Kavya, et al.
Publicado: (2026)
An Analytical Approach to Privacy and Performance Trade-Offs in Healthcare Data Sharing
por: Wei, Yusi, et al.
Publicado: (2025)
por: Wei, Yusi, et al.
Publicado: (2025)
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
por: Dorner, Florian E., et al.
Publicado: (2024)
por: Dorner, Florian E., et al.
Publicado: (2024)
Allocation Requires Prediction Only if Inequality Is Low
por: Shirali, Ali, et al.
Publicado: (2024)
por: Shirali, Ali, et al.
Publicado: (2024)
Computational Discovery of Microstructured Composites with Optimal Stiffness-Toughness Trade-Offs
por: Li, Beichen, et al.
Publicado: (2023)
por: Li, Beichen, et al.
Publicado: (2023)
Pruning Extensions and Efficiency Trade-Offs for Sustainable Time Series Classification
por: Fischer, Raphael, et al.
Publicado: (2026)
por: Fischer, Raphael, et al.
Publicado: (2026)
Sharp Trade-Offs in High-Dimensional Inference via 2-Level SLOPE
por: Bu, Zhiqi, et al.
Publicado: (2025)
por: Bu, Zhiqi, et al.
Publicado: (2025)
Least but not Last: Fine-tuning Intermediate Principal Components for Better Performance-Forgetting Trade-Offs
por: Quercia, Alessio, et al.
Publicado: (2026)
por: Quercia, Alessio, et al.
Publicado: (2026)
Utility-Fairness Trade-Offs and How to Find Them
por: Dehdashtian, Sepehr, et al.
Publicado: (2024)
por: Dehdashtian, Sepehr, et al.
Publicado: (2024)
Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
por: Voria, Gianmario, et al.
Publicado: (2024)
por: Voria, Gianmario, et al.
Publicado: (2024)
Evaluating language models as risk scores
por: Cruz, André F., et al.
Publicado: (2024)
por: Cruz, André F., et al.
Publicado: (2024)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
por: Liang, Kaiqu, et al.
Publicado: (2024)
por: Liang, Kaiqu, et al.
Publicado: (2024)
Limits to Predicting Online Speech Using Large Language Models
por: Remeli, Mina, et al.
Publicado: (2024)
por: Remeli, Mina, et al.
Publicado: (2024)
A Multi-Objective Evaluation Framework for Analyzing Utility-Fairness Trade-Offs in Machine Learning Systems
por: Özbulak, Gökhan, et al.
Publicado: (2025)
por: Özbulak, Gökhan, et al.
Publicado: (2025)
Spectral Clustering for Crowdsourcing with Inherently Distinct Task Types
por: Mandal, Saptarshi, et al.
Publicado: (2023)
por: Mandal, Saptarshi, et al.
Publicado: (2023)
FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
por: Kim, Jaemin, et al.
Publicado: (2025)
por: Kim, Jaemin, et al.
Publicado: (2025)
Exploring Performance-Complexity Trade-Offs in Sound Event Detection Models
por: Morocutti, Tobias, et al.
Publicado: (2025)
por: Morocutti, Tobias, et al.
Publicado: (2025)
Beyond Task Diversity: Provable Representation Transfer for Sequential Multi-Task Linear Bandits
por: Duong, Thang, et al.
Publicado: (2025)
por: Duong, Thang, et al.
Publicado: (2025)
First-Explore, then Exploit: Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs
por: Norman, Ben, et al.
Publicado: (2023)
por: Norman, Ben, et al.
Publicado: (2023)
A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs
por: Tanner, Kasimir, et al.
Publicado: (2024)
por: Tanner, Kasimir, et al.
Publicado: (2024)
Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM
por: Bianco, Francesca, et al.
Publicado: (2026)
por: Bianco, Francesca, et al.
Publicado: (2026)
A New Rejection Sampling Approach to $k$-$\mathtt{means}$++ With Improved Trade-Offs
por: Shah, Poojan, et al.
Publicado: (2025)
por: Shah, Poojan, et al.
Publicado: (2025)
Ejemplares similares
-
How Benchmark Prediction from Fewer Data Misses the Mark
por: Zhang, Guanhua, et al.
Publicado: (2025) -
Leaderboard Incentives: Model Rankings under Strategic Post-Training
por: Chen, Yatong, et al.
Publicado: (2026) -
Train-before-Test Harmonizes Language Model Rankings
por: Zhang, Guanhua, et al.
Publicado: (2025) -
Good Allocations from Bad Estimates
por: Casacuberta, Sílvia, et al.
Publicado: (2026) -
Test-Time Training on Nearest Neighbors for Large Language Models
por: Hardt, Moritz, et al.
Publicado: (2023)