Inherent Trade-Offs between Diversity and Stability in Multi-Task Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Guanhua, Hardt, Moritz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Benchmark Prediction from Fewer Data Misses the Mark
von: Zhang, Guanhua, et al.
Veröffentlicht: (2025)
von: Zhang, Guanhua, et al.
Veröffentlicht: (2025)
Leaderboard Incentives: Model Rankings under Strategic Post-Training
von: Chen, Yatong, et al.
Veröffentlicht: (2026)
von: Chen, Yatong, et al.
Veröffentlicht: (2026)
Train-before-Test Harmonizes Language Model Rankings
von: Zhang, Guanhua, et al.
Veröffentlicht: (2025)
von: Zhang, Guanhua, et al.
Veröffentlicht: (2025)
Good Allocations from Bad Estimates
von: Casacuberta, Sílvia, et al.
Veröffentlicht: (2026)
von: Casacuberta, Sílvia, et al.
Veröffentlicht: (2026)
Test-Time Training on Nearest Neighbors for Large Language Models
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
von: Dorner, Florian E., et al.
Veröffentlicht: (2024)
von: Dorner, Florian E., et al.
Veröffentlicht: (2024)
Do causal predictors generalize better to new domains?
von: Nastl, Vivian Y., et al.
Veröffentlicht: (2024)
von: Nastl, Vivian Y., et al.
Veröffentlicht: (2024)
Performative Prediction: Past and Future
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
Is your model predicting the past?
von: Hardt, Moritz, et al.
Veröffentlicht: (2022)
von: Hardt, Moritz, et al.
Veröffentlicht: (2022)
Training on the Test Task Confounds Evaluation and Emergence
von: Dominguez-Olmedo, Ricardo, et al.
Veröffentlicht: (2024)
von: Dominguez-Olmedo, Ricardo, et al.
Veröffentlicht: (2024)
Trade-Offs of Diagonal Fisher Information Matrix Estimators
von: Soen, Alexander, et al.
Veröffentlicht: (2024)
von: Soen, Alexander, et al.
Veröffentlicht: (2024)
ImageNot: A contrast with ImageNet preserves model rankings
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2024)
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2024)
What Makes ImageNet Look Unlike LAION
von: Shirali, Ali, et al.
Veröffentlicht: (2023)
von: Shirali, Ali, et al.
Veröffentlicht: (2023)
Unprocessing Seven Years of Algorithmic Fairness
von: Cruz, André F., et al.
Veröffentlicht: (2023)
von: Cruz, André F., et al.
Veröffentlicht: (2023)
Fairness-Accuracy Trade-Offs: A Causal Perspective
von: Plecko, Drago, et al.
Veröffentlicht: (2024)
von: Plecko, Drago, et al.
Veröffentlicht: (2024)
A Trajectory-Based Bayesian Approach to Multi-Objective Hyperparameter Optimization with Epoch-Aware Trade-Offs
von: Wang, Wenyu, et al.
Veröffentlicht: (2024)
von: Wang, Wenyu, et al.
Veröffentlicht: (2024)
Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation
von: Cohen, Nadav Z., et al.
Veröffentlicht: (2024)
von: Cohen, Nadav Z., et al.
Veröffentlicht: (2024)
Computational Arbitrage in AI Model Markets
von: Olmedo, Ricardo, et al.
Veröffentlicht: (2026)
von: Olmedo, Ricardo, et al.
Veröffentlicht: (2026)
First-See-Then-Design: A Multi-Stakeholder View for Optimal Performance-Fairness Trade-Offs
von: Gupta, Kavya, et al.
Veröffentlicht: (2026)
von: Gupta, Kavya, et al.
Veröffentlicht: (2026)
An Analytical Approach to Privacy and Performance Trade-Offs in Healthcare Data Sharing
von: Wei, Yusi, et al.
Veröffentlicht: (2025)
von: Wei, Yusi, et al.
Veröffentlicht: (2025)
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
von: Dorner, Florian E., et al.
Veröffentlicht: (2024)
von: Dorner, Florian E., et al.
Veröffentlicht: (2024)
Allocation Requires Prediction Only if Inequality Is Low
von: Shirali, Ali, et al.
Veröffentlicht: (2024)
von: Shirali, Ali, et al.
Veröffentlicht: (2024)
Computational Discovery of Microstructured Composites with Optimal Stiffness-Toughness Trade-Offs
von: Li, Beichen, et al.
Veröffentlicht: (2023)
von: Li, Beichen, et al.
Veröffentlicht: (2023)
Pruning Extensions and Efficiency Trade-Offs for Sustainable Time Series Classification
von: Fischer, Raphael, et al.
Veröffentlicht: (2026)
von: Fischer, Raphael, et al.
Veröffentlicht: (2026)
Sharp Trade-Offs in High-Dimensional Inference via 2-Level SLOPE
von: Bu, Zhiqi, et al.
Veröffentlicht: (2025)
von: Bu, Zhiqi, et al.
Veröffentlicht: (2025)
Least but not Last: Fine-tuning Intermediate Principal Components for Better Performance-Forgetting Trade-Offs
von: Quercia, Alessio, et al.
Veröffentlicht: (2026)
von: Quercia, Alessio, et al.
Veröffentlicht: (2026)
Utility-Fairness Trade-Offs and How to Find Them
von: Dehdashtian, Sepehr, et al.
Veröffentlicht: (2024)
von: Dehdashtian, Sepehr, et al.
Veröffentlicht: (2024)
Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
von: Voria, Gianmario, et al.
Veröffentlicht: (2024)
von: Voria, Gianmario, et al.
Veröffentlicht: (2024)
Evaluating language models as risk scores
von: Cruz, André F., et al.
Veröffentlicht: (2024)
von: Cruz, André F., et al.
Veröffentlicht: (2024)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
von: Liang, Kaiqu, et al.
Veröffentlicht: (2024)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2024)
Limits to Predicting Online Speech Using Large Language Models
von: Remeli, Mina, et al.
Veröffentlicht: (2024)
von: Remeli, Mina, et al.
Veröffentlicht: (2024)
A Multi-Objective Evaluation Framework for Analyzing Utility-Fairness Trade-Offs in Machine Learning Systems
von: Özbulak, Gökhan, et al.
Veröffentlicht: (2025)
von: Özbulak, Gökhan, et al.
Veröffentlicht: (2025)
Spectral Clustering for Crowdsourcing with Inherently Distinct Task Types
von: Mandal, Saptarshi, et al.
Veröffentlicht: (2023)
von: Mandal, Saptarshi, et al.
Veröffentlicht: (2023)
FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
Exploring Performance-Complexity Trade-Offs in Sound Event Detection Models
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
Beyond Task Diversity: Provable Representation Transfer for Sequential Multi-Task Linear Bandits
von: Duong, Thang, et al.
Veröffentlicht: (2025)
von: Duong, Thang, et al.
Veröffentlicht: (2025)
First-Explore, then Exploit: Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs
von: Norman, Ben, et al.
Veröffentlicht: (2023)
von: Norman, Ben, et al.
Veröffentlicht: (2023)
A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs
von: Tanner, Kasimir, et al.
Veröffentlicht: (2024)
von: Tanner, Kasimir, et al.
Veröffentlicht: (2024)
Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM
von: Bianco, Francesca, et al.
Veröffentlicht: (2026)
von: Bianco, Francesca, et al.
Veröffentlicht: (2026)
A New Rejection Sampling Approach to $k$-$\mathtt{means}$++ With Improved Trade-Offs
von: Shah, Poojan, et al.
Veröffentlicht: (2025)
von: Shah, Poojan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Benchmark Prediction from Fewer Data Misses the Mark
von: Zhang, Guanhua, et al.
Veröffentlicht: (2025) -
Leaderboard Incentives: Model Rankings under Strategic Post-Training
von: Chen, Yatong, et al.
Veröffentlicht: (2026) -
Train-before-Test Harmonizes Language Model Rankings
von: Zhang, Guanhua, et al.
Veröffentlicht: (2025) -
Good Allocations from Bad Estimates
von: Casacuberta, Sílvia, et al.
Veröffentlicht: (2026) -
Test-Time Training on Nearest Neighbors for Large Language Models
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)