Saved in:
| Main Authors: | Dorner, Florian E., Nastl, Vivian Y., Hardt, Moritz |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.13341 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do causal predictors generalize better to new domains?
by: Nastl, Vivian Y., et al.
Published: (2024)
by: Nastl, Vivian Y., et al.
Published: (2024)
Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
by: Dorner, Florian E., et al.
Published: (2024)
by: Dorner, Florian E., et al.
Published: (2024)
How Benchmark Prediction from Fewer Data Misses the Mark
by: Zhang, Guanhua, et al.
Published: (2025)
by: Zhang, Guanhua, et al.
Published: (2025)
Training on the Test Task Confounds Evaluation and Emergence
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
Causal Inference from Competing Treatments
by: Stoica, Ana-Andreea, et al.
Published: (2024)
by: Stoica, Ana-Andreea, et al.
Published: (2024)
"If we don't they won't"
by: Ehrhardt, Harryette B.
Published: (1969)
by: Ehrhardt, Harryette B.
Published: (1969)
Britain. Will he, won't he?
Published: (2002)
Published: (2002)
Inherent Trade-Offs between Diversity and Stability in Multi-Task Benchmarks
by: Zhang, Guanhua, et al.
Published: (2024)
by: Zhang, Guanhua, et al.
Published: (2024)
Limits to Predicting Online Speech Using Large Language Models
by: Remeli, Mina, et al.
Published: (2024)
by: Remeli, Mina, et al.
Published: (2024)
Good Allocations from Bad Estimates
by: Casacuberta, Sílvia, et al.
Published: (2026)
by: Casacuberta, Sílvia, et al.
Published: (2026)
Test-Time Training on Nearest Neighbors for Large Language Models
by: Hardt, Moritz, et al.
Published: (2023)
by: Hardt, Moritz, et al.
Published: (2023)
Performative Prediction: Past and Future
by: Hardt, Moritz, et al.
Published: (2023)
by: Hardt, Moritz, et al.
Published: (2023)
Is your model predicting the past?
by: Hardt, Moritz, et al.
Published: (2022)
by: Hardt, Moritz, et al.
Published: (2022)
Where a Pill won't reach
Published: (2003)
Published: (2003)
Price lists won't solve the problem
by: Lucy Dobree
Published: (2025)
by: Lucy Dobree
Published: (2025)
ImageNot: A contrast with ImageNet preserves model rankings
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
What Makes ImageNet Look Unlike LAION
by: Shirali, Ali, et al.
Published: (2023)
by: Shirali, Ali, et al.
Published: (2023)
Unprocessing Seven Years of Algorithmic Fairness
by: Cruz, André F., et al.
Published: (2023)
by: Cruz, André F., et al.
Published: (2023)
American survey. Politics into economics won't go
Published: (1996)
Published: (1996)
Scars that won't heal. The neurobiology of child abuse
Published: (2002)
Published: (2002)
Computational Arbitrage in AI Model Markets
by: Olmedo, Ricardo, et al.
Published: (2026)
by: Olmedo, Ricardo, et al.
Published: (2026)
Allocation Requires Prediction Only if Inequality Is Low
by: Shirali, Ali, et al.
Published: (2024)
by: Shirali, Ali, et al.
Published: (2024)
Leaderboard Incentives: Model Rankings under Strategic Post-Training
by: Chen, Yatong, et al.
Published: (2026)
by: Chen, Yatong, et al.
Published: (2026)
Train-before-Test Harmonizes Language Model Rankings
by: Zhang, Guanhua, et al.
Published: (2025)
by: Zhang, Guanhua, et al.
Published: (2025)
Asia. Reform won't change Japan after all
Published: (1996)
Published: (1996)
Special report zimbabwe. Hell, no, I won't go
Published: (2002)
Published: (2002)
So they won't just be blown by the wind / Elaine Levine
by: Levine, Elaine
by: Levine, Elaine
United States. Old Alabama won't leave politely
Published: (2002)
Published: (2002)
FDA reverses course, won't approve leucovorin for autism
Published: (2026)
Published: (2026)
Evaluating language models as risk scores
by: Cruz, André F., et al.
Published: (2024)
by: Cruz, André F., et al.
Published: (2024)
Best Arm Identification with LLM Judges and Limited Human
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
ROC-n-reroll: How verifier imperfection affects test-time scaling
by: Dorner, Florian E., et al.
Published: (2025)
by: Dorner, Florian E., et al.
Published: (2025)
Scaling Open-Ended Reasoning to Predict the Future
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Algorithmic Collective Action in Machine Learning
by: Hardt, Moritz, et al.
Published: (2023)
by: Hardt, Moritz, et al.
Published: (2023)
Incentivizing Honesty among Competitors in Collaborative Learning and Optimization
by: Dorner, Florian E., et al.
Published: (2023)
by: Dorner, Florian E., et al.
Published: (2023)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Doctor credentialing process in Mass. won't include invasive MH questions
by: Gary Enos
Published: (2024)
by: Gary Enos
Published: (2024)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
by: Hübotter, Jonas, et al.
Published: (2025)
by: Hübotter, Jonas, et al.
Published: (2025)
Just read twice: closing the recall gap for recurrent language models
by: Arora, Simran, et al.
Published: (2024)
by: Arora, Simran, et al.
Published: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Similar Items
-
Do causal predictors generalize better to new domains?
by: Nastl, Vivian Y., et al.
Published: (2024) -
Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
by: Dorner, Florian E., et al.
Published: (2024) -
How Benchmark Prediction from Fewer Data Misses the Mark
by: Zhang, Guanhua, et al.
Published: (2025) -
Training on the Test Task Confounds Evaluation and Emergence
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024) -
Causal Inference from Competing Treatments
by: Stoica, Ana-Andreea, et al.
Published: (2024)