Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
Fuente:
arXiv
Saved in:
| Main Authors: | Salaudeen, Olawale, Chiou, Nicole, Weng, Shiny, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causally Inspired Regularization Enables Domain General Representations
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
Toward an Evaluation Science for Generative AI Systems
by: Weidinger, Laura, et al.
Published: (2025)
by: Weidinger, Laura, et al.
Published: (2025)
Proxy Methods for Domain Adaptation
by: Tsai, Katherine, et al.
Published: (2024)
by: Tsai, Katherine, et al.
Published: (2024)
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
by: Stromberg, Nathan, et al.
Published: (2024)
by: Stromberg, Nathan, et al.
Published: (2024)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
by: Zhu, Junzhe, et al.
Published: (2023)
by: Zhu, Junzhe, et al.
Published: (2023)
Optimization and Generalization Guarantees for Weight Normalization
by: Cisneros-Velarde, Pedro, et al.
Published: (2024)
by: Cisneros-Velarde, Pedro, et al.
Published: (2024)
Quantifying Variance in Evaluation Benchmarks
by: Madaan, Lovish, et al.
Published: (2024)
by: Madaan, Lovish, et al.
Published: (2024)
A Framework for Objective-Driven Dynamical Stochastic Fields
by: Zhang, Yibo Jacky, et al.
Published: (2025)
by: Zhang, Yibo Jacky, et al.
Published: (2025)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
by: Chen, Edward, et al.
Published: (2025)
by: Chen, Edward, et al.
Published: (2025)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
by: Gupta, Isha, et al.
Published: (2025)
by: Gupta, Isha, et al.
Published: (2025)
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
by: Sühr, Tom, et al.
Published: (2025)
by: Sühr, Tom, et al.
Published: (2025)
Position: Model Collapse Does Not Mean What You Think
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026)
by: Ahmed, Ahmed, et al.
Published: (2026)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
by: Kazdan, Joshua, et al.
Published: (2024)
by: Kazdan, Joshua, et al.
Published: (2024)
Decision from Suboptimal Classifiers: Excess Risk Pre- and Post-Calibration
by: Perez-Lebel, Alexandre, et al.
Published: (2025)
by: Perez-Lebel, Alexandre, et al.
Published: (2025)
Latent Adversarial Regularization for Offline Preference Optimization
by: Jiang, Enyi, et al.
Published: (2026)
by: Jiang, Enyi, et al.
Published: (2026)
On Fairness of Low-Rank Adaptation of Large Models
by: Ding, Zhoujie, et al.
Published: (2024)
by: Ding, Zhoujie, et al.
Published: (2024)
ImageNot: A contrast with ImageNet preserves model rankings
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
ALMo: Interactive Aim-Limit-Defined, Multi-Objective System for Personalized High-Dose-Rate Brachytherapy Treatment Planning and Visualization for Cervical Cancer
by: Chen, Edward, et al.
Published: (2026)
by: Chen, Edward, et al.
Published: (2026)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
by: Denisov-Blanch, Yegor, et al.
Published: (2026)
by: Denisov-Blanch, Yegor, et al.
Published: (2026)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
by: Zhou, Zhanke, et al.
Published: (2025)
by: Zhou, Zhanke, et al.
Published: (2025)
DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity
by: Zhong, Joey, et al.
Published: (2026)
by: Zhong, Joey, et al.
Published: (2026)
Efficient Prediction of Pass@k Scaling in Large Language Models
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
Neural Nonmyopic Bayesian Optimization in Dynamic Cost Settings
by: Truong, Sang T., et al.
Published: (2026)
by: Truong, Sang T., et al.
Published: (2026)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
The Robustness of Differentiable Causal Discovery in Misspecified Scenarios
by: Yi, Huiyang, et al.
Published: (2025)
by: Yi, Huiyang, et al.
Published: (2025)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Aligning Compound AI Systems via System-level DPO
by: Wang, Xiangwen, et al.
Published: (2025)
by: Wang, Xiangwen, et al.
Published: (2025)
Globalizing Fairness Attributes in Machine Learning: A Case Study on Health in Africa
by: Asiedu, Mercy Nyamewaa, et al.
Published: (2023)
by: Asiedu, Mercy Nyamewaa, et al.
Published: (2023)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
by: Belakaria, Syrine, et al.
Published: (2025)
by: Belakaria, Syrine, et al.
Published: (2025)
Modeling Multi-Objective Tradeoffs with Monotonic Utility Functions
by: Chen, Edward, et al.
Published: (2024)
by: Chen, Edward, et al.
Published: (2024)
What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity from Incidental Causes
by: Lecomte, Victor, et al.
Published: (2023)
by: Lecomte, Victor, et al.
Published: (2023)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion
by: Karnik, Sathwik, et al.
Published: (2026)
by: Karnik, Sathwik, et al.
Published: (2026)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
Scale Dependent Data Duplication
by: Kazdan, Joshua, et al.
Published: (2026)
by: Kazdan, Joshua, et al.
Published: (2026)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
by: Obbad, Elyas, et al.
Published: (2024)
by: Obbad, Elyas, et al.
Published: (2024)
Similar Items
-
Causally Inspired Regularization Enables Domain General Representations
by: Salaudeen, Olawale, et al.
Published: (2024) -
Toward an Evaluation Science for Generative AI Systems
by: Weidinger, Laura, et al.
Published: (2025) -
Proxy Methods for Domain Adaptation
by: Tsai, Katherine, et al.
Published: (2024) -
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
by: Stromberg, Nathan, et al.
Published: (2024) -
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
by: Zhu, Junzhe, et al.
Published: (2023)