Quantifying Variance in Evaluation Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Madaan, Lovish, Singh, Aaditya K., Schaeffer, Rylan, Poulton, Andrew, Koyejo, Sanmi, Stenetorp, Pontus, Narang, Sharan, Hupkes, Dieuwke |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
por: Roberts, Nicholas, et al.
Publicado: (2025)
por: Roberts, Nicholas, et al.
Publicado: (2025)
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
por: Madaan, Lovish, et al.
Publicado: (2024)
por: Madaan, Lovish, et al.
Publicado: (2024)
Position: Model Collapse Does Not Mean What You Think
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
por: Gupta, Isha, et al.
Publicado: (2025)
por: Gupta, Isha, et al.
Publicado: (2025)
In-Context Learning of Energy Functions
por: Schaeffer, Rylan, et al.
Publicado: (2024)
por: Schaeffer, Rylan, et al.
Publicado: (2024)
What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity from Incidental Causes
por: Lecomte, Victor, et al.
Publicado: (2023)
por: Lecomte, Victor, et al.
Publicado: (2023)
Efficient Prediction of Pass@k Scaling in Large Language Models
por: Kazdan, Joshua, et al.
Publicado: (2025)
por: Kazdan, Joshua, et al.
Publicado: (2025)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
por: Kazdan, Joshua, et al.
Publicado: (2024)
por: Kazdan, Joshua, et al.
Publicado: (2024)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
por: Denisov-Blanch, Yegor, et al.
Publicado: (2026)
por: Denisov-Blanch, Yegor, et al.
Publicado: (2026)
Pretraining Scaling Laws for Generative Evaluations of Language Models
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
por: Obbad, Elyas, et al.
Publicado: (2024)
por: Obbad, Elyas, et al.
Publicado: (2024)
Investigating Data Contamination for Pre-training Language Models
por: Jiang, Minhao, et al.
Publicado: (2024)
por: Jiang, Minhao, et al.
Publicado: (2024)
Scale Dependent Data Duplication
por: Kazdan, Joshua, et al.
Publicado: (2026)
por: Kazdan, Joshua, et al.
Publicado: (2026)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
por: Kazdan, Joshua, et al.
Publicado: (2025)
por: Kazdan, Joshua, et al.
Publicado: (2025)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
por: Miranda, Brando, et al.
Publicado: (2023)
por: Miranda, Brando, et al.
Publicado: (2023)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
por: Salaudeen, Olawale, et al.
Publicado: (2025)
por: Salaudeen, Olawale, et al.
Publicado: (2025)
How Do Large Language Monkeys Get Their Power (Laws)?
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
por: Schaeffer, Rylan, et al.
Publicado: (2024)
por: Schaeffer, Rylan, et al.
Publicado: (2024)
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
por: Nathani, Deepak, et al.
Publicado: (2025)
por: Nathani, Deepak, et al.
Publicado: (2025)
Best-of-N Jailbreaking
por: Hughes, John, et al.
Publicado: (2024)
por: Hughes, John, et al.
Publicado: (2024)
Quantifying Generative Media Bias with a Corpus of Real-world and Generated News Articles
por: Trhlik, Filip, et al.
Publicado: (2024)
por: Trhlik, Filip, et al.
Publicado: (2024)
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
por: Zhu, Junzhe, et al.
Publicado: (2023)
por: Zhu, Junzhe, et al.
Publicado: (2023)
HARP: A challenging human-annotated math reasoning benchmark
por: Yue, Albert S., et al.
Publicado: (2024)
por: Yue, Albert S., et al.
Publicado: (2024)
Reliable and Efficient Amortized Model-based Evaluation
por: Truong, Sang, et al.
Publicado: (2025)
por: Truong, Sang, et al.
Publicado: (2025)
A Framework for Objective-Driven Dynamical Stochastic Fields
por: Zhang, Yibo Jacky, et al.
Publicado: (2025)
por: Zhang, Yibo Jacky, et al.
Publicado: (2025)
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
por: Stromberg, Nathan, et al.
Publicado: (2024)
por: Stromberg, Nathan, et al.
Publicado: (2024)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
por: Chen, Edward, et al.
Publicado: (2025)
por: Chen, Edward, et al.
Publicado: (2025)
Why Do Safety Guardrails Degrade Across Languages?
por: Zhang, Max, et al.
Publicado: (2026)
por: Zhang, Max, et al.
Publicado: (2026)
Logits are All We Need to Adapt Closed Models
por: Hiranandani, Gaurush, et al.
Publicado: (2025)
por: Hiranandani, Gaurush, et al.
Publicado: (2025)
Quantifying the Importance of Data Alignment in Downstream Model Performance
por: Chawla, Krrish, et al.
Publicado: (2025)
por: Chawla, Krrish, et al.
Publicado: (2025)
Jet Expansions of Residual Computation
por: Chen, Yihong, et al.
Publicado: (2024)
por: Chen, Yihong, et al.
Publicado: (2024)
Optimization and Generalization Guarantees for Weight Normalization
por: Cisneros-Velarde, Pedro, et al.
Publicado: (2024)
por: Cisneros-Velarde, Pedro, et al.
Publicado: (2024)
Extracting books from production language models
por: Ahmed, Ahmed, et al.
Publicado: (2026)
por: Ahmed, Ahmed, et al.
Publicado: (2026)
Position: Beyond Sensitive Attributes, ML Fairness Should Quantify Structural Injustice via Social Determinants
por: Tang, Zeyu, et al.
Publicado: (2025)
por: Tang, Zeyu, et al.
Publicado: (2025)
Decision from Suboptimal Classifiers: Excess Risk Pre- and Post-Calibration
por: Perez-Lebel, Alexandre, et al.
Publicado: (2025)
por: Perez-Lebel, Alexandre, et al.
Publicado: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
por: Schaeffer, Rylan, et al.
Publicado: (2026)
por: Schaeffer, Rylan, et al.
Publicado: (2026)
Rethinking Thinking Tokens: LLMs as Improvement Operators
por: Madaan, Lovish, et al.
Publicado: (2025)
por: Madaan, Lovish, et al.
Publicado: (2025)
Latent Adversarial Regularization for Offline Preference Optimization
por: Jiang, Enyi, et al.
Publicado: (2026)
por: Jiang, Enyi, et al.
Publicado: (2026)
Ejemplares similares
-
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
por: Schaeffer, Rylan, et al.
Publicado: (2025) -
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
por: Roberts, Nicholas, et al.
Publicado: (2025) -
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
por: Madaan, Lovish, et al.
Publicado: (2024) -
Position: Model Collapse Does Not Mean What You Think
por: Schaeffer, Rylan, et al.
Publicado: (2025) -
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
por: Gupta, Isha, et al.
Publicado: (2025)