Quantifying the Importance of Data Alignment in Downstream Model Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Chawla, Krrish, Sahai, Aryan, DePavia, Mario, Sundar, Sudharsan, Miranda, Brando, Obbad, Elyas, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
by: Chan, Willy, et al.
Published: (2025)
by: Chan, Willy, et al.
Published: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
by: Obbad, Elyas, et al.
Published: (2024)
by: Obbad, Elyas, et al.
Published: (2024)
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
by: Gulati, Aryan, et al.
Published: (2025)
by: Gulati, Aryan, et al.
Published: (2025)
Scaling Laws for Downstream Task Performance of Large Language Models
by: Isik, Berivan, et al.
Published: (2024)
by: Isik, Berivan, et al.
Published: (2024)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
An Evaluation Benchmark for Autoformalization in Lean4
by: Gulati, Aryan, et al.
Published: (2024)
by: Gulati, Aryan, et al.
Published: (2024)
Discovering Implicit Large Language Model Alignment Objectives
by: Chen, Edward, et al.
Published: (2026)
by: Chen, Edward, et al.
Published: (2026)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Is Pre-training Truly Better Than Meta-Learning?
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
by: Schaeffer, Rylan, et al.
Published: (2026)
by: Schaeffer, Rylan, et al.
Published: (2026)
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
by: Wang, Angelina, et al.
Published: (2025)
by: Wang, Angelina, et al.
Published: (2025)
SpecEval: Evaluating Model Adherence to Behavior Specifications
by: Ahmed, Ahmed, et al.
Published: (2025)
by: Ahmed, Ahmed, et al.
Published: (2025)
DafnyMPI: A Dafny Library for Verifying Message-Passing Concurrent Programs
by: Fedchin, Aleksandr, et al.
Published: (2025)
by: Fedchin, Aleksandr, et al.
Published: (2025)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
Yale-DM-Lab at ArchEHR-QA 2026: Deterministic Grounding and Multi-Pass Evidence Alignment for EHR Question Answering
by: Irankhah, Elyas, et al.
Published: (2026)
by: Irankhah, Elyas, et al.
Published: (2026)
EditLens: Quantifying the Extent of AI Editing in Text
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
by: Wang, Angelina, et al.
Published: (2025)
by: Wang, Angelina, et al.
Published: (2025)
SUQL: Conversational Search over Structured and Unstructured Data with Large Language Models
by: Liu, Shicheng, et al.
Published: (2023)
by: Liu, Shicheng, et al.
Published: (2023)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
A Performance Model for Warp Specialization Kernels
by: Liu, Zhengyang, et al.
Published: (2025)
by: Liu, Zhengyang, et al.
Published: (2025)
Converting IEC 61131-3 LD into SFC Using Large Language Model: Dataset and Testing
by: Zhang, Yimin, et al.
Published: (2025)
by: Zhang, Yimin, et al.
Published: (2025)
Parsing TTree Formula in Python
by: Roy, Aryan, et al.
Published: (2025)
by: Roy, Aryan, et al.
Published: (2025)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Are Large Language Models Good Data Preprocessors?
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
Investigating Data Contamination for Pre-training Language Models
by: Jiang, Minhao, et al.
Published: (2024)
by: Jiang, Minhao, et al.
Published: (2024)
Data Models of German Lute Tablature With TScore
by: Lepper, Markus, et al.
Published: (2024)
by: Lepper, Markus, et al.
Published: (2024)
KestRel: Relational Verification Using E-Graphs for Program Alignment
by: Dickerson, Robert, et al.
Published: (2024)
by: Dickerson, Robert, et al.
Published: (2024)
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026)
by: Ahmed, Ahmed, et al.
Published: (2026)
Scalable Ensembling For Mitigating Reward Overoptimisation
by: Ahmed, Ahmed M., et al.
Published: (2024)
by: Ahmed, Ahmed M., et al.
Published: (2024)
Correctness is Demanding, Performance is Frustrating
by: Sinkarovs, Artjoms, et al.
Published: (2024)
by: Sinkarovs, Artjoms, et al.
Published: (2024)
Quantifier Elimination and Craig Interpolation, Quantitatively
by: Batz, Kevin, et al.
Published: (2025)
by: Batz, Kevin, et al.
Published: (2025)
Exploring LLM Support for Generating IEC 61131-3 Graphic Language Programs
by: Zhang, Yimin, et al.
Published: (2024)
by: Zhang, Yimin, et al.
Published: (2024)
Skitter: A Distributed Stream Processing Framework with Pluggable Distribution Strategies
by: Saey, Mathijs, et al.
Published: (2025)
by: Saey, Mathijs, et al.
Published: (2025)
Reactive Programming without Functions
by: Oeyen, Bjarno, et al.
Published: (2024)
by: Oeyen, Bjarno, et al.
Published: (2024)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
by: Zhou, Zhanke, et al.
Published: (2025)
by: Zhou, Zhanke, et al.
Published: (2025)
Similar Items
-
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
by: Miranda, Brando, et al.
Published: (2023) -
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
by: Chan, Willy, et al.
Published: (2025) -
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
by: Obbad, Elyas, et al.
Published: (2024) -
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
by: Gulati, Aryan, et al.
Published: (2025) -
Scaling Laws for Downstream Task Performance of Large Language Models
by: Isik, Berivan, et al.
Published: (2024)