Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
Fuente:
arXiv
Salvato in:
| Autori principali: | Chan, Willy, Souliman, Michael, Nordhagen, Jakob, Miranda, Brando, Obbad, Elyas, Koyejo, Kai Fronsdal Sanmi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
di: Miranda, Brando, et al.
Pubblicazione: (2023)
di: Miranda, Brando, et al.
Pubblicazione: (2023)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
di: Obbad, Elyas, et al.
Pubblicazione: (2024)
di: Obbad, Elyas, et al.
Pubblicazione: (2024)
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
di: Gulati, Aryan, et al.
Pubblicazione: (2025)
di: Gulati, Aryan, et al.
Pubblicazione: (2025)
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
Is Pre-training Truly Better Than Meta-Learning?
di: Miranda, Brando, et al.
Pubblicazione: (2023)
di: Miranda, Brando, et al.
Pubblicazione: (2023)
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
An Evaluation Benchmark for Autoformalization in Lean4
di: Gulati, Aryan, et al.
Pubblicazione: (2024)
di: Gulati, Aryan, et al.
Pubblicazione: (2024)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
di: Vo, Truong, et al.
Pubblicazione: (2025)
di: Vo, Truong, et al.
Pubblicazione: (2025)
Reasoning Models Don't Just Think Longer, They Move Differently
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
di: Tatariya, Kushal, et al.
Pubblicazione: (2024)
di: Tatariya, Kushal, et al.
Pubblicazione: (2024)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
Model-Based Quality Assessment for Massively Multilingual Parallel Data
di: Ibrahim, Abdelaziz M. A., et al.
Pubblicazione: (2026)
di: Ibrahim, Abdelaziz M. A., et al.
Pubblicazione: (2026)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
di: Fronsdal, Kai, et al.
Pubblicazione: (2024)
di: Fronsdal, Kai, et al.
Pubblicazione: (2024)
Discovering Implicit Large Language Model Alignment Objectives
di: Chen, Edward, et al.
Pubblicazione: (2026)
di: Chen, Edward, et al.
Pubblicazione: (2026)
Pantograph: A Machine-to-Machine Interaction Interface for Advanced Theorem Proving, High Level Reasoning, and Data Extraction in Lean 4
di: Aniva, Leni, et al.
Pubblicazione: (2024)
di: Aniva, Leni, et al.
Pubblicazione: (2024)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
di: Negoita, Vlad, et al.
Pubblicazione: (2025)
di: Negoita, Vlad, et al.
Pubblicazione: (2025)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
di: Turki, Yassine, et al.
Pubblicazione: (2026)
di: Turki, Yassine, et al.
Pubblicazione: (2026)
SpecEval: Evaluating Model Adherence to Behavior Specifications
di: Ahmed, Ahmed, et al.
Pubblicazione: (2025)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2025)
The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
di: Wu, Minghao, et al.
Pubblicazione: (2024)
di: Wu, Minghao, et al.
Pubblicazione: (2024)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
di: Liu, Fengze, et al.
Pubblicazione: (2025)
di: Liu, Fengze, et al.
Pubblicazione: (2025)
Are Large Language Models Good Data Preprocessors?
di: Meguellati, Elyas, et al.
Pubblicazione: (2025)
di: Meguellati, Elyas, et al.
Pubblicazione: (2025)
M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets
di: Zhao, Chunguang, et al.
Pubblicazione: (2025)
di: Zhao, Chunguang, et al.
Pubblicazione: (2025)
SampleMix: A Sample-wise Pre-training Data Mixing Strategey by Coordinating Data Quality and Diversity
di: Xi, Xiangyu, et al.
Pubblicazione: (2025)
di: Xi, Xiangyu, et al.
Pubblicazione: (2025)
Investigating Data Contamination for Pre-training Language Models
di: Jiang, Minhao, et al.
Pubblicazione: (2024)
di: Jiang, Minhao, et al.
Pubblicazione: (2024)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
di: Schaeffer, Rylan, et al.
Pubblicazione: (2026)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2026)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026)
di: Zhang, Max, et al.
Pubblicazione: (2026)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
CritiQ: Mining Data Quality Criteria from Human Preferences
di: Guo, Honglin, et al.
Pubblicazione: (2025)
di: Guo, Honglin, et al.
Pubblicazione: (2025)
Test Set Quality in Multilingual LLM Evaluation
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation
di: Iyer, Vivek, et al.
Pubblicazione: (2024)
di: Iyer, Vivek, et al.
Pubblicazione: (2024)
A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility and Quality Perspectives
di: Rao, Hanshu, et al.
Pubblicazione: (2025)
di: Rao, Hanshu, et al.
Pubblicazione: (2025)
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
di: Lau, Mingfei, et al.
Pubblicazione: (2025)
di: Lau, Mingfei, et al.
Pubblicazione: (2025)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
di: Kopiczko, Dawid J., et al.
Pubblicazione: (2026)
di: Kopiczko, Dawid J., et al.
Pubblicazione: (2026)
Data Quality Enhancement on the Basis of Diversity with Large Language Models for Text Classification: Uncovered, Difficult, and Noisy
di: Zeng, Min, et al.
Pubblicazione: (2024)
di: Zeng, Min, et al.
Pubblicazione: (2024)
Extracting books from production language models
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets
di: Samardzic, Tanja, et al.
Pubblicazione: (2024)
di: Samardzic, Tanja, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
di: Miranda, Brando, et al.
Pubblicazione: (2023) -
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
di: Obbad, Elyas, et al.
Pubblicazione: (2024) -
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025) -
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
di: Gulati, Aryan, et al.
Pubblicazione: (2025) -
Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)