Learning from Saturated Data: Signals Beyond Correctness for LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hiss, Hanno, Dekoninck, Jasper, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
von: Petrov, Ivo, et al.
Veröffentlicht: (2026)
von: Petrov, Ivo, et al.
Veröffentlicht: (2026)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
A Unified Approach to Routing and Cascading for LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
ConStat: Performance-Based Contamination Detection in Large Language Models
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
Controlled Text Generation via Language Model Arithmetic
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2023)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2023)
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
von: Mündler, Niels, et al.
Veröffentlicht: (2025)
von: Mündler, Niels, et al.
Veröffentlicht: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
Evading Data Contamination Detection for Language Models is (too) Easy
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
Adaptive Generation of Bias-Eliciting Questions for LLMs
von: Staab, Robin, et al.
Veröffentlicht: (2025)
von: Staab, Robin, et al.
Veröffentlicht: (2025)
The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2025)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2025)
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
Beyond Public Access in LLM Pre-Training Data
von: Rosenblat, Sruly, et al.
Veröffentlicht: (2025)
von: Rosenblat, Sruly, et al.
Veröffentlicht: (2025)
Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2024)
von: Beurer-Kellner, Luca, et al.
Veröffentlicht: (2024)
Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2026)
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2026)
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
von: Chen, Jiamin, et al.
Veröffentlicht: (2026)
von: Chen, Jiamin, et al.
Veröffentlicht: (2026)
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Metric-Dependent Annotation Saturation for Learning from Label Distributions
von: Kohli, Guneet
Veröffentlicht: (2026)
von: Kohli, Guneet
Veröffentlicht: (2026)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
von: Guldimann, Philipp, et al.
Veröffentlicht: (2024)
von: Guldimann, Philipp, et al.
Veröffentlicht: (2024)
Tailored Truths: Optimizing LLM Persuasion with Personalization and Fabricated Statistics
von: Timm, Jasper, et al.
Veröffentlicht: (2025)
von: Timm, Jasper, et al.
Veröffentlicht: (2025)
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
SyncThink: A Training-Free Strategy to Align Inference Termination with Reasoning Saturation
von: Li, Gengyang, et al.
Veröffentlicht: (2026)
von: Li, Gengyang, et al.
Veröffentlicht: (2026)
A Training-free LLM-based Approach to General Chinese Character Error Correction
von: Zhou, Houquan, et al.
Veröffentlicht: (2025)
von: Zhou, Houquan, et al.
Veröffentlicht: (2025)
A Synthetic Dataset for Personal Attribute Inference
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024)
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
Saturation-Driven Dataset Generation for LLM Mathematical Reasoning in the TPTP Ecosystem
von: Quesnel, Valentin, et al.
Veröffentlicht: (2025)
von: Quesnel, Valentin, et al.
Veröffentlicht: (2025)
The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
CuTS: Customizable Tabular Synthetic Data Generation
von: Vero, Mark, et al.
Veröffentlicht: (2023)
von: Vero, Mark, et al.
Veröffentlicht: (2023)
Beyond Correctness: Learning Robust Reasoning via Transfer
von: Lee, Hyunseok, et al.
Veröffentlicht: (2026)
von: Lee, Hyunseok, et al.
Veröffentlicht: (2026)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
von: Tong, Chaodong, et al.
Veröffentlicht: (2025)
von: Tong, Chaodong, et al.
Veröffentlicht: (2025)
Learning from Mistakes: Self-correct Adversarial Training for Chinese Unnatural Text Correction
von: Feng, Xuan, et al.
Veröffentlicht: (2024)
von: Feng, Xuan, et al.
Veröffentlicht: (2024)
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
von: Li, Bolian, et al.
Veröffentlicht: (2026)
von: Li, Bolian, et al.
Veröffentlicht: (2026)
Reinforcement Learning for LLM Post-Training: A Survey
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
von: Zhang, Zhihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2024)
Detecting Non-Membership in LLM Training Data via Rank Correlations
von: Shetty, Pranav, et al.
Veröffentlicht: (2026)
von: Shetty, Pranav, et al.
Veröffentlicht: (2026)
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
von: Langlais, Pierre-Carl, et al.
Veröffentlicht: (2025)
von: Langlais, Pierre-Carl, et al.
Veröffentlicht: (2025)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025)
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
von: Petrov, Ivo, et al.
Veröffentlicht: (2026) -
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024) -
A Unified Approach to Routing and Cascading for LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024) -
ConStat: Performance-Based Contamination Detection in Large Language Models
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024) -
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)