Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
Fuente:
arXiv
Saved in:
| Main Authors: | An, Chenyang, Imani, Shima, Yao, Feng, Dong, Chengyu, Abbasi, Ali, Shrivastava, Harsh, Buss, Samuel, Shang, Jingbo, Mahalingam, Gayathri, Sharma, Pramod, Diesendruck, Maurice |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion-Augmented Coreset Expansion for Scalable Dataset Distillation
by: Abbasi, Ali, et al.
Published: (2024)
by: Abbasi, Ali, et al.
Published: (2024)
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
by: Diesendruck, Maurice, et al.
Published: (2024)
by: Diesendruck, Maurice, et al.
Published: (2024)
Are uGLAD? Time will tell!
by: Imani, Shima, et al.
Published: (2023)
by: Imani, Shima, et al.
Published: (2023)
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022)
by: Dong, Chengyu, et al.
Published: (2022)
Linear Correlation in LM's Compositional Generalization and Hallucination
by: Peng, Letian, et al.
Published: (2025)
by: Peng, Letian, et al.
Published: (2025)
Generative Kaleidoscopic Networks
by: Shrivastava, Harsh
Published: (2024)
by: Shrivastava, Harsh
Published: (2024)
Exploring Group and Symmetry Principles in Large Language Models
by: Imani, Shima, et al.
Published: (2024)
by: Imani, Shima, et al.
Published: (2024)
Correlation and Navigation in the Vocabulary Key Representation Space of Language Models
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
BatchPrompt: Accomplish more with less
by: Lin, Jianzhe, et al.
Published: (2023)
by: Lin, Jianzhe, et al.
Published: (2023)
Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval
by: Zhang, Yuwei, et al.
Published: (2025)
by: Zhang, Yuwei, et al.
Published: (2025)
Neural Graph Revealers
by: Shrivastava, Harsh, et al.
Published: (2023)
by: Shrivastava, Harsh, et al.
Published: (2023)
Knowledge Propagation over Conditional Independence Graphs
by: Chajewska, Urszula, et al.
Published: (2023)
by: Chajewska, Urszula, et al.
Published: (2023)
Federated Learning with Neural Graphical Models
by: Chajewska, Urszula, et al.
Published: (2023)
by: Chajewska, Urszula, et al.
Published: (2023)
Methods for Recovering Conditional Independence Graphs: A Survey
by: Shrivastava, Harsh, et al.
Published: (2022)
by: Shrivastava, Harsh, et al.
Published: (2022)
When is the consistent prediction likely to be a correct prediction?
by: Nguyen, Alex, et al.
Published: (2024)
by: Nguyen, Alex, et al.
Published: (2024)
Evaluating the Smooth Control of Attribute Intensity in Text Generation with LLMs
by: Zhou, Shang, et al.
Published: (2024)
by: Zhou, Shang, et al.
Published: (2024)
DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences
by: Zhou, Xiaoxiao, et al.
Published: (2025)
by: Zhou, Xiaoxiao, et al.
Published: (2025)
Mechanochemical Diversity in Block Copolymers
by: Hang Zhang, et al.
Published: (2024)
by: Hang Zhang, et al.
Published: (2024)
Raising the Bar or Training Library Technicians To Assume Reference Responsibilities.
by: Brandys, Barbara, et al.
Published: (2002)
by: Brandys, Barbara, et al.
Published: (2002)
Reasoning Bias of Next Token Prediction Training
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
Optical, Thermal Studies on Binary and Ternary Hydrogen-Bonded Liquid Crystal Complexes
by: T. Mahalingam
Published: (2016)
by: T. Mahalingam
Published: (2016)
Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text Classification
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
by: Jain, Neel, et al.
Published: (2024)
by: Jain, Neel, et al.
Published: (2024)
Synthesizing Grasps and Regrasps for Complex Manipulation Tasks
by: Patankar, Aditya, et al.
Published: (2025)
by: Patankar, Aditya, et al.
Published: (2025)
The Price of Format: Diversity Collapse in LLMs
by: Yun, Longfei, et al.
Published: (2025)
by: Yun, Longfei, et al.
Published: (2025)
Text Generation Beyond Discrete Token Sampling
by: Zhuang, Yufan, et al.
Published: (2025)
by: Zhuang, Yufan, et al.
Published: (2025)
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
Smaller Language Models are capable of selecting Instruction-Tuning Training Data for Larger Language Models
by: Mekala, Dheeraj, et al.
Published: (2024)
by: Mekala, Dheeraj, et al.
Published: (2024)
Chain‐Folding Effects on Polymer Film Gas Permeabilities
by: Iris Agami, et al.
Published: (2025)
by: Iris Agami, et al.
Published: (2025)
TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
by: Pan, Liang, et al.
Published: (2025)
by: Pan, Liang, et al.
Published: (2025)
Watson-Crick strong bi-catenation on words
by: Mahalingam, Kalpana
Published: (2025)
by: Mahalingam, Kalpana
Published: (2025)
A Logspace Constructive Proof of L=SL
by: Buss, Sam, et al.
Published: (2025)
by: Buss, Sam, et al.
Published: (2025)
OAT: Ordered Action Tokenization
by: Liu, Chaoqi, et al.
Published: (2026)
by: Liu, Chaoqi, et al.
Published: (2026)
A New Klebsiella planticola Strain (Cd-1) Grows Anaerobically at High Cadmium Concentrations and Precipitates Cadmium Sulfide. / Pramod K. Sharma
by: Sharma, Pramod K
Published: (2000)
by: Sharma, Pramod K
Published: (2000)
On the Concept of Inherent Core Damage Frequency: a Framework for Residual Risk Floors in Seismic Probabilistic Safety Assessment
by: Pramod Kumar Sharma
Published: (2026)
by: Pramod Kumar Sharma
Published: (2026)
On the BCSS Proof of the Fundamental Theorem of Algebra
by: Rojas, J. Maurice
Published: (2024)
by: Rojas, J. Maurice
Published: (2024)
VLTP: Vision-Language Guided Token Pruning for Task-Oriented Segmentation
by: Chen, Hanning, et al.
Published: (2024)
by: Chen, Hanning, et al.
Published: (2024)
Efficient Training of Language Models with Compact and Consistent Next Token Distributions
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering
by: Nguyen, Alex, et al.
Published: (2024)
by: Nguyen, Alex, et al.
Published: (2024)
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
Similar Items
-
Diffusion-Augmented Coreset Expansion for Scalable Dataset Distillation
by: Abbasi, Ali, et al.
Published: (2024) -
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
by: Diesendruck, Maurice, et al.
Published: (2024) -
Are uGLAD? Time will tell!
by: Imani, Shima, et al.
Published: (2023) -
Toward Student-Oriented Teacher Network Training For Knowledge Distillation
by: Dong, Chengyu, et al.
Published: (2022) -
Linear Correlation in LM's Compositional Generalization and Hallucination
by: Peng, Letian, et al.
Published: (2025)