Bridging the Training-Inference Gap in LLMs by Leveraging Self-Generated Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Cen, Zhepeng, Liu, Yao, Zeng, Siliang, Chaudhari, Pratik, Rangwala, Huzefa, Karypis, George, Fakoor, Rasool |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025)
by: Zeng, Siliang, et al.
Published: (2025)
Budgeting Counterfactual for Offline RL
by: Liu, Yao, et al.
Published: (2023)
by: Liu, Yao, et al.
Published: (2023)
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
by: Yang, Ke, et al.
Published: (2024)
by: Yang, Ke, et al.
Published: (2024)
Time-Varying Propensity Score to Bridge the Gap between the Past and Present
by: Fakoor, Rasool, et al.
Published: (2022)
by: Fakoor, Rasool, et al.
Published: (2022)
Offline Learning and Forgetting for Reasoning with Large Language Models
by: Ni, Tianwei, et al.
Published: (2025)
by: Ni, Tianwei, et al.
Published: (2025)
AutoG: Towards automatic graph construction from tabular data
by: Chen, Zhikai, et al.
Published: (2025)
by: Chen, Zhikai, et al.
Published: (2025)
Long-context Protein Language Modeling Using Bidirectional Mamba with Shared Projection Layers
by: Wang, Yingheng, et al.
Published: (2024)
by: Wang, Yingheng, et al.
Published: (2024)
DeCaf: A Causal Decoupling Framework for OOD Generalization on Node Classification
by: Han, Xiaoxue, et al.
Published: (2024)
by: Han, Xiaoxue, et al.
Published: (2024)
MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
by: Ou, Jiefu, et al.
Published: (2026)
by: Ou, Jiefu, et al.
Published: (2026)
Protein Structure Tokenization: Benchmarking and New Recipe
by: Yuan, Xinyu, et al.
Published: (2025)
by: Yuan, Xinyu, et al.
Published: (2025)
Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent Space
by: Zhang, Hengrui, et al.
Published: (2023)
by: Zhang, Hengrui, et al.
Published: (2023)
OpenTab: Advancing Large Language Models as Open-domain Table Reasoners
by: Kong, Kezhi, et al.
Published: (2024)
by: Kong, Kezhi, et al.
Published: (2024)
DispaRisk: Auditing Fairness Through Usable Information
by: Vasquez, Jonathan, et al.
Published: (2024)
by: Vasquez, Jonathan, et al.
Published: (2024)
BioBridge: Bridging Biomedical Foundation Models via Knowledge Graphs
by: Wang, Zifeng, et al.
Published: (2023)
by: Wang, Zifeng, et al.
Published: (2023)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
Scalable Prompt Routing via Fine-Grained Latent Task Discovery
by: Zhang, Yunyi, et al.
Published: (2026)
by: Zhang, Yunyi, et al.
Published: (2026)
Extending Input Contexts of Language Models through Training on Segmented Sequences
by: Karypis, Petros, et al.
Published: (2023)
by: Karypis, Petros, et al.
Published: (2023)
Feasibility Consistent Representation Learning for Safe Reinforcement Learning
by: Cen, Zhepeng, et al.
Published: (2024)
by: Cen, Zhepeng, et al.
Published: (2024)
TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks
by: Tschalzev, Andrej, et al.
Published: (2026)
by: Tschalzev, Andrej, et al.
Published: (2026)
Learning the Target Network in Function Space
by: Asadi, Kavosh, et al.
Published: (2024)
by: Asadi, Kavosh, et al.
Published: (2024)
Pushing the Limits of All-Atom Geometric Graph Neural Networks: Pre-Training, Scaling and Zero-Shot Transfer
by: Pengmei, Zihan, et al.
Published: (2024)
by: Pengmei, Zihan, et al.
Published: (2024)
An Effective Gram Matrix Characterizes Generalization in Deep Networks
by: Yang, Rubing, et al.
Published: (2025)
by: Yang, Rubing, et al.
Published: (2025)
Language Modeling with Learned Meta-Tokens
by: Shah, Alok N., et al.
Published: (2025)
by: Shah, Alok N., et al.
Published: (2025)
Behavior Injection: Preparing Language Models for Reinforcement Learning
by: Cen, Zhepeng, et al.
Published: (2025)
by: Cen, Zhepeng, et al.
Published: (2025)
An Information-Geometric Distance on the Space of Tasks
by: Gao, Yansong, et al.
Published: (2020)
by: Gao, Yansong, et al.
Published: (2020)
Model Zoo: A Growing "Brain" That Learns Continually
by: Ramesh, Rahul, et al.
Published: (2021)
by: Ramesh, Rahul, et al.
Published: (2021)
TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models
by: Liu, Zuxin, et al.
Published: (2023)
by: Liu, Zuxin, et al.
Published: (2023)
Relatron: Automating Relational Machine Learning over Relational Databases
by: Chen, Zhikai, et al.
Published: (2026)
by: Chen, Zhikai, et al.
Published: (2026)
Bridging the Gap between Learning and Inference for Diffusion-Based Molecule Generation
by: Liu, Peidong, et al.
Published: (2024)
by: Liu, Peidong, et al.
Published: (2024)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
by: Zhang, Jesse, et al.
Published: (2024)
by: Zhang, Jesse, et al.
Published: (2024)
Learning from Sparse Offline Datasets via Conservative Density Estimation
by: Cen, Zhepeng, et al.
Published: (2024)
by: Cen, Zhepeng, et al.
Published: (2024)
Understanding Silent Data Corruption in LLM Training
by: Ma, Jeffrey, et al.
Published: (2025)
by: Ma, Jeffrey, et al.
Published: (2025)
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
Understanding the Challenges in Iterative Generative Optimization with LLMs
by: Nie, Allen, et al.
Published: (2026)
by: Nie, Allen, et al.
Published: (2026)
Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
by: Zhang, Xiyuan, et al.
Published: (2025)
by: Zhang, Xiyuan, et al.
Published: (2025)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
by: Gai, Jiading, et al.
Published: (2026)
by: Gai, Jiading, et al.
Published: (2026)
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
by: Pengmei, Zihan, et al.
Published: (2025)
by: Pengmei, Zihan, et al.
Published: (2025)
FeatNavigator: Automatic Feature Augmentation on Tabular Data
by: Liang, Jiaming, et al.
Published: (2024)
by: Liang, Jiaming, et al.
Published: (2024)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
by: Yao, Yihang, et al.
Published: (2023)
by: Yao, Yihang, et al.
Published: (2023)
ReSyn: Autonomously Scaling Synthetic Environments for Reasoning Models
by: He, Andre, et al.
Published: (2026)
by: He, Andre, et al.
Published: (2026)
Similar Items
-
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025) -
Budgeting Counterfactual for Offline RL
by: Liu, Yao, et al.
Published: (2023) -
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
by: Yang, Ke, et al.
Published: (2024) -
Time-Varying Propensity Score to Bridge the Gap between the Past and Present
by: Fakoor, Rasool, et al.
Published: (2022) -
Offline Learning and Forgetting for Reasoning with Large Language Models
by: Ni, Tianwei, et al.
Published: (2025)