Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Long, Yunbo, Afonja, Tejumade, Hao, Guangya, Brintrup, Alexandra, Fritz, Mario |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Topological Federated Clustering via Gravitational Potential Fields under Local Differential Privacy
by: Long, Yunbo, et al.
Published: (2025)
by: Long, Yunbo, et al.
Published: (2025)
LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion
by: Long, Yunbo, et al.
Published: (2025)
by: Long, Yunbo, et al.
Published: (2025)
DP-2Stage: Adapting Language Models as Differentially Private Tabular Data Generators
by: Afonja, Tejumade, et al.
Published: (2024)
by: Afonja, Tejumade, et al.
Published: (2024)
Evaluating Inter-Column Logical Relationships in Synthetic Tabular Data Generation
by: Long, Yunbo, et al.
Published: (2025)
by: Long, Yunbo, et al.
Published: (2025)
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
by: Chen, Jinhao, et al.
Published: (2025)
by: Chen, Jinhao, et al.
Published: (2025)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
by: Liang, Yu, et al.
Published: (2026)
by: Liang, Yu, et al.
Published: (2026)
Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Multi-Agent LLMs
by: Long, Yunbo, et al.
Published: (2026)
by: Long, Yunbo, et al.
Published: (2026)
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
by: Yang, Shidong, et al.
Published: (2026)
by: Yang, Shidong, et al.
Published: (2026)
SynDelay: A Synthetic Dataset for Delivery Delay Prediction
by: Xu, Liming, et al.
Published: (2025)
by: Xu, Liming, et al.
Published: (2025)
Random Walk Guided Hyperbolic Graph Distillation
by: Long, Yunbo, et al.
Published: (2025)
by: Long, Yunbo, et al.
Published: (2025)
Efficient and Privacy-Preserved Link Prediction via Condensed Graphs
by: Long, Yunbo, et al.
Published: (2025)
by: Long, Yunbo, et al.
Published: (2025)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
by: Salimi, Moein, et al.
Published: (2026)
by: Salimi, Moein, et al.
Published: (2026)
Sampling for Quality: Training-Free Reward-Guided LLM Decoding via Sequential Monte Carlo
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
VeriTrace: Evolving Mental Models for Deep Research Agents
by: Zhao, Haolang, et al.
Published: (2026)
by: Zhao, Haolang, et al.
Published: (2026)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
by: Ren, Tao, et al.
Published: (2025)
by: Ren, Tao, et al.
Published: (2025)
Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design
by: Su, Xingyu, et al.
Published: (2025)
by: Su, Xingyu, et al.
Published: (2025)
GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
Active Tabular Augmentation via Policy-Guided Diffusion Inpainting
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
Efficient Process Reward Model Training via Active Learning
by: Duan, Keyu, et al.
Published: (2025)
by: Duan, Keyu, et al.
Published: (2025)
Reward Model Overoptimisation in Iterated RLHF
by: Wolf, Lorenz, et al.
Published: (2025)
by: Wolf, Lorenz, et al.
Published: (2025)
Binning as a Pretext Task: Improving Self-Supervised Learning in Tabular Domains
by: Lee, Kyungeun, et al.
Published: (2024)
by: Lee, Kyungeun, et al.
Published: (2024)
RewardHarness: Self-Evolving Agentic Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
by: Long, Yunbo, et al.
Published: (2026)
by: Long, Yunbo, et al.
Published: (2026)
Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
by: Gupta, Prakhar, et al.
Published: (2025)
by: Gupta, Prakhar, et al.
Published: (2025)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
by: Li, Mengqi, et al.
Published: (2025)
by: Li, Mengqi, et al.
Published: (2025)
Improving Quantization with Post-Training Model Expansion
by: Franco, Giuseppe, et al.
Published: (2025)
by: Franco, Giuseppe, et al.
Published: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
LLM4GRN: Discovering Causal Gene Regulatory Networks with LLMs -- Evaluation through Synthetic Data Generation
by: Afonja, Tejumade, et al.
Published: (2024)
by: Afonja, Tejumade, et al.
Published: (2024)
Towards Biologically Plausible and Private Gene Expression Data Generation
by: Chen, Dingfan, et al.
Published: (2024)
by: Chen, Dingfan, et al.
Published: (2024)
Adversarial Training for Process Reward Models
by: Juneja, Gurusha, et al.
Published: (2025)
by: Juneja, Gurusha, et al.
Published: (2025)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
by: Zhang, Junkai, et al.
Published: (2025)
by: Zhang, Junkai, et al.
Published: (2025)
Improving LLM Group Fairness on Tabular Data via In-Context Learning
by: Cherepanova, Valeriia, et al.
Published: (2024)
by: Cherepanova, Valeriia, et al.
Published: (2024)
Interpreting Language Reward Models via Contrastive Explanations
by: Jiang, Junqi, et al.
Published: (2024)
by: Jiang, Junqi, et al.
Published: (2024)
Understanding Post-Training Structural Changes in Large Language Models
by: He, Xinyu, et al.
Published: (2025)
by: He, Xinyu, et al.
Published: (2025)
Post Training Quantization of Large Language Models with Microscaling Formats
by: Sharify, Sayeh, et al.
Published: (2024)
by: Sharify, Sayeh, et al.
Published: (2024)
West-of-N: Synthetic Preferences for Self-Improving Reward Models
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
GEM-T: Generative Tabular Data via Fitting Moments
by: Li, Miao, et al.
Published: (2025)
by: Li, Miao, et al.
Published: (2025)
Self-Supervised Pre-Training for Precipitation Post-Processor
by: An, Sojung, et al.
Published: (2023)
by: An, Sojung, et al.
Published: (2023)
Similar Items
-
Topological Federated Clustering via Gravitational Potential Fields under Local Differential Privacy
by: Long, Yunbo, et al.
Published: (2025) -
LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion
by: Long, Yunbo, et al.
Published: (2025) -
DP-2Stage: Adapting Language Models as Differentially Private Tabular Data Generators
by: Afonja, Tejumade, et al.
Published: (2024) -
Evaluating Inter-Column Logical Relationships in Synthetic Tabular Data Generation
by: Long, Yunbo, et al.
Published: (2025) -
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
by: Chen, Jinhao, et al.
Published: (2025)