Self-Trained Verification for Training- and Test-Time Self-Improvement
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Chen Henry, Raghunathan, Aditi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mode-Conditioning Unlocks Superior Test-Time Scaling
by: Wu, Chen Henry, et al.
Published: (2025)
by: Wu, Chen Henry, et al.
Published: (2025)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
Query-Conditioned Test-Time Self-Training for Large Language Models
by: Song, Chaehee, et al.
Published: (2026)
by: Song, Chaehee, et al.
Published: (2026)
TTSR: Test-Time Self-Reflection for Continual Reasoning Improvement
by: He, Haoyang, et al.
Published: (2026)
by: He, Haoyang, et al.
Published: (2026)
In-Place Test-Time Training
by: Feng, Guhao, et al.
Published: (2026)
by: Feng, Guhao, et al.
Published: (2026)
TSO: Self-Training with Scaled Preference Optimization
by: Chen, Kaihui, et al.
Published: (2024)
by: Chen, Kaihui, et al.
Published: (2024)
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
by: Hu, Zhangyi, et al.
Published: (2026)
by: Hu, Zhangyi, et al.
Published: (2026)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
by: Finlayson, Matthew, et al.
Published: (2025)
by: Finlayson, Matthew, et al.
Published: (2025)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
Test-Time Detoxification without Training or Learning Anything
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Self-Improving LLM Agents at Test-Time
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
by: Schröder, Christopher, et al.
Published: (2024)
by: Schröder, Christopher, et al.
Published: (2024)
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
by: Chen, Maximillian, et al.
Published: (2024)
by: Chen, Maximillian, et al.
Published: (2024)
SR-TTT: Surprisal-Aware Residual Test-Time Training
by: P, Swamynathan V
Published: (2026)
by: P, Swamynathan V
Published: (2026)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
by: Xu, Fangzhi, et al.
Published: (2025)
by: Xu, Fangzhi, et al.
Published: (2025)
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025)
by: Shah, Rishi Rajesh, et al.
Published: (2025)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
by: Xi, Zhiheng, et al.
Published: (2024)
by: Xi, Zhiheng, et al.
Published: (2024)
Self-Training Elicits Concise Reasoning in Large Language Models
by: Munkhbat, Tergel, et al.
Published: (2025)
by: Munkhbat, Tergel, et al.
Published: (2025)
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
TTCS: Test-Time Curriculum Synthesis for Self-Evolving
by: Yang, Chengyi, et al.
Published: (2026)
by: Yang, Chengyi, et al.
Published: (2026)
Efficient Test-Time Scaling via Self-Calibration
by: Huang, Chengsong, et al.
Published: (2025)
by: Huang, Chengsong, et al.
Published: (2025)
Self-Improvement in Language Models: The Sharpening Mechanism
by: Huang, Audrey, et al.
Published: (2024)
by: Huang, Audrey, et al.
Published: (2024)
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
by: von Oswald, Johannes, et al.
Published: (2025)
by: von Oswald, Johannes, et al.
Published: (2025)
Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models
by: Pinneri, Cristina, et al.
Published: (2025)
by: Pinneri, Cristina, et al.
Published: (2025)
Self-Improvement as Coherence Optimization: A Theoretical Account
by: Qiu, Tianyi, et al.
Published: (2026)
by: Qiu, Tianyi, et al.
Published: (2026)
Latent Principle Discovery for Language Model Self-Improvement
by: Ramji, Keshav, et al.
Published: (2025)
by: Ramji, Keshav, et al.
Published: (2025)
Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
by: Subramaniam, Vighnesh, et al.
Published: (2025)
by: Subramaniam, Vighnesh, et al.
Published: (2025)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
by: Koh, Woosung, et al.
Published: (2025)
by: Koh, Woosung, et al.
Published: (2025)
CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
by: Shen, Zhanming, et al.
Published: (2025)
by: Shen, Zhanming, et al.
Published: (2025)
Base Models Look Human To AI Detectors
by: Xu, Yixuan Even, et al.
Published: (2026)
by: Xu, Yixuan Even, et al.
Published: (2026)
Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
by: Long, Quanyu, et al.
Published: (2026)
by: Long, Quanyu, et al.
Published: (2026)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
by: Lv, Xingtai, et al.
Published: (2026)
by: Lv, Xingtai, et al.
Published: (2026)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
by: Liang, Yu, et al.
Published: (2026)
by: Liang, Yu, et al.
Published: (2026)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
by: Lee, Jaehyeok, et al.
Published: (2024)
by: Lee, Jaehyeok, et al.
Published: (2024)
A Self-matching Training Method with Annotation Embedding Models for Ontology Subsumption Prediction
by: Shiraishi, Yukihiro, et al.
Published: (2024)
by: Shiraishi, Yukihiro, et al.
Published: (2024)
Memorization Sinks: Isolating Memorization during LLM Training
by: Ghosal, Gaurav R., et al.
Published: (2025)
by: Ghosal, Gaurav R., et al.
Published: (2025)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
Training on the Test Task Confounds Evaluation and Emergence
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
by: Lin, Weizhe, et al.
Published: (2025)
by: Lin, Weizhe, et al.
Published: (2025)
Similar Items
-
Mode-Conditioning Unlocks Superior Test-Time Scaling
by: Wu, Chen Henry, et al.
Published: (2025) -
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025) -
Query-Conditioned Test-Time Self-Training for Large Language Models
by: Song, Chaehee, et al.
Published: (2026) -
TTSR: Test-Time Self-Reflection for Continual Reasoning Improvement
by: He, Haoyang, et al.
Published: (2026) -
In-Place Test-Time Training
by: Feng, Guhao, et al.
Published: (2026)