Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection
Fuente:
arXiv
Guardado en:
| Autores principales: | Moradi, Mohammad Mahdi, Amer, Hossam, Mudur, Sudhir, Zhang, Weiwei, Liu, Yang, Ahmed, Walid |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DiSCTT: Consensus-Guided Self-Curriculum for Efficient Test-Time Adaptation in Reasoning
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2026)
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2026)
Balancing Computation Load and Representation Expressivity in Parallel Hybrid Neural Networks
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2025)
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2025)
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2025)
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2025)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
por: Amer, Hossam, et al.
Publicado: (2026)
por: Amer, Hossam, et al.
Publicado: (2026)
Distributed Hybrid Parallelism for Large Language Models: Comparative Study and System Design Guide
por: Amer, Hossam, et al.
Publicado: (2026)
por: Amer, Hossam, et al.
Publicado: (2026)
Bayesian Mixture of Experts For Large Language Models
por: Dialameh, Maryam, et al.
Publicado: (2025)
por: Dialameh, Maryam, et al.
Publicado: (2025)
On-Device Emoji Classifier Trained with GPT-based Data Augmentation for a Mobile Keyboard
por: Amer, Hossam, et al.
Publicado: (2024)
por: Amer, Hossam, et al.
Publicado: (2024)
Right for Right Reasons: Large Language Models for Verifiable Commonsense Knowledge Graph Question Answering
por: Toroghi, Armin, et al.
Publicado: (2024)
por: Toroghi, Armin, et al.
Publicado: (2024)
Self-Trained Verification for Training- and Test-Time Self-Improvement
por: Wu, Chen Henry, et al.
Publicado: (2026)
por: Wu, Chen Henry, et al.
Publicado: (2026)
Self-Improvement in Multimodal Large Language Models: A Survey
por: Deng, Shijian, et al.
Publicado: (2025)
por: Deng, Shijian, et al.
Publicado: (2025)
TTSR: Test-Time Self-Reflection for Continual Reasoning Improvement
por: He, Haoyang, et al.
Publicado: (2026)
por: He, Haoyang, et al.
Publicado: (2026)
DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models
por: Guo, YiQiu, et al.
Publicado: (2025)
por: Guo, YiQiu, et al.
Publicado: (2025)
Self-Improvement of Large Language Models: A Technical Overview and Future Outlook
por: Yang, Haoyan, et al.
Publicado: (2026)
por: Yang, Haoyan, et al.
Publicado: (2026)
Dynamic Noise Preference Optimization: Self-Improvement of Large Language Models with Self-Synthetic Data
por: Yang, Haoyan, et al.
Publicado: (2025)
por: Yang, Haoyan, et al.
Publicado: (2025)
Leveraging Test Driven Development with Large Language Models for Reliable and Verifiable Spreadsheet Code Generation: A Research Framework
por: Thorne, Simon, et al.
Publicado: (2025)
por: Thorne, Simon, et al.
Publicado: (2025)
Query-Conditioned Test-Time Self-Training for Large Language Models
por: Song, Chaehee, et al.
Publicado: (2026)
por: Song, Chaehee, et al.
Publicado: (2026)
Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
por: Chang, Kaiyan, et al.
Publicado: (2025)
por: Chang, Kaiyan, et al.
Publicado: (2025)
Test-time Recursive Thinking: Self-Improvement without External Feedback
por: Zhuang, Yufan, et al.
Publicado: (2026)
por: Zhuang, Yufan, et al.
Publicado: (2026)
ETT: Expanding the Long Context Understanding Capability of LLMs at Test-Time
por: Zahirnia, Kiarash, et al.
Publicado: (2025)
por: Zahirnia, Kiarash, et al.
Publicado: (2025)
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
por: Mirbagheri, Mohammad Reza, et al.
Publicado: (2025)
por: Mirbagheri, Mohammad Reza, et al.
Publicado: (2025)
Mid-Training of Large Language Models: A Survey
por: Mo, Kaixiang, et al.
Publicado: (2025)
por: Mo, Kaixiang, et al.
Publicado: (2025)
The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models
por: Jamshidi, Saeid, et al.
Publicado: (2025)
por: Jamshidi, Saeid, et al.
Publicado: (2025)
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
por: Zhang, Jun, et al.
Publicado: (2023)
por: Zhang, Jun, et al.
Publicado: (2023)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
por: Pezeshkpour, Pouya, et al.
Publicado: (2026)
por: Pezeshkpour, Pouya, et al.
Publicado: (2026)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
por: Song, Yuda, et al.
Publicado: (2024)
por: Song, Yuda, et al.
Publicado: (2024)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
por: Hajimolahoseini, Habib, et al.
Publicado: (2023)
por: Hajimolahoseini, Habib, et al.
Publicado: (2023)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
por: Heakl, Ahmed, et al.
Publicado: (2024)
por: Heakl, Ahmed, et al.
Publicado: (2024)
Enabling On-Device Large Language Model Personalization with Self-Supervised Data Selection and Synthesis
por: Qin, Ruiyang, et al.
Publicado: (2023)
por: Qin, Ruiyang, et al.
Publicado: (2023)
SLOT: Sample-specific Language Model Optimization at Test-time
por: Hu, Yang, et al.
Publicado: (2025)
por: Hu, Yang, et al.
Publicado: (2025)
DragD3D: Realistic Mesh Editing with Rigidity Control Driven by 2D Diffusion Priors
por: Xie, Tianhao, et al.
Publicado: (2023)
por: Xie, Tianhao, et al.
Publicado: (2023)
Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA
por: Tang, Xing, et al.
Publicado: (2026)
por: Tang, Xing, et al.
Publicado: (2026)
Verifiable Format Control for Large Language Model Generations
por: Wang, Zhaoyang, et al.
Publicado: (2025)
por: Wang, Zhaoyang, et al.
Publicado: (2025)
Enabling Language Models to Implicitly Learn Self-Improvement
por: Wang, Ziqi, et al.
Publicado: (2023)
por: Wang, Ziqi, et al.
Publicado: (2023)
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
por: Xu, Xin, et al.
Publicado: (2026)
por: Xu, Xin, et al.
Publicado: (2026)
Method-Based Reasoning for Large Language Models: Extraction, Reuse, and Continuous Improvement
por: Su, Hong
Publicado: (2025)
por: Su, Hong
Publicado: (2025)
Generating Diverse Training Samples for Relation Extraction with Large Language Models
por: Li, Zexuan, et al.
Publicado: (2025)
por: Li, Zexuan, et al.
Publicado: (2025)
DAST: Difficulty-Aware Self-Training on Large Language Models
por: Xue, Boyang, et al.
Publicado: (2025)
por: Xue, Boyang, et al.
Publicado: (2025)
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
por: Zhong, Qihuang, et al.
Publicado: (2026)
por: Zhong, Qihuang, et al.
Publicado: (2026)
From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
por: Zhu, Xudong, et al.
Publicado: (2025)
por: Zhu, Xudong, et al.
Publicado: (2025)
Self-Training Large Language Models with Confident Reasoning
por: Jang, Hyosoon, et al.
Publicado: (2025)
por: Jang, Hyosoon, et al.
Publicado: (2025)
Ejemplares similares
-
DiSCTT: Consensus-Guided Self-Curriculum for Efficient Test-Time Adaptation in Reasoning
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2026) -
Balancing Computation Load and Representation Expressivity in Parallel Hybrid Neural Networks
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2025) -
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance
por: Moradi, Mohammad Mahdi, et al.
Publicado: (2025) -
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
por: Amer, Hossam, et al.
Publicado: (2026) -
Distributed Hybrid Parallelism for Large Language Models: Comparative Study and System Design Guide
por: Amer, Hossam, et al.
Publicado: (2026)