Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Ru, Huang, Wei, Cao, Qi, Iwasawa, Yusuke, Matsuo, Yutaka, Guo, Jiaxian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
von: Cao, Qi, et al.
Veröffentlicht: (2026)
von: Cao, Qi, et al.
Veröffentlicht: (2026)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
von: Yang, Bo, et al.
Veröffentlicht: (2025)
von: Yang, Bo, et al.
Veröffentlicht: (2025)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
von: Hu, Zhangyi, et al.
Veröffentlicht: (2026)
von: Hu, Zhangyi, et al.
Veröffentlicht: (2026)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
von: Kubo, Kenji, et al.
Veröffentlicht: (2026)
von: Kubo, Kenji, et al.
Veröffentlicht: (2026)
Self-Distilled Agentic Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
von: Sokota, Samuel, et al.
Veröffentlicht: (2025)
von: Sokota, Samuel, et al.
Veröffentlicht: (2025)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
Self-Supervised Multimodal Learning: A Survey
von: Zong, Yongshuo, et al.
Veröffentlicht: (2023)
von: Zong, Yongshuo, et al.
Veröffentlicht: (2023)
Efficient Test-Time Scaling via Self-Calibration
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
Self-Trained Verification for Training- and Test-Time Self-Improvement
von: Wu, Chen Henry, et al.
Veröffentlicht: (2026)
von: Wu, Chen Henry, et al.
Veröffentlicht: (2026)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
von: Zhang, Ru, et al.
Veröffentlicht: (2026)
von: Zhang, Ru, et al.
Veröffentlicht: (2026)
TTCS: Test-Time Curriculum Synthesis for Self-Evolving
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning
von: Yang, Chengcao
Veröffentlicht: (2026)
von: Yang, Chengcao
Veröffentlicht: (2026)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
Knowledge Graph Reasoning with Self-supervised Reinforcement Learning
von: Ma, Ying, et al.
Veröffentlicht: (2024)
von: Ma, Ying, et al.
Veröffentlicht: (2024)
Self-Hinting Language Models Enhance Reinforcement Learning
von: Liao, Baohao, et al.
Veröffentlicht: (2026)
von: Liao, Baohao, et al.
Veröffentlicht: (2026)
Self-Improving LLM Agents at Test-Time
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
von: Ding, Zifeng, et al.
Veröffentlicht: (2025)
Adaptive Self-Supervised Learning Strategies for Dynamic On-Device LLM Personalization
von: Mendoza, Rafael, et al.
Veröffentlicht: (2024)
von: Mendoza, Rafael, et al.
Veröffentlicht: (2024)
Understanding Emergent Misalignment via Feature Superposition Geometry
von: Minegishi, Gouki, et al.
Veröffentlicht: (2026)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2026)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4
von: Guo, Jiaxian, et al.
Veröffentlicht: (2023)
von: Guo, Jiaxian, et al.
Veröffentlicht: (2023)
Clustering Properties of Self-Supervised Learning
von: Weng, Xi, et al.
Veröffentlicht: (2025)
von: Weng, Xi, et al.
Veröffentlicht: (2025)
Self-Supervised Prompt Optimization
von: Xiang, Jinyu, et al.
Veröffentlicht: (2025)
von: Xiang, Jinyu, et al.
Veröffentlicht: (2025)
Self-Play Preference Optimization for Language Model Alignment
von: Wu, Yue, et al.
Veröffentlicht: (2024)
von: Wu, Yue, et al.
Veröffentlicht: (2024)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
von: Yang, Wenjie, et al.
Veröffentlicht: (2025)
von: Yang, Wenjie, et al.
Veröffentlicht: (2025)
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Linear Complexity Self-Supervised Learning for Music Understanding with Random Quantizer
von: Vavaroutsos, Petros, et al.
Veröffentlicht: (2026)
von: Vavaroutsos, Petros, et al.
Veröffentlicht: (2026)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
von: Kim, Bum Jun, et al.
Veröffentlicht: (2026)
von: Kim, Bum Jun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024) -
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
von: Cao, Qi, et al.
Veröffentlicht: (2026) -
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
von: Yang, Bo, et al.
Veröffentlicht: (2025) -
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025) -
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)