Training Large Language Models to Reason in a Continuous Latent Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hao, Shibo, Sukhbaatar, Sainbayar, Su, DiJia, Li, Xian, Hu, Zhiting, Weston, Jason, Tian, Yuandong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
von: Su, DiJia, et al.
Veröffentlicht: (2024)
von: Su, DiJia, et al.
Veröffentlicht: (2024)
Self-Challenging Language Model Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
von: Su, DiJia, et al.
Veröffentlicht: (2025)
von: Su, DiJia, et al.
Veröffentlicht: (2025)
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Multi-Token Attention
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
Contextual Position Encoding: Learning to Count What's Important
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
von: Xu, Jing, et al.
Veröffentlicht: (2023)
von: Xu, Jing, et al.
Veröffentlicht: (2023)
Self-Rewarding Language Models
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping
von: Lehnert, Lucas, et al.
Veröffentlicht: (2024)
von: Lehnert, Lucas, et al.
Veröffentlicht: (2024)
Adaptive Decoding via Latent Preference Optimization
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
Iterative Reasoning Preference Optimization
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
SPICE: Self-Play In Corpus Environments Improves Reasoning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
Thinking LLMs: General Instruction Following with Thought Generation
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
LLM Pretraining with Continuous Concepts
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
R.I.P.: Better Models by Survival of the Fittest Prompts
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Following Length Constraints in Instructions
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
Diverse Preference Optimization
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
von: Su, DiJia, et al.
Veröffentlicht: (2025)
von: Su, DiJia, et al.
Veröffentlicht: (2025)
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
von: Hao, Shibo, et al.
Veröffentlicht: (2023)
von: Hao, Shibo, et al.
Veröffentlicht: (2023)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
von: Yuan, Weizhe, et al.
Veröffentlicht: (2025)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2025)
Bridging Offline and Online Reinforcement Learning for LLMs
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
von: Lanchantin, Jack, et al.
Veröffentlicht: (2025)
SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models
von: Wang, Chenyu, et al.
Veröffentlicht: (2025)
von: Wang, Chenyu, et al.
Veröffentlicht: (2025)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language Models
von: Fan, Zhiting, et al.
Veröffentlicht: (2025)
von: Fan, Zhiting, et al.
Veröffentlicht: (2025)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
Composing Global Solutions to Reasoning Tasks via Algebraic Objects in Neural Nets
von: Tian, Yuandong
Veröffentlicht: (2024)
von: Tian, Yuandong
Veröffentlicht: (2024)
Param$Δ$ for Direct Weight Mixing: Post-Train Large Language Model at Zero Cost
von: Cao, Sheng, et al.
Veröffentlicht: (2025)
von: Cao, Sheng, et al.
Veröffentlicht: (2025)
Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
von: Wang, Xinyuan, et al.
Veröffentlicht: (2025)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2025)
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
von: Lin, Yen-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Yen-Ting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
von: Su, DiJia, et al.
Veröffentlicht: (2024) -
Self-Challenging Language Model Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2025) -
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
von: Su, DiJia, et al.
Veröffentlicht: (2025) -
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024) -
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)