PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Boyi, Li, He, Song, Shixiang, Wang, Yixuan, Wang, Zitong, He, Ziwei, Wang, Xinbing, Lin, Zhouhan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PonderLM: Pretraining Language Models to Ponder in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
PonderLM-3: Adaptive Token-Wise Pondering with Differentiable Masking
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
by: Song, Shixiang, et al.
Published: (2026)
by: Song, Shixiang, et al.
Published: (2026)
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
by: Zeng, Boyi, et al.
Published: (2026)
by: Zeng, Boyi, et al.
Published: (2026)
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
by: Kai, Jushi, et al.
Published: (2025)
by: Kai, Jushi, et al.
Published: (2025)
Learning to Ponder: Adaptive Reasoning in Latent Space
by: He, Yixin, et al.
Published: (2025)
by: He, Yixin, et al.
Published: (2025)
Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
by: He, Ziwei, et al.
Published: (2023)
by: He, Ziwei, et al.
Published: (2023)
HuRef: HUman-REadable Fingerprint for Large Language Models
by: Zeng, Boyi, et al.
Published: (2023)
by: Zeng, Boyi, et al.
Published: (2023)
VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
by: Wang, Yixuan, et al.
Published: (2026)
by: Wang, Yixuan, et al.
Published: (2026)
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
by: Shen, Junzhe, et al.
Published: (2026)
by: Shen, Junzhe, et al.
Published: (2026)
Towards Controlled Table-to-Text Generation with Scientific Reasoning
by: Guo, Zhixin, et al.
Published: (2023)
by: Guo, Zhixin, et al.
Published: (2023)
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
by: Liu, Yuliang, et al.
Published: (2026)
by: Liu, Yuliang, et al.
Published: (2026)
BAMBINO-LM: (Bilingual-)Human-Inspired Continual Pretraining of BabyLM
by: Shen, Zhewen, et al.
Published: (2024)
by: Shen, Zhewen, et al.
Published: (2024)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
by: Li, Jeffrey, et al.
Published: (2025)
by: Li, Jeffrey, et al.
Published: (2025)
MLP Memory: A Retriever-Pretrained Memory for Large Language Models
by: Wei, Rubin, et al.
Published: (2025)
by: Wei, Rubin, et al.
Published: (2025)
FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation
by: Zhang, Runzhe, et al.
Published: (2026)
by: Zhang, Runzhe, et al.
Published: (2026)
ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
ControlLM: Crafting Diverse Personalities for Language Models
by: Weng, Yixuan, et al.
Published: (2024)
by: Weng, Yixuan, et al.
Published: (2024)
Adapting Knowledge for Few-shot Table-to-Text Generation
by: Guo, Zhixin, et al.
Published: (2023)
by: Guo, Zhixin, et al.
Published: (2023)
Context-level Language Modeling by Learning Predictive Context Embeddings
by: Dai, Beiya, et al.
Published: (2025)
by: Dai, Beiya, et al.
Published: (2025)
KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning
by: Yu, Peng, et al.
Published: (2024)
by: Yu, Peng, et al.
Published: (2024)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
by: Cao, Jiaqi, et al.
Published: (2025)
by: Cao, Jiaqi, et al.
Published: (2025)
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
LM-Combiner: A Contextual Rewriting Model for Chinese Grammatical Error Correction
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text
by: Ginn, Michael, et al.
Published: (2024)
by: Ginn, Michael, et al.
Published: (2024)
Steer LLM Latents for Hallucination Detection
by: Park, Seongheon, et al.
Published: (2025)
by: Park, Seongheon, et al.
Published: (2025)
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
by: Ishibashi, Yoichi, et al.
Published: (2025)
by: Ishibashi, Yoichi, et al.
Published: (2025)
LLM Pretraining with Continuous Concepts
by: Tack, Jihoon, et al.
Published: (2025)
by: Tack, Jihoon, et al.
Published: (2025)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Training-free LLM-generated Text Detection by Mining Token Probability Sequences
by: Xu, Yihuai, et al.
Published: (2024)
by: Xu, Yihuai, et al.
Published: (2024)
MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
by: Dong, Haoyu, et al.
Published: (2025)
by: Dong, Haoyu, et al.
Published: (2025)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
by: Wang, Yidong, et al.
Published: (2023)
by: Wang, Yidong, et al.
Published: (2023)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
by: Lu, Liming, et al.
Published: (2026)
by: Lu, Liming, et al.
Published: (2026)
RADAR: Reasoning as Discrimination with Aligned Representations for LLM-based Knowledge Graph Reasoning
by: Xue, Bo, et al.
Published: (2026)
by: Xue, Bo, et al.
Published: (2026)
SPOT: Span-level Pause-of-Thought for Efficient and Interpretable Latent Reasoning in Large Language Models
by: Chu, Yunlong, et al.
Published: (2026)
by: Chu, Yunlong, et al.
Published: (2026)
GeoGalactica: A Scientific Large Language Model in Geoscience
by: Lin, Zhouhan, et al.
Published: (2023)
by: Lin, Zhouhan, et al.
Published: (2023)
Checkpoint Merging via Bayesian Optimization in LLM Pretraining
by: Liu, Deyuan, et al.
Published: (2024)
by: Liu, Deyuan, et al.
Published: (2024)
KG-BiLM: Knowledge Graph Embedding via Bidirectional Language Models
by: Chen, Zirui, et al.
Published: (2025)
by: Chen, Zirui, et al.
Published: (2025)
Similar Items
-
PonderLM: Pretraining Language Models to Ponder in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025) -
PonderLM-3: Adaptive Token-Wise Pondering with Differentiable Masking
by: Li, He, et al.
Published: (2026) -
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
by: Song, Shixiang, et al.
Published: (2026) -
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
by: Zeng, Boyi, et al.
Published: (2026) -
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
by: Zeng, Boyi, et al.
Published: (2025)