Saved in:
| Main Authors: | Li, He, Song, Feichen, Zeng, Boyi, Song, Shixiang, Xu, Zhiqin John, He, Ziwei, Lin, Zhouhan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.02023 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
by: Song, Shixiang, et al.
Published: (2026)
by: Song, Shixiang, et al.
Published: (2026)
PonderLM: Pretraining Language Models to Ponder in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
by: Zeng, Boyi, et al.
Published: (2026)
by: Zeng, Boyi, et al.
Published: (2026)
Learning to Ponder: Adaptive Reasoning in Latent Space
by: He, Yixin, et al.
Published: (2025)
by: He, Yixin, et al.
Published: (2025)
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
by: Lu, Liming, et al.
Published: (2026)
by: Lu, Liming, et al.
Published: (2026)
When to Ponder: Adaptive Compute Allocation for Code Generation via Test-Time Training
by: Sim, Gihyeon
Published: (2025)
by: Sim, Gihyeon
Published: (2025)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
by: Kai, Jushi, et al.
Published: (2025)
by: Kai, Jushi, et al.
Published: (2025)
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
by: Shen, Junzhe, et al.
Published: (2026)
by: Shen, Junzhe, et al.
Published: (2026)
VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
by: Wang, Yixuan, et al.
Published: (2026)
by: Wang, Yixuan, et al.
Published: (2026)
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
by: Pan, Yanjie, et al.
Published: (2025)
by: Pan, Yanjie, et al.
Published: (2025)
FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation
by: Zhang, Runzhe, et al.
Published: (2026)
by: Zhang, Runzhe, et al.
Published: (2026)
Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows
by: Lehmann, Fabian, et al.
Published: (2024)
by: Lehmann, Fabian, et al.
Published: (2024)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
by: Zhu, Haoyi, et al.
Published: (2023)
by: Zhu, Haoyi, et al.
Published: (2023)
HuRef: HUman-REadable Fingerprint for Large Language Models
by: Zeng, Boyi, et al.
Published: (2023)
by: Zeng, Boyi, et al.
Published: (2023)
Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
by: He, Ziwei, et al.
Published: (2023)
by: He, Ziwei, et al.
Published: (2023)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models
by: Zheng, Kangjie, et al.
Published: (2025)
by: Zheng, Kangjie, et al.
Published: (2025)
Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM
by: Liu, Mengjie, et al.
Published: (2025)
by: Liu, Mengjie, et al.
Published: (2025)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
by: Cao, Guiming, et al.
Published: (2024)
by: Cao, Guiming, et al.
Published: (2024)
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
Training-free LLM-generated Text Detection by Mining Token Probability Sequences
by: Xu, Yihuai, et al.
Published: (2024)
by: Xu, Yihuai, et al.
Published: (2024)
Replicating ReLM Results: Validating Large Language Models with ReLM
by: Adamson, Reece, et al.
Published: (2025)
by: Adamson, Reece, et al.
Published: (2025)
Towards Controlled Table-to-Text Generation with Scientific Reasoning
by: Guo, Zhixin, et al.
Published: (2023)
by: Guo, Zhixin, et al.
Published: (2023)
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
by: Cai, Zhuoxuan, et al.
Published: (2025)
by: Cai, Zhuoxuan, et al.
Published: (2025)
Dynamic Pondering Sparsity-aware Mixture-of-Experts Transformer for Event Stream based Visual Object Tracking
by: Wang, Shiao, et al.
Published: (2026)
by: Wang, Shiao, et al.
Published: (2026)
Rewiring the Transformer with Depth-Wise LSTMs
by: Xu, Hongfei, et al.
Published: (2020)
by: Xu, Hongfei, et al.
Published: (2020)
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
by: He, Yexiao, et al.
Published: (2025)
by: He, Yexiao, et al.
Published: (2025)
TreeRare: Syntax Tree-Guided Retrieval and Reasoning for Knowledge-Intensive Question Answering
by: Zhang, Boyi, et al.
Published: (2025)
by: Zhang, Boyi, et al.
Published: (2025)
Token Masking Improves Transformer-Based Text Classification
by: Xu, Xianglong, et al.
Published: (2025)
by: Xu, Xianglong, et al.
Published: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
by: Taniguchi, Rei, et al.
Published: (2026)
by: Taniguchi, Rei, et al.
Published: (2026)
AntLM: Bridging Causal and Masked Language Models
by: Yu, Xinru, et al.
Published: (2024)
by: Yu, Xinru, et al.
Published: (2024)
How to Alleviate Catastrophic Forgetting in LLMs Finetuning? Hierarchical Layer-Wise and Element-Wise Regularization
by: Song, Shezheng, et al.
Published: (2025)
by: Song, Shezheng, et al.
Published: (2025)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
by: Li, Xue, et al.
Published: (2025)
by: Li, Xue, et al.
Published: (2025)
BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous Diffusion
by: Zhuang, Shaobin, et al.
Published: (2026)
by: Zhuang, Shaobin, et al.
Published: (2026)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
by: Jo, Daejin, et al.
Published: (2025)
by: Jo, Daejin, et al.
Published: (2025)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
by: Zuo, Youhui, et al.
Published: (2025)
by: Zuo, Youhui, et al.
Published: (2025)
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
by: Wang, Huanyu, et al.
Published: (2025)
by: Wang, Huanyu, et al.
Published: (2025)
Gumbel Reranking: Differentiable End-to-End Reranker Optimization
by: Huang, Siyuan, et al.
Published: (2025)
by: Huang, Siyuan, et al.
Published: (2025)
Similar Items
-
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
by: Song, Shixiang, et al.
Published: (2026) -
PonderLM: Pretraining Language Models to Ponder in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025) -
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025) -
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
by: Zeng, Boyi, et al.
Published: (2026) -
Learning to Ponder: Adaptive Reasoning in Latent Space
by: He, Yixin, et al.
Published: (2025)