Guardado en:
| Autores principales: | Li, He, Song, Feichen, Zeng, Boyi, Song, Shixiang, Xu, Zhiqin John, He, Ziwei, Lin, Zhouhan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.02023 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
por: Song, Shixiang, et al.
Publicado: (2026)
por: Song, Shixiang, et al.
Publicado: (2026)
PonderLM: Pretraining Language Models to Ponder in Continuous Space
por: Zeng, Boyi, et al.
Publicado: (2025)
por: Zeng, Boyi, et al.
Publicado: (2025)
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
por: Zeng, Boyi, et al.
Publicado: (2025)
por: Zeng, Boyi, et al.
Publicado: (2025)
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
por: Zeng, Boyi, et al.
Publicado: (2026)
por: Zeng, Boyi, et al.
Publicado: (2026)
Learning to Ponder: Adaptive Reasoning in Latent Space
por: He, Yixin, et al.
Publicado: (2025)
por: He, Yixin, et al.
Publicado: (2025)
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
por: Zeng, Boyi, et al.
Publicado: (2025)
por: Zeng, Boyi, et al.
Publicado: (2025)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
por: Lu, Liming, et al.
Publicado: (2026)
por: Lu, Liming, et al.
Publicado: (2026)
When to Ponder: Adaptive Compute Allocation for Code Generation via Test-Time Training
por: Sim, Gihyeon
Publicado: (2025)
por: Sim, Gihyeon
Publicado: (2025)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
por: Kai, Jushi, et al.
Publicado: (2025)
por: Kai, Jushi, et al.
Publicado: (2025)
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
por: Shen, Junzhe, et al.
Publicado: (2026)
por: Shen, Junzhe, et al.
Publicado: (2026)
VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
por: Wang, Yixuan, et al.
Publicado: (2026)
por: Wang, Yixuan, et al.
Publicado: (2026)
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
por: Pan, Yanjie, et al.
Publicado: (2025)
por: Pan, Yanjie, et al.
Publicado: (2025)
FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation
por: Zhang, Runzhe, et al.
Publicado: (2026)
por: Zhang, Runzhe, et al.
Publicado: (2026)
Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows
por: Lehmann, Fabian, et al.
Publicado: (2024)
por: Lehmann, Fabian, et al.
Publicado: (2024)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
por: Zhu, Haoyi, et al.
Publicado: (2023)
por: Zhu, Haoyi, et al.
Publicado: (2023)
HuRef: HUman-REadable Fingerprint for Large Language Models
por: Zeng, Boyi, et al.
Publicado: (2023)
por: Zeng, Boyi, et al.
Publicado: (2023)
Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
por: He, Ziwei, et al.
Publicado: (2023)
por: He, Ziwei, et al.
Publicado: (2023)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
por: Oba, Daisuke, et al.
Publicado: (2026)
por: Oba, Daisuke, et al.
Publicado: (2026)
ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models
por: Zheng, Kangjie, et al.
Publicado: (2025)
por: Zheng, Kangjie, et al.
Publicado: (2025)
Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM
por: Liu, Mengjie, et al.
Publicado: (2025)
por: Liu, Mengjie, et al.
Publicado: (2025)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
por: Cao, Guiming, et al.
Publicado: (2024)
por: Cao, Guiming, et al.
Publicado: (2024)
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
por: Wang, Yiqin, et al.
Publicado: (2024)
por: Wang, Yiqin, et al.
Publicado: (2024)
Training-free LLM-generated Text Detection by Mining Token Probability Sequences
por: Xu, Yihuai, et al.
Publicado: (2024)
por: Xu, Yihuai, et al.
Publicado: (2024)
Replicating ReLM Results: Validating Large Language Models with ReLM
por: Adamson, Reece, et al.
Publicado: (2025)
por: Adamson, Reece, et al.
Publicado: (2025)
Towards Controlled Table-to-Text Generation with Scientific Reasoning
por: Guo, Zhixin, et al.
Publicado: (2023)
por: Guo, Zhixin, et al.
Publicado: (2023)
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
por: Cai, Zhuoxuan, et al.
Publicado: (2025)
por: Cai, Zhuoxuan, et al.
Publicado: (2025)
Dynamic Pondering Sparsity-aware Mixture-of-Experts Transformer for Event Stream based Visual Object Tracking
por: Wang, Shiao, et al.
Publicado: (2026)
por: Wang, Shiao, et al.
Publicado: (2026)
Rewiring the Transformer with Depth-Wise LSTMs
por: Xu, Hongfei, et al.
Publicado: (2020)
por: Xu, Hongfei, et al.
Publicado: (2020)
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
por: He, Yexiao, et al.
Publicado: (2025)
por: He, Yexiao, et al.
Publicado: (2025)
TreeRare: Syntax Tree-Guided Retrieval and Reasoning for Knowledge-Intensive Question Answering
por: Zhang, Boyi, et al.
Publicado: (2025)
por: Zhang, Boyi, et al.
Publicado: (2025)
Token Masking Improves Transformer-Based Text Classification
por: Xu, Xianglong, et al.
Publicado: (2025)
por: Xu, Xianglong, et al.
Publicado: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
por: Taniguchi, Rei, et al.
Publicado: (2026)
por: Taniguchi, Rei, et al.
Publicado: (2026)
AntLM: Bridging Causal and Masked Language Models
por: Yu, Xinru, et al.
Publicado: (2024)
por: Yu, Xinru, et al.
Publicado: (2024)
How to Alleviate Catastrophic Forgetting in LLMs Finetuning? Hierarchical Layer-Wise and Element-Wise Regularization
por: Song, Shezheng, et al.
Publicado: (2025)
por: Song, Shezheng, et al.
Publicado: (2025)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
por: Li, Xue, et al.
Publicado: (2025)
por: Li, Xue, et al.
Publicado: (2025)
BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous Diffusion
por: Zhuang, Shaobin, et al.
Publicado: (2026)
por: Zhuang, Shaobin, et al.
Publicado: (2026)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
por: Jo, Daejin, et al.
Publicado: (2025)
por: Jo, Daejin, et al.
Publicado: (2025)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
por: Zuo, Youhui, et al.
Publicado: (2025)
por: Zuo, Youhui, et al.
Publicado: (2025)
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
por: Wang, Huanyu, et al.
Publicado: (2025)
por: Wang, Huanyu, et al.
Publicado: (2025)
Gumbel Reranking: Differentiable End-to-End Reranker Optimization
por: Huang, Siyuan, et al.
Publicado: (2025)
por: Huang, Siyuan, et al.
Publicado: (2025)
Ejemplares similares
-
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
por: Song, Shixiang, et al.
Publicado: (2026) -
PonderLM: Pretraining Language Models to Ponder in Continuous Space
por: Zeng, Boyi, et al.
Publicado: (2025) -
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
por: Zeng, Boyi, et al.
Publicado: (2025) -
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
por: Zeng, Boyi, et al.
Publicado: (2026) -
Learning to Ponder: Adaptive Reasoning in Latent Space
por: He, Yixin, et al.
Publicado: (2025)