Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer
Fuente:
arXiv
Saved in:
| Main Authors: | Kuang, Penghao, Wu, Haoyi, Tu, Kewei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
by: Wu, You, et al.
Published: (2024)
by: Wu, You, et al.
Published: (2024)
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
by: Wu, Haoyi, et al.
Published: (2024)
by: Wu, Haoyi, et al.
Published: (2024)
Parallel Continuous Chain-of-Thought with Jacobi Iteration
by: Wu, Haoyi, et al.
Published: (2025)
by: Wu, Haoyi, et al.
Published: (2025)
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale
by: Hu, Xiang, et al.
Published: (2024)
by: Hu, Xiang, et al.
Published: (2024)
YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference
by: Wu, You, et al.
Published: (2026)
by: Wu, You, et al.
Published: (2026)
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
by: Peng, Junjie, et al.
Published: (2026)
by: Peng, Junjie, et al.
Published: (2026)
Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language Models
by: Zhao, Yida, et al.
Published: (2024)
by: Zhao, Yida, et al.
Published: (2024)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
by: Lou, Chao, et al.
Published: (2024)
by: Lou, Chao, et al.
Published: (2024)
Augmenting Transformers with Recursively Composed Multi-grained Representations
by: Hu, Xiang, et al.
Published: (2023)
by: Hu, Xiang, et al.
Published: (2023)
Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
by: Hu, Xiang, et al.
Published: (2024)
by: Hu, Xiang, et al.
Published: (2024)
GiLT: Augmenting Transformer Language Models with Dependency Graphs
by: Huang, Tianyu, et al.
Published: (2026)
by: Huang, Tianyu, et al.
Published: (2026)
RoT: Enhancing Large Language Models with Reflection on Search Trees
by: Hui, Wenyang, et al.
Published: (2024)
by: Hui, Wenyang, et al.
Published: (2024)
A Systematic Study of Compositional Syntactic Transformer Language Models
by: Zhao, Yida, et al.
Published: (2025)
by: Zhao, Yida, et al.
Published: (2025)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access
by: Hu, Xiang, et al.
Published: (2025)
by: Hu, Xiang, et al.
Published: (2025)
Parallel Loop Transformer for Efficient Test-Time Computation Scaling
by: Wu, Bohong, et al.
Published: (2025)
by: Wu, Bohong, et al.
Published: (2025)
Using Interpretation Methods for Model Enhancement
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
by: Kuang, Peng, et al.
Published: (2025)
by: Kuang, Peng, et al.
Published: (2025)
Multilingual Test-Time Scaling via Initial Thought Transfer
by: Bajpai, Prasoon, et al.
Published: (2025)
by: Bajpai, Prasoon, et al.
Published: (2025)
OptScale: Probabilistic Optimality for Inference-time Scaling
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
Efficient Pretraining Length Scaling
by: Wu, Bohong, et al.
Published: (2025)
by: Wu, Bohong, et al.
Published: (2025)
A Probabilistic Inference Scaling Theory for LLM Self-Correction
by: Yang, Zhe, et al.
Published: (2025)
by: Yang, Zhe, et al.
Published: (2025)
Efficient Multimodal Planning Agent for Visual Question-Answering
by: Chen, Zhuo, et al.
Published: (2026)
by: Chen, Zhuo, et al.
Published: (2026)
LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language Models
by: Song, Haoyi, et al.
Published: (2025)
by: Song, Haoyi, et al.
Published: (2025)
Zero-Shot Performance Prediction for Probabilistic Scaling Laws
by: Schram, Viktoria, et al.
Published: (2025)
by: Schram, Viktoria, et al.
Published: (2025)
Efficiently Aligned Cross-Lingual Transfer Learning for Conversational Tasks using Prompt-Tuning
by: Tu, Lifu, et al.
Published: (2023)
by: Tu, Lifu, et al.
Published: (2023)
HRM-Text: Efficient Pretraining Beyond Scaling
by: Wang, Guan, et al.
Published: (2026)
by: Wang, Guan, et al.
Published: (2026)
Z1: Efficient Test-time Scaling with Code
by: Yu, Zhaojian, et al.
Published: (2025)
by: Yu, Zhaojian, et al.
Published: (2025)
Unsupervised Morphological Tree Tokenizer
by: Zhu, Qingyang, et al.
Published: (2024)
by: Zhu, Qingyang, et al.
Published: (2024)
Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models
by: Tan, Yuqiao, et al.
Published: (2025)
by: Tan, Yuqiao, et al.
Published: (2025)
What Scales in Cross-Entropy Scaling Law?
by: Yan, Junxi, et al.
Published: (2025)
by: Yan, Junxi, et al.
Published: (2025)
Exploring the Potential of Probabilistic Transformer for Time Series Modeling: A Report on the ST-PT Framework
by: Xiong, Zhangzhi, et al.
Published: (2026)
by: Xiong, Zhangzhi, et al.
Published: (2026)
Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
by: Lialin, Vladislav, et al.
Published: (2023)
by: Lialin, Vladislav, et al.
Published: (2023)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
by: Ko, Jongwoo, et al.
Published: (2026)
by: Ko, Jongwoo, et al.
Published: (2026)
Efficient Data Selection at Scale via Influence Distillation
by: Nikdan, Mahdi, et al.
Published: (2025)
by: Nikdan, Mahdi, et al.
Published: (2025)
Language Fusion for Parameter-Efficient Cross-lingual Transfer
by: Borchert, Philipp, et al.
Published: (2025)
by: Borchert, Philipp, et al.
Published: (2025)
DeepDiver: Adaptive Search Intensity Scaling via Open-Web Reinforcement Learning
by: Shi, Wenxuan, et al.
Published: (2025)
by: Shi, Wenxuan, et al.
Published: (2025)
Scaling Efficient LLMs
by: Kausik, B. N.
Published: (2024)
by: Kausik, B. N.
Published: (2024)
Similar Items
-
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
by: Wu, You, et al.
Published: (2024) -
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
by: Wu, Haoyi, et al.
Published: (2024) -
Parallel Continuous Chain-of-Thought with Jacobi Iteration
by: Wu, Haoyi, et al.
Published: (2025) -
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale
by: Hu, Xiang, et al.
Published: (2024) -
YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference
by: Wu, You, et al.
Published: (2026)