Martingale Foresight Sampling: A Principled Approach to Inference-Time LLM Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Huayu, He, ZhengXiao, Tian, Siyuan, Wen, Jinghao, Li, Ao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NeuroHD-RA: Neural-distilled Hyperdimensional Model with Rhythm Alignment
by: He, ZhengXiao, et al.
Published: (2025)
by: He, ZhengXiao, et al.
Published: (2025)
MedMamba: Recasting Mamba for Medical Time Series Classification
by: He, ZhengXiao, et al.
Published: (2026)
by: He, ZhengXiao, et al.
Published: (2026)
Learning Fingerprints for Medical Time Series with Redundancy-Constrained Information Maximization
by: Li, Huayu, et al.
Published: (2026)
by: Li, Huayu, et al.
Published: (2026)
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
by: Xu, Fangzhi, et al.
Published: (2025)
by: Xu, Fangzhi, et al.
Published: (2025)
MTS-LOF: Medical Time-Series Representation Learning via Occlusion-Invariant Features
by: Li, Huayu, et al.
Published: (2023)
by: Li, Huayu, et al.
Published: (2023)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
by: Jeon, Wonseok, et al.
Published: (2024)
by: Jeon, Wonseok, et al.
Published: (2024)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
by: He, Zhonghao, et al.
Published: (2025)
by: He, Zhonghao, et al.
Published: (2025)
RAP: Runtime Adaptive Pruning for LLM Inference
by: Liu, Huanrong, et al.
Published: (2025)
by: Liu, Huanrong, et al.
Published: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
by: Wen, Zhuofan, et al.
Published: (2024)
by: Wen, Zhuofan, et al.
Published: (2024)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
by: Li, Ruixiao, et al.
Published: (2025)
by: Li, Ruixiao, et al.
Published: (2025)
Multimodal Cardiovascular Risk Profiling Using Self-Supervised Learning of Polysomnography
by: He, Zhengxiao, et al.
Published: (2025)
by: He, Zhengxiao, et al.
Published: (2025)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
by: Dong, Yanhao, et al.
Published: (2025)
by: Dong, Yanhao, et al.
Published: (2025)
Dr-CiK: A Testbed for Foresight-Driven Agents
by: Tang, Yihong, et al.
Published: (2026)
by: Tang, Yihong, et al.
Published: (2026)
Current Agents Fail to Leverage World Model as Tool for Foresight
by: Qian, Cheng, et al.
Published: (2026)
by: Qian, Cheng, et al.
Published: (2026)
Martingale-Consistent Self-Supervised Learning
by: Gögl, Moritz, et al.
Published: (2026)
by: Gögl, Moritz, et al.
Published: (2026)
EEG-EMG FAConformer: Frequency Aware Conv-Transformer for the fusion of EEG and EMG
by: He, ZhengXiao, et al.
Published: (2024)
by: He, ZhengXiao, et al.
Published: (2024)
Decocted Experience Improves Test-Time Inference in LLM Agents
by: Shen, Maohao, et al.
Published: (2026)
by: Shen, Maohao, et al.
Published: (2026)
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
by: Ye, Wen, et al.
Published: (2025)
by: Ye, Wen, et al.
Published: (2025)
Note on Martingale Theory and Applications
by: Zou, Xiandong
Published: (2026)
by: Zou, Xiandong
Published: (2026)
Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights
by: Parashar, Shubham, et al.
Published: (2025)
by: Parashar, Shubham, et al.
Published: (2025)
IdealTSF: Can Non-Ideal Data Contribute to Enhancing the Performance of Time Series Forecasting Models?
by: Wang, Hua, et al.
Published: (2025)
by: Wang, Hua, et al.
Published: (2025)
A Contextual Combinatorial Bandit Approach to Negotiation
by: Li, Yexin, et al.
Published: (2024)
by: Li, Yexin, et al.
Published: (2024)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
by: Trivedi, Prashant, et al.
Published: (2025)
by: Trivedi, Prashant, et al.
Published: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024)
by: He, Junhui, et al.
Published: (2024)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
Advancing time series completion via RFAMoE and MDFF
by: Zhang, Ci, et al.
Published: (2025)
by: Zhang, Ci, et al.
Published: (2025)
Thermodynamic Focusing for Inference-Time Search: Practical Methods for Target-Conditioned Sampling and Prompted Inference
by: Zhang, Zhan
Published: (2025)
by: Zhang, Zhan
Published: (2025)
Sampling for Quality: Training-Free Reward-Guided LLM Decoding via Sequential Monte Carlo
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
RelayLLM: Efficient Reasoning via Collaborative Decoding
by: Huang, Chengsong, et al.
Published: (2026)
by: Huang, Chengsong, et al.
Published: (2026)
WebLLM: A High-Performance In-Browser LLM Inference Engine
by: Ruan, Charlie F., et al.
Published: (2024)
by: Ruan, Charlie F., et al.
Published: (2024)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
by: Zhu, Xiao, et al.
Published: (2026)
by: Zhu, Xiao, et al.
Published: (2026)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
SDMixer: Sparse Dual-Mixer for Time Series Forecasting
by: Ao, Xiang
Published: (2026)
by: Ao, Xiang
Published: (2026)
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
by: Liu, Zhihan, et al.
Published: (2023)
by: Liu, Zhihan, et al.
Published: (2023)
Adaptive Layer Splitting for Wireless LLM Inference in Edge Computing: A Model-Based Reinforcement Learning Approach
by: Chen, Yuxuan, et al.
Published: (2024)
by: Chen, Yuxuan, et al.
Published: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
by: Timor, Nadav, et al.
Published: (2025)
by: Timor, Nadav, et al.
Published: (2025)
The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning
by: Zhuang, Ren, et al.
Published: (2026)
by: Zhuang, Ren, et al.
Published: (2026)
WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales
by: Prinster, Drew, et al.
Published: (2025)
by: Prinster, Drew, et al.
Published: (2025)
A Versatile Graph Learning Approach through LLM-based Agent
by: Wei, Lanning, et al.
Published: (2023)
by: Wei, Lanning, et al.
Published: (2023)
Similar Items
-
NeuroHD-RA: Neural-distilled Hyperdimensional Model with Rhythm Alignment
by: He, ZhengXiao, et al.
Published: (2025) -
MedMamba: Recasting Mamba for Medical Time Series Classification
by: He, ZhengXiao, et al.
Published: (2026) -
Learning Fingerprints for Medical Time Series with Redundancy-Constrained Information Maximization
by: Li, Huayu, et al.
Published: (2026) -
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
by: Xu, Fangzhi, et al.
Published: (2025) -
MTS-LOF: Medical Time-Series Representation Learning via Occlusion-Invariant Features
by: Li, Huayu, et al.
Published: (2023)