The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, Tu, Zimmer, Matthieu, Tutunov, Rasul, Ji, Xiaotong, Ammar, Haitham Bou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
von: Bourigault, Pauline, et al.
Veröffentlicht: (2026)
von: Bourigault, Pauline, et al.
Veröffentlicht: (2026)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Model-Based and Sample-Efficient AI-Assisted Math Discovery in Sphere Packing
von: Tutunov, Rasul, et al.
Veröffentlicht: (2025)
von: Tutunov, Rasul, et al.
Veröffentlicht: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025)
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025)
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
von: Tutnov, Rasul, et al.
Veröffentlicht: (2025)
von: Tutnov, Rasul, et al.
Veröffentlicht: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
von: Fountas, Zafeirios, et al.
Veröffentlicht: (2026)
von: Fountas, Zafeirios, et al.
Veröffentlicht: (2026)
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
von: Diwan, Anish, et al.
Veröffentlicht: (2026)
von: Diwan, Anish, et al.
Veröffentlicht: (2026)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
von: Wieser, Frederico, et al.
Veröffentlicht: (2025)
von: Wieser, Frederico, et al.
Veröffentlicht: (2025)
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
von: Hazard, Hugo, et al.
Veröffentlicht: (2025)
von: Hazard, Hugo, et al.
Veröffentlicht: (2025)
Path-Guided Particle-based Sampling
von: Fan, Mingzhou, et al.
Veröffentlicht: (2024)
von: Fan, Mingzhou, et al.
Veröffentlicht: (2024)
Al-Khwarizmi: Discovering Physical Laws with Foundation Models
von: Mower, Christopher E., et al.
Veröffentlicht: (2025)
von: Mower, Christopher E., et al.
Veröffentlicht: (2025)
FASTER: Value-Guided Sampling for Fast RL
von: Dong, Perry, et al.
Veröffentlicht: (2026)
von: Dong, Perry, et al.
Veröffentlicht: (2026)
Why Can Large Language Models Generate Correct Chain-of-Thoughts?
von: Tutunov, Rasul, et al.
Veröffentlicht: (2023)
von: Tutunov, Rasul, et al.
Veröffentlicht: (2023)
Human-inspired Episodic Memory for Infinite Context LLMs
von: Fountas, Zafeirios, et al.
Veröffentlicht: (2024)
von: Fountas, Zafeirios, et al.
Veröffentlicht: (2024)
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
von: Liu, Zhenhua, et al.
Veröffentlicht: (2025)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2025)
Sampling for Quality: Training-Free Reward-Guided LLM Decoding via Sequential Monte Carlo
von: Markovic-Voronov, Jelena, et al.
Veröffentlicht: (2026)
von: Markovic-Voronov, Jelena, et al.
Veröffentlicht: (2026)
Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling
von: Liu, Yuejiang, et al.
Veröffentlicht: (2024)
von: Liu, Yuejiang, et al.
Veröffentlicht: (2024)
Spend Search Where It Pays: Value-Guided Structured Sampling and Optimization for Generative Recommendation
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
PowerPM: Foundation Model for Power Systems
von: Tu, Shihao, et al.
Veröffentlicht: (2024)
von: Tu, Shihao, et al.
Veröffentlicht: (2024)
Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding
von: Li, Xiner, et al.
Veröffentlicht: (2024)
von: Li, Xiner, et al.
Veröffentlicht: (2024)
When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning
von: Hao, Chenjie, et al.
Veröffentlicht: (2026)
von: Hao, Chenjie, et al.
Veröffentlicht: (2026)
Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
von: Du, Xiaodan, et al.
Veröffentlicht: (2023)
von: Du, Xiaodan, et al.
Veröffentlicht: (2023)
VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models
von: Xu, Hefei, et al.
Veröffentlicht: (2026)
von: Xu, Hefei, et al.
Veröffentlicht: (2026)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
von: Park, Young-Jin, et al.
Veröffentlicht: (2025)
von: Park, Young-Jin, et al.
Veröffentlicht: (2025)
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
von: Zimmer, Max, et al.
Veröffentlicht: (2026)
von: Zimmer, Max, et al.
Veröffentlicht: (2026)
Epistemic Deep Learning: Enabling Machine Learning Models to Know When They Do Not Know
von: Manchingal, Shireen Kudukkil
Veröffentlicht: (2025)
von: Manchingal, Shireen Kudukkil
Veröffentlicht: (2025)
Ark: An Open-source Python-based Framework for Robot Learning
von: Dierking, Magnus, et al.
Veröffentlicht: (2025)
von: Dierking, Magnus, et al.
Veröffentlicht: (2025)
Optimizing Deep Neural Networks using Safety-Guided Self Compression
von: Zbeeb, Mohammad, et al.
Veröffentlicht: (2025)
von: Zbeeb, Mohammad, et al.
Veröffentlicht: (2025)
Fine-Tuning Diffusion Models for Molecular Generation via Reinforcement Learning and Fast Sampling
von: Lin, Guang, et al.
Veröffentlicht: (2026)
von: Lin, Guang, et al.
Veröffentlicht: (2026)
Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging
von: Zimmer, Max, et al.
Veröffentlicht: (2023)
von: Zimmer, Max, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026) -
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026) -
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
von: Roy, Amartya, et al.
Veröffentlicht: (2026) -
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025) -
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)