FlashSampling: Fast and Memory-Efficient Exact Sampling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ruiz, Tomas, Qin, Zhen, Zhang, Yifan, Shen, Xuyang, Zhong, Yiran, Wang, Mengdi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Higher-order Linear Attention
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
von: Wang, Ziqing, et al.
Veröffentlicht: (2026)
von: Wang, Ziqing, et al.
Veröffentlicht: (2026)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024)
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024)
AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tuning
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principles
von: Wei, Qingyan, et al.
Veröffentlicht: (2025)
von: Wei, Qingyan, et al.
Veröffentlicht: (2025)
Sample-Efficient Alignment for LLMs
von: Liu, Zichen, et al.
Veröffentlicht: (2024)
von: Liu, Zichen, et al.
Veröffentlicht: (2024)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
von: Pan, Rui, et al.
Veröffentlicht: (2024)
von: Pan, Rui, et al.
Veröffentlicht: (2024)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
von: Zhang, Kongcheng, et al.
Veröffentlicht: (2025)
von: Zhang, Kongcheng, et al.
Veröffentlicht: (2025)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
von: Sharma, Agniv, et al.
Veröffentlicht: (2024)
von: Sharma, Agniv, et al.
Veröffentlicht: (2024)
Fast Controlled Generation from Language Models with Adaptive Weighted Rejection Sampling
von: Lipkin, Benjamin, et al.
Veröffentlicht: (2025)
von: Lipkin, Benjamin, et al.
Veröffentlicht: (2025)
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
Iterative Deepening Sampling as Efficient Test-Time Scaling
von: Chen, Weizhe, et al.
Veröffentlicht: (2025)
von: Chen, Weizhe, et al.
Veröffentlicht: (2025)
Hallucination Detection in LLMs: Fast and Memory-Efficient Fine-Tuned Models
von: Arteaga, Gabriel Y., et al.
Veröffentlicht: (2024)
von: Arteaga, Gabriel Y., et al.
Veröffentlicht: (2024)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
PRISM: Parametrically Refactoring Inference for Speculative Sampling Draft Models
von: Wang, Xuliang, et al.
Veröffentlicht: (2026)
von: Wang, Xuliang, et al.
Veröffentlicht: (2026)
MemRerank: Preference Memory for Personalized Product Reranking
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2026)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
von: Hu, Haiquan, et al.
Veröffentlicht: (2025)
von: Hu, Haiquan, et al.
Veröffentlicht: (2025)
OpenClaw-RL: Train Any Agent Simply by Talking
von: Wang, Yinjie, et al.
Veröffentlicht: (2026)
von: Wang, Yinjie, et al.
Veröffentlicht: (2026)
A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
von: Yuan, Hui, et al.
Veröffentlicht: (2024)
von: Yuan, Hui, et al.
Veröffentlicht: (2024)
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
von: Wang, Fei, et al.
Veröffentlicht: (2025)
von: Wang, Fei, et al.
Veröffentlicht: (2025)
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
von: Wang, Zhanliang, et al.
Veröffentlicht: (2026)
von: Wang, Zhanliang, et al.
Veröffentlicht: (2026)
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
Sample-aware Adaptive Structured Pruning for Large Language Models
von: Kong, Jun, et al.
Veröffentlicht: (2025)
von: Kong, Jun, et al.
Veröffentlicht: (2025)
Deep Delta Learning
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
SLOT: Sample-specific Language Model Optimization at Test-time
von: Hu, Yang, et al.
Veröffentlicht: (2025)
von: Hu, Yang, et al.
Veröffentlicht: (2025)
GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
von: Zhang, Yunan, et al.
Veröffentlicht: (2025)
von: Zhang, Yunan, et al.
Veröffentlicht: (2025)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
von: Guo, Jizhou, et al.
Veröffentlicht: (2025)
von: Guo, Jizhou, et al.
Veröffentlicht: (2025)
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step
von: Dugan, Owen, et al.
Veröffentlicht: (2024)
von: Dugan, Owen, et al.
Veröffentlicht: (2024)
Constrained Adaptive Rejection Sampling
von: Parys, Paweł, et al.
Veröffentlicht: (2025)
von: Parys, Paweł, et al.
Veröffentlicht: (2025)
Federated Data-Efficient Instruction Tuning for Large Language Models
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
von: Schröder, Christopher, et al.
Veröffentlicht: (2024)
von: Schröder, Christopher, et al.
Veröffentlicht: (2024)
LAMPO: Large Language Models as Preference Machines for Few-shot Ordinal Classification
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
Boosting Protein Language Models with Negative Sample Mining
von: Xu, Yaoyao, et al.
Veröffentlicht: (2024)
von: Xu, Yaoyao, et al.
Veröffentlicht: (2024)
Linear Attention Sequence Parallelism
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Higher-order Linear Attention
von: Zhang, Yifan, et al.
Veröffentlicht: (2025) -
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
von: Wang, Ziqing, et al.
Veröffentlicht: (2026) -
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024) -
AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tuning
von: Yang, Yifan, et al.
Veröffentlicht: (2024) -
Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principles
von: Wei, Qingyan, et al.
Veröffentlicht: (2025)