DISA: Offline Importance Sampling for Distribution-Matching LLM-RL
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shaobo, Chen, Yujie, Sun, Yafeng, Qiu, Wenjie, Xie, Zhihui, Li, Sihang, Li, Yucheng, Jiang, Huiqiang, Ren, Xingzhang, Hu, Xuming, Liu, Dayiheng, Zhang, Linfeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates
by: Wang, Shaobo, et al.
Published: (2026)
by: Wang, Shaobo, et al.
Published: (2026)
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
by: Zhou, Yufa, et al.
Published: (2025)
by: Zhou, Yufa, et al.
Published: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
by: Wang, Shaobo, et al.
Published: (2026)
by: Wang, Shaobo, et al.
Published: (2026)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
by: Li, Yucheng, et al.
Published: (2025)
by: Li, Yucheng, et al.
Published: (2025)
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
by: Tian, Yuxin, et al.
Published: (2024)
by: Tian, Yuxin, et al.
Published: (2024)
FlowRL: Matching Reward Distributions for LLM Reasoning
by: Zhu, Xuekai, et al.
Published: (2025)
by: Zhu, Xuekai, et al.
Published: (2025)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
by: Zhao, Anhao, et al.
Published: (2026)
by: Zhao, Anhao, et al.
Published: (2026)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Are Expressive Models Truly Necessary for Offline RL?
by: Wang, Guan, et al.
Published: (2024)
by: Wang, Guan, et al.
Published: (2024)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
AIS: Adaptive Importance Sampling for Quantized RL
by: Zhou, Jiajun, et al.
Published: (2026)
by: Zhou, Jiajun, et al.
Published: (2026)
Digital Imaging South Africa (DISA): A Case Study
by: Saunders, Christopher
Published: (2005)
by: Saunders, Christopher
Published: (2005)
Multi-Agent Path Finding via Offline RL and LLM Collaboration
by: Atasever, Merve, et al.
Published: (2025)
by: Atasever, Merve, et al.
Published: (2025)
Score-Regularized Joint Sampling with Importance Weights for Flow Matching
by: Liu, Xinshuang, et al.
Published: (2025)
by: Liu, Xinshuang, et al.
Published: (2025)
Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning
by: Li, Na, et al.
Published: (2025)
by: Li, Na, et al.
Published: (2025)
Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
Not All Samples Should Be Utilized Equally: Towards Understanding and Improving Dataset Distillation
by: Wang, Shaobo, et al.
Published: (2024)
by: Wang, Shaobo, et al.
Published: (2024)
Offline Map Matching Based on Localization Error Distribution Modeling
by: Xu, Ruilin, et al.
Published: (2025)
by: Xu, Ruilin, et al.
Published: (2025)
Gnothi Seauton: Empowering Faithful Self-Interpretability in Black-Box Transformers
by: Wang, Shaobo, et al.
Published: (2024)
by: Wang, Shaobo, et al.
Published: (2024)
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
by: Wang, Jiakang, et al.
Published: (2025)
by: Wang, Jiakang, et al.
Published: (2025)
Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift
by: Li, Bochao, et al.
Published: (2026)
by: Li, Bochao, et al.
Published: (2026)
Dataset Distillation with Neural Characteristic Function: A Minmax Perspective
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
Improving Zero-Shot Offline RL via Behavioral Task Sampling
by: Bendib, Nazim, et al.
Published: (2026)
by: Bendib, Nazim, et al.
Published: (2026)
AdamO: A Collapse-Suppressed Optimizer for Offline RL
by: Qiao, Nan, et al.
Published: (2026)
by: Qiao, Nan, et al.
Published: (2026)
Unified Data Selection for LLM Reasoning
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
Less is More: Clustered Cross-Covariance Control for Offline RL
by: Qiao, Nan, et al.
Published: (2026)
by: Qiao, Nan, et al.
Published: (2026)
Mitigating Distribution Shift in Model-based Offline RL via Shifts-aware Reward Learning
by: Luo, Wang, et al.
Published: (2024)
by: Luo, Wang, et al.
Published: (2024)
STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order
by: Gu, Chengyang, et al.
Published: (2026)
by: Gu, Chengyang, et al.
Published: (2026)
Uni-RCM: Unified Reference-guided Cross-modal Mapping for Multi-Class Anomaly Detection
by: Wu, Yangchen, et al.
Published: (2026)
by: Wu, Yangchen, et al.
Published: (2026)
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
by: Ning, Shan, et al.
Published: (2026)
by: Ning, Shan, et al.
Published: (2026)
Budgeting Counterfactual for Offline RL
by: Liu, Yao, et al.
Published: (2023)
by: Liu, Yao, et al.
Published: (2023)
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
by: Wang, Hanlin, et al.
Published: (2025)
by: Wang, Hanlin, et al.
Published: (2025)
Learning from Random Demonstrations: Offline Reinforcement Learning with Importance-Sampled Diffusion Models
by: Fang, Zeyu, et al.
Published: (2024)
by: Fang, Zeyu, et al.
Published: (2024)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Similar Items
-
Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning
by: Wang, Shaobo, et al.
Published: (2025) -
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
by: Wang, Shaobo, et al.
Published: (2025) -
The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates
by: Wang, Shaobo, et al.
Published: (2026) -
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
by: Zhou, Yufa, et al.
Published: (2025) -
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)