Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Shaohua, Huang, Pengcheng, Li, Xinze, Liu, Zhenghao, Yi, Xiaoyuan, Yan, Yukun, Wang, Shuo, Gu, Yu, Yu, Ge, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
by: Liu, Zhenghao, et al.
Published: (2026)
by: Liu, Zhenghao, et al.
Published: (2026)
MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization
by: Xin, Haidong, et al.
Published: (2026)
by: Xin, Haidong, et al.
Published: (2026)
Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
by: Liu, Zhenghao, et al.
Published: (2025)
by: Liu, Zhenghao, et al.
Published: (2025)
Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
by: Wu, Zhuoyang, et al.
Published: (2025)
by: Wu, Zhuoyang, et al.
Published: (2025)
Revealing the Attention Floating Mechanism in Masked Diffusion Models
by: Dai, Xin, et al.
Published: (2026)
by: Dai, Xin, et al.
Published: (2026)
HIPPO: Enhancing the Table Understanding Capability of LLMs through Hybrid-Modal Preference Optimization
by: Wang, Haolan, et al.
Published: (2025)
by: Wang, Haolan, et al.
Published: (2025)
ExpandR: Teaching Dense Retrievers Beyond Queries with LLM Guidance
by: Yao, Sijia, et al.
Published: (2025)
by: Yao, Sijia, et al.
Published: (2025)
Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
by: Peng, Chunyi, et al.
Published: (2025)
by: Peng, Chunyi, et al.
Published: (2025)
ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
by: Jin, Zhensheng, et al.
Published: (2025)
by: Jin, Zhensheng, et al.
Published: (2025)
Structured Knowledge Representation through Contextual Pages for Retrieval-Augmented Generation
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
Legal$Δ$: Enhancing Legal Reasoning in LLMs via Reinforcement Learning with Chain-of-Thought Guided Information Gain
by: Dai, Xin, et al.
Published: (2025)
by: Dai, Xin, et al.
Published: (2025)
Finding What Matters: Anchoring Context Knowledge with Evolving Indices for Iterative Retrieval
by: Wu, Mingyan, et al.
Published: (2026)
by: Wu, Mingyan, et al.
Published: (2026)
Say More with Less: Understanding Prompt Learning Behaviors through Gist Compression
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
Building A Coding Assistant via the Retrieval-Augmented Language Model
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
by: Huang, Pengcheng, et al.
Published: (2025)
by: Huang, Pengcheng, et al.
Published: (2025)
RankCoT: Refining Knowledge for Retrieval-Augmented Generation through Ranking Chain-of-Thoughts
by: Wu, Mingyan, et al.
Published: (2025)
by: Wu, Mingyan, et al.
Published: (2025)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis
by: Yang, Weiqing, et al.
Published: (2024)
by: Yang, Weiqing, et al.
Published: (2024)
LISRec: Modeling User Preferences with Learned Item Shortcuts for Sequential Recommendation
by: Xin, Haidong, et al.
Published: (2025)
by: Xin, Haidong, et al.
Published: (2025)
Context Attribution with Multi-Armed Bandit Optimization
by: Pan, Deng, et al.
Published: (2025)
by: Pan, Deng, et al.
Published: (2025)
The Unreasonable Effectiveness of Greedy Algorithms in Multi-Armed Bandit with Many Arms
by: Bayati, Mohsen, et al.
Published: (2020)
by: Bayati, Mohsen, et al.
Published: (2020)
LegalDuet: Learning Fine-grained Representations for Legal Judgment Prediction via a Dual-View Contrastive Learning
by: Xu, Buqiang, et al.
Published: (2024)
by: Xu, Buqiang, et al.
Published: (2024)
Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
by: Gao, Xingjie, et al.
Published: (2026)
by: Gao, Xingjie, et al.
Published: (2026)
Meet Me at the Arm: The Cooperative Multi-Armed Bandits Problem with Shareable Arms
by: Hu, Xinyi, et al.
Published: (2025)
by: Hu, Xinyi, et al.
Published: (2025)
Empirical Analysis of Decoding Biases in Masked Diffusion Models
by: Huang, Pengcheng, et al.
Published: (2025)
by: Huang, Pengcheng, et al.
Published: (2025)
Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
by: Bai, Yuzhuo, et al.
Published: (2025)
by: Bai, Yuzhuo, et al.
Published: (2025)
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Learning Refined Document Representations for Dense Retrieval via Deliberate Thinking
by: Ji, Yifan, et al.
Published: (2025)
by: Ji, Yifan, et al.
Published: (2025)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
by: Sun, Yubo, et al.
Published: (2025)
by: Sun, Yubo, et al.
Published: (2025)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
by: Li, Weizhuo, et al.
Published: (2024)
by: Li, Weizhuo, et al.
Published: (2024)
KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs
by: Wu, Dingjun, et al.
Published: (2025)
by: Wu, Dingjun, et al.
Published: (2025)
Multi-Evidence based Fact Verification via A Confidential Graph Neural Network
by: Lan, Yuqing, et al.
Published: (2024)
by: Lan, Yuqing, et al.
Published: (2024)
Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains
by: Xiong, Yuqi, et al.
Published: (2026)
by: Xiong, Yuqi, et al.
Published: (2026)
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
by: Ran, Dezhi, et al.
Published: (2025)
by: Ran, Dezhi, et al.
Published: (2025)
Similar Items
-
Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
by: Liu, Zhenghao, et al.
Published: (2026) -
MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization
by: Xin, Haidong, et al.
Published: (2026) -
Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
by: Liu, Zhenghao, et al.
Published: (2025) -
Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
by: Wu, Zhuoyang, et al.
Published: (2025) -
Revealing the Attention Floating Mechanism in Masked Diffusion Models
by: Dai, Xin, et al.
Published: (2026)