ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Zhexin, Wang, Li, Wang, Xiaohan, Chai, Jiajun, Guo, Xiaojun, Lin, Wei, Yin, Guojun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
von: Wei, Zhenlin, et al.
Veröffentlicht: (2026)
von: Wei, Zhenlin, et al.
Veröffentlicht: (2026)
SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
von: Guo, Xiaojun, et al.
Veröffentlicht: (2025)
von: Guo, Xiaojun, et al.
Veröffentlicht: (2025)
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
von: Chai, Jiajun, et al.
Veröffentlicht: (2025)
von: Chai, Jiajun, et al.
Veröffentlicht: (2025)
MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL
von: Xu, Zekun, et al.
Veröffentlicht: (2025)
von: Xu, Zekun, et al.
Veröffentlicht: (2025)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning
von: Sun, Jingbo, et al.
Veröffentlicht: (2026)
von: Sun, Jingbo, et al.
Veröffentlicht: (2026)
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
von: Shianifar, Jonaid, et al.
Veröffentlicht: (2026)
von: Shianifar, Jonaid, et al.
Veröffentlicht: (2026)
GeneZip: Region-Aware Compression for Long Context DNA Modeling
von: Zhao, Jianan, et al.
Veröffentlicht: (2026)
von: Zhao, Jianan, et al.
Veröffentlicht: (2026)
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
von: Romio, Gabriel, et al.
Veröffentlicht: (2026)
von: Romio, Gabriel, et al.
Veröffentlicht: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
von: Ding, Liang
Veröffentlicht: (2026)
von: Ding, Liang
Veröffentlicht: (2026)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
Cluster-based Sampling in Hindsight Experience Replay for Robotic Tasks (Student Abstract)
von: Kim, Taeyoung, et al.
Veröffentlicht: (2022)
von: Kim, Taeyoung, et al.
Veröffentlicht: (2022)
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
von: Liu, Dengcan, et al.
Veröffentlicht: (2026)
von: Liu, Dengcan, et al.
Veröffentlicht: (2026)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
von: Wang, Ke, et al.
Veröffentlicht: (2024)
von: Wang, Ke, et al.
Veröffentlicht: (2024)
Promoting Efficient Reasoning with Verifiable Stepwise Reward
von: Yue, Chuhuai, et al.
Veröffentlicht: (2025)
von: Yue, Chuhuai, et al.
Veröffentlicht: (2025)
Adaptable Hindsight Experience Replay for Search-Based Learning
von: Vazaios, Alexandros, et al.
Veröffentlicht: (2025)
von: Vazaios, Alexandros, et al.
Veröffentlicht: (2025)
Training Multi-Image Vision Agents via End2End Reinforcement Learning
von: Dong, Chengqi, et al.
Veröffentlicht: (2025)
von: Dong, Chengqi, et al.
Veröffentlicht: (2025)
ACR: Adaptive Context Refactoring via Context Refactoring Operators for Multi-Turn Dialogue
von: Shen, Jiawei, et al.
Veröffentlicht: (2026)
von: Shen, Jiawei, et al.
Veröffentlicht: (2026)
RLAE: Reinforcement Learning-Assisted Ensemble for LLMs
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
MRHER: Model-based Relay Hindsight Experience Replay for Sequential Object Manipulation Tasks with Sparse Rewards
von: Huang, Yuming, et al.
Veröffentlicht: (2023)
von: Huang, Yuming, et al.
Veröffentlicht: (2023)
Human-Aware Robot Navigation via Reinforcement Learning with Hindsight Experience Replay and Curriculum Learning
von: Li, Keyu, et al.
Veröffentlicht: (2021)
von: Li, Keyu, et al.
Veröffentlicht: (2021)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
AIS: Adaptive Importance Sampling for Quantized RL
von: Zhou, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhou, Jiajun, et al.
Veröffentlicht: (2026)
Kevin: Multi-Turn RL for Generating CUDA Kernels
von: Baronio, Carlo, et al.
Veröffentlicht: (2025)
von: Baronio, Carlo, et al.
Veröffentlicht: (2025)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
von: Guo, Yipin, et al.
Veröffentlicht: (2026)
von: Guo, Yipin, et al.
Veröffentlicht: (2026)
H$^2$R: Hierarchical Hindsight Reflection for Multi-Task LLM Agents
von: Ye, Shicheng, et al.
Veröffentlicht: (2025)
von: Ye, Shicheng, et al.
Veröffentlicht: (2025)
Beyond More Context: Retrieval Diversity Boosts Multi-Turn Intent Understanding
von: Lin, Zhiming
Veröffentlicht: (2025)
von: Lin, Zhiming
Veröffentlicht: (2025)
Enhancing RAG Efficiency with Adaptive Context Compression
von: Guo, Shuyu, et al.
Veröffentlicht: (2025)
von: Guo, Shuyu, et al.
Veröffentlicht: (2025)
Prioritized Replay for RL Post-training
von: Fatemi, Mehdi
Veröffentlicht: (2026)
von: Fatemi, Mehdi
Veröffentlicht: (2026)
AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
von: Zhang, Hanchen, et al.
Veröffentlicht: (2025)
von: Zhang, Hanchen, et al.
Veröffentlicht: (2025)
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
von: He, Hang, et al.
Veröffentlicht: (2025)
von: He, Hang, et al.
Veröffentlicht: (2025)
Hindsight Preference Optimization for Financial Time Series Advisory
von: Cui, Yanwei, et al.
Veröffentlicht: (2026)
von: Cui, Yanwei, et al.
Veröffentlicht: (2026)
SAC-GLAM: Improving Online RL for LLM agents with Soft Actor-Critic and Hindsight Relabeling
von: Gaven, Loris, et al.
Veröffentlicht: (2024)
von: Gaven, Loris, et al.
Veröffentlicht: (2024)
AlphaZip: Neural Network-Enhanced Lossless Text Compression
von: Narashiman, Swathi Shree, et al.
Veröffentlicht: (2024)
von: Narashiman, Swathi Shree, et al.
Veröffentlicht: (2024)
MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games
von: Xie, Yunfei, et al.
Veröffentlicht: (2026)
von: Xie, Yunfei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
von: Wei, Zhenlin, et al.
Veröffentlicht: (2026) -
SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
von: Guo, Xiaojun, et al.
Veröffentlicht: (2025) -
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
von: Chen, Hao, et al.
Veröffentlicht: (2025) -
SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training
von: Zhang, Qi, et al.
Veröffentlicht: (2026) -
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
von: Chai, Jiajun, et al.
Veröffentlicht: (2025)