TVCACHE: A Stateful Tool-Value Cache for Post-Training LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Abhishek Vijaya, Kataria, Bhaskar, Oh, Byungsoo, Manzoor, Emaad, Singh, Rachee |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
by: Ding, Eric, et al.
Published: (2026)
by: Ding, Eric, et al.
Published: (2026)
FlashMoE: Fast Distributed MoE in a Single Kernel
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
Eliminating Hidden Serialization in Multi-Node Megakernel Communication
by: Oh, Byungsoo, et al.
Published: (2026)
by: Oh, Byungsoo, et al.
Published: (2026)
LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
Efficient AllReduce with Stragglers
by: Devraj, Arjun, et al.
Published: (2025)
by: Devraj, Arjun, et al.
Published: (2025)
Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
Learning When to Quit in Sales Conversations
by: Manzoor, Emaad, et al.
Published: (2025)
by: Manzoor, Emaad, et al.
Published: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
PCCL: Photonic circuit-switched collective communication for distributed ML
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
Photonic Rails in ML Datacenters with Opus
by: Ding, Eric, et al.
Published: (2026)
by: Ding, Eric, et al.
Published: (2026)
Wireless Communication Enhanced Value Decomposition for Multi-Agent Reinforcement Learning
by: Hu, Diyi, et al.
Published: (2026)
by: Hu, Diyi, et al.
Published: (2026)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026)
by: Rank, Ben, et al.
Published: (2026)
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025)
by: Bhatia, Mehar, et al.
Published: (2025)
Chip-to-chip photonic connectivity in multi-accelerator servers for ML
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
LLM Agents for Bargaining with Utility-based Feedback
by: Oh, Jihwan
Published: (2025)
by: Oh, Jihwan
Published: (2025)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
by: Hu, Pingbang, et al.
Published: (2026)
by: Hu, Pingbang, et al.
Published: (2026)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
by: Afroz, Sabiha, et al.
Published: (2025)
by: Afroz, Sabiha, et al.
Published: (2025)
RDAR: Reward-Driven Agent Relevance Estimation for Autonomous Driving
by: Bosio, Carlo, et al.
Published: (2025)
by: Bosio, Carlo, et al.
Published: (2025)
Digi-Q: Learning Q-Value Functions for Training Device-Control Agents
by: Bai, Hao, et al.
Published: (2025)
by: Bai, Hao, et al.
Published: (2025)
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
by: Chai, Jiajun, et al.
Published: (2025)
by: Chai, Jiajun, et al.
Published: (2025)
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
by: Jeon, Hyesung, et al.
Published: (2026)
by: Jeon, Hyesung, et al.
Published: (2026)
DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2026)
by: Zhu, Dingwei, et al.
Published: (2026)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
by: Yuan, Yurun, et al.
Published: (2026)
by: Yuan, Yurun, et al.
Published: (2026)
GraphFLEx: Structure Learning Framework for Large Expanding Graphs
by: Kataria, Mohit, et al.
Published: (2025)
by: Kataria, Mohit, et al.
Published: (2025)
CacheFormer: High Attention-Based Segment Caching
by: Singh, Sushant, et al.
Published: (2025)
by: Singh, Sushant, et al.
Published: (2025)
Training Transformers for KV Cache Compressibility
by: Gelberg, Yoav, et al.
Published: (2026)
by: Gelberg, Yoav, et al.
Published: (2026)
SHARe-KAN: Post-Training Vector Quantization for Cache-Resident KAN Inference
by: Smith, Jeff
Published: (2025)
by: Smith, Jeff
Published: (2025)
Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
by: Ma, Qingsen, et al.
Published: (2025)
by: Ma, Qingsen, et al.
Published: (2025)
How Many Tools Should an LLM Agent See? A Chance-Corrected Answer
by: Repantis, Vyzantinos, et al.
Published: (2026)
by: Repantis, Vyzantinos, et al.
Published: (2026)
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training
by: Thede, Lukas, et al.
Published: (2026)
by: Thede, Lukas, et al.
Published: (2026)
Consolidating Rewarded Perturbations for LLM Post-Training
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
Automatic Configuration of LLM Post-Training Pipelines
by: Chwa, Channe, et al.
Published: (2026)
by: Chwa, Channe, et al.
Published: (2026)
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents
by: Chu, Kexin, et al.
Published: (2026)
by: Chu, Kexin, et al.
Published: (2026)
Memory-Induced Tool-Drift in LLM Agents
by: Dabas, Mahavir, et al.
Published: (2026)
by: Dabas, Mahavir, et al.
Published: (2026)
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
by: Acikgoz, Emre Can, et al.
Published: (2026)
by: Acikgoz, Emre Can, et al.
Published: (2026)
GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
by: Regmi, Sajal, et al.
Published: (2024)
by: Regmi, Sajal, et al.
Published: (2024)
Learning to Evict from Key-Value Cache
by: Moschella, Luca, et al.
Published: (2026)
by: Moschella, Luca, et al.
Published: (2026)
LLM Agents Making Agent Tools
by: Wölflein, Georg, et al.
Published: (2025)
by: Wölflein, Georg, et al.
Published: (2025)
Similar Items
-
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
by: Ding, Eric, et al.
Published: (2026) -
FlashMoE: Fast Distributed MoE in a Single Kernel
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025) -
Eliminating Hidden Serialization in Multi-Node Megakernel Communication
by: Oh, Byungsoo, et al.
Published: (2026) -
LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics
by: Kumar, Abhishek Vijaya, et al.
Published: (2025) -
Efficient AllReduce with Stragglers
by: Devraj, Arjun, et al.
Published: (2025)