Distilling Feedback into Memory-as-a-Tool
Fuente:
arXiv
Saved in:
| Main Author: | Gallego, Víctor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distilled Self-Critique of LLMs with Synthetic Data: a Bayesian Perspective
by: Gallego, Victor
Published: (2023)
by: Gallego, Victor
Published: (2023)
Beyond Scalar Rewards: Dense Feedback for LLM Policy Synthesis in Sequential Social Dilemmas
by: Gallego, Víctor
Published: (2026)
by: Gallego, Víctor
Published: (2026)
Discovering Agentic Safety Specifications from 1-Bit Danger Signals
by: Gallego, Víctor
Published: (2026)
by: Gallego, Víctor
Published: (2026)
Configurable Preference Tuning with Rubric-Guided Synthetic Data
by: Gallego, Víctor
Published: (2025)
by: Gallego, Víctor
Published: (2025)
Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
by: Gallego, Víctor
Published: (2024)
by: Gallego, Víctor
Published: (2024)
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
by: Gallego, Víctor
Published: (2025)
by: Gallego, Víctor
Published: (2025)
Merging Improves Self-Critique Against Jailbreak Attacks
by: Gallego, Victor
Published: (2024)
by: Gallego, Victor
Published: (2024)
MetaSC: Test-Time Safety Specification Optimization for Language Models
by: Gallego, Víctor
Published: (2025)
by: Gallego, Víctor
Published: (2025)
Configurable Safety Tuning of Language Models with Synthetic Preference Data
by: Gallego, Victor
Published: (2024)
by: Gallego, Victor
Published: (2024)
MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
by: Xiao, Yunzhong, et al.
Published: (2025)
by: Xiao, Yunzhong, et al.
Published: (2025)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
by: Hu, Yushi, et al.
Published: (2023)
by: Hu, Yushi, et al.
Published: (2023)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
by: Mou, Yutao, et al.
Published: (2026)
by: Mou, Yutao, et al.
Published: (2026)
Small But Funny: A Feedback-Driven Approach to Humor Distillation
by: Ravi, Sahithya, et al.
Published: (2024)
by: Ravi, Sahithya, et al.
Published: (2024)
FlashMem: Distilling Intrinsic Latent Memory via Computation Reuse
by: Hou, Yubo, et al.
Published: (2026)
by: Hou, Yubo, et al.
Published: (2026)
MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents
by: Li, Zihan, et al.
Published: (2026)
by: Li, Zihan, et al.
Published: (2026)
ToolPlanner: A Tool Augmented LLM for Multi Granularity Instructions with Path Planning and Feedback
by: Wu, Qinzhuo, et al.
Published: (2024)
by: Wu, Qinzhuo, et al.
Published: (2024)
Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management
by: Ma, Weitao, et al.
Published: (2026)
by: Ma, Weitao, et al.
Published: (2026)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
by: Ma, Junyu, et al.
Published: (2025)
by: Ma, Junyu, et al.
Published: (2025)
When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors
by: Yang, Chenghao, et al.
Published: (2026)
by: Yang, Chenghao, et al.
Published: (2026)
Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph Translation
by: Yin, Fan, et al.
Published: (2025)
by: Yin, Fan, et al.
Published: (2025)
Distilling LLM Agent into Small Models with Retrieval and Code Tools
by: Kang, Minki, et al.
Published: (2025)
by: Kang, Minki, et al.
Published: (2025)
MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
by: Lumer, Elias, et al.
Published: (2025)
by: Lumer, Elias, et al.
Published: (2025)
Enhancing Tool Retrieval with Iterative Feedback from Large Language Models
by: Xu, Qiancheng, et al.
Published: (2024)
by: Xu, Qiancheng, et al.
Published: (2024)
Toward Efficient Agents: Memory, Tool learning, and Planning
by: Yang, Xiaofang, et al.
Published: (2026)
by: Yang, Xiaofang, et al.
Published: (2026)
Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
by: Yang, Ning, et al.
Published: (2026)
by: Yang, Ning, et al.
Published: (2026)
VeriAgent: A Tool-Integrated Multi-Agent System with Evolving Memory for PPA-Aware RTL Code Generation
by: Wang, Yaoxiang, et al.
Published: (2026)
by: Wang, Yaoxiang, et al.
Published: (2026)
Unifying Dynamic Tool Creation and Cross-Task Experience Sharing through Cognitive Memory Architecture
by: Liu, Jiarun, et al.
Published: (2025)
by: Liu, Jiarun, et al.
Published: (2025)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
by: Wang, Xingyao, et al.
Published: (2023)
by: Wang, Xingyao, et al.
Published: (2023)
MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications
by: Sultan, Oren, et al.
Published: (2024)
by: Sultan, Oren, et al.
Published: (2024)
Minerva: A Programmable Memory Test Benchmark for Language Models
by: Xia, Menglin, et al.
Published: (2025)
by: Xia, Menglin, et al.
Published: (2025)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
by: Zhou, Xiaofeng, et al.
Published: (2025)
by: Zhou, Xiaofeng, et al.
Published: (2025)
Making Language Models Better Tool Learners with Execution Feedback
by: Qiao, Shuofei, et al.
Published: (2023)
by: Qiao, Shuofei, et al.
Published: (2023)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025)
by: Csizmadia, Daniel, et al.
Published: (2025)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
by: Li, Zeju, et al.
Published: (2026)
by: Li, Zeju, et al.
Published: (2026)
TSUBASA: Improving Long-Horizon Personalization via Evolving Memory and Self-Learning with Context Distillation
by: Zhang, Xinliang Frederick, et al.
Published: (2026)
by: Zhang, Xinliang Frederick, et al.
Published: (2026)
Similar Items
-
Distilled Self-Critique of LLMs with Synthetic Data: a Bayesian Perspective
by: Gallego, Victor
Published: (2023) -
Beyond Scalar Rewards: Dense Feedback for LLM Policy Synthesis in Sequential Social Dilemmas
by: Gallego, Víctor
Published: (2026) -
Discovering Agentic Safety Specifications from 1-Bit Danger Signals
by: Gallego, Víctor
Published: (2026) -
Configurable Preference Tuning with Rubric-Guided Synthetic Data
by: Gallego, Víctor
Published: (2025) -
Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
by: Gallego, Víctor
Published: (2024)