Offline Training of Language Model Agents with Functions as Learnable Weights
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Shaokun, Zhang, Jieyu, Liu, Jiale, Song, Linxin, Wang, Chi, Krishna, Ranjay, Wu, Qingyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive In-conversation Team Building for Language Model Agents
von: Song, Linxin, et al.
Veröffentlicht: (2024)
von: Song, Linxin, et al.
Veröffentlicht: (2024)
Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025)
StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows
von: Wu, Yiran, et al.
Veröffentlicht: (2024)
von: Wu, Yiran, et al.
Veröffentlicht: (2024)
EcoAct: Economic Agent Determines When to Register What Action
von: Zhang, Shaokun, et al.
Veröffentlicht: (2024)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2024)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
Memory-Augmented Agent Training for Business Document Understanding
von: Liu, Jiale, et al.
Veröffentlicht: (2024)
von: Liu, Jiale, et al.
Veröffentlicht: (2024)
Video-Based Reward Modeling for Computer-Use Agents
von: Song, Linxin, et al.
Veröffentlicht: (2026)
von: Song, Linxin, et al.
Veröffentlicht: (2026)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
von: Yu, Yijiong, et al.
Veröffentlicht: (2025)
von: Yu, Yijiong, et al.
Veröffentlicht: (2025)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
von: Wang, Kangrui, et al.
Veröffentlicht: (2025)
von: Wang, Kangrui, et al.
Veröffentlicht: (2025)
Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025)
Retrospex: Language Agent Meets Offline Reinforcement Learning Critic
von: Xiang, Yufei, et al.
Veröffentlicht: (2025)
von: Xiang, Yufei, et al.
Veröffentlicht: (2025)
MentalArena: Self-play Training of Language Models for Diagnosis and Treatment of Mental Health Disorders
von: Li, Cheng, et al.
Veröffentlicht: (2024)
von: Li, Cheng, et al.
Veröffentlicht: (2024)
MindCube: Spatial Mental Modeling from Limited Views
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
von: Zhang, Pingyue, et al.
Veröffentlicht: (2026)
von: Zhang, Pingyue, et al.
Veröffentlicht: (2026)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
Safer-Instruct: Aligning Language Models with Automated Preference Data
von: Shi, Taiwei, et al.
Veröffentlicht: (2023)
von: Shi, Taiwei, et al.
Veröffentlicht: (2023)
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
Importance Weighting Can Help Large Language Models Self-Improve
von: Jiang, Chunyang, et al.
Veröffentlicht: (2024)
von: Jiang, Chunyang, et al.
Veröffentlicht: (2024)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
von: Yang, Yinuo, et al.
Veröffentlicht: (2026)
von: Yang, Yinuo, et al.
Veröffentlicht: (2026)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
von: Long, Yunbo, et al.
Veröffentlicht: (2026)
von: Long, Yunbo, et al.
Veröffentlicht: (2026)
Embodied LLM Agents Learn to Cooperate in Organized Teams
von: Guo, Xudong, et al.
Veröffentlicht: (2024)
von: Guo, Xudong, et al.
Veröffentlicht: (2024)
Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models
von: Zhang, Shaokun, et al.
Veröffentlicht: (2023)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2023)
Brain in a Vat: On Missing Pieces Towards Artificial General Intelligence in Large Language Models
von: Ma, Yuxi, et al.
Veröffentlicht: (2023)
von: Ma, Yuxi, et al.
Veröffentlicht: (2023)
DLP: Dynamic Layerwise Pruning in Large Language Models
von: Chen, Yuli, et al.
Veröffentlicht: (2025)
von: Chen, Yuli, et al.
Veröffentlicht: (2025)
Online Training of Large Language Models: Learn while chatting
von: Liang, Juhao, et al.
Veröffentlicht: (2024)
von: Liang, Juhao, et al.
Veröffentlicht: (2024)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
Large Language Models are Learnable Planners for Long-Term Recommendation
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
PIORS: Personalized Intelligent Outpatient Reception based on Large Language Model with Multi-Agents Medical Scenario Simulation
von: Bao, Zhijie, et al.
Veröffentlicht: (2024)
von: Bao, Zhijie, et al.
Veröffentlicht: (2024)
Training Agents with Weakly Supervised Feedback from Large Language Models
von: Gong, Dihong, et al.
Veröffentlicht: (2024)
von: Gong, Dihong, et al.
Veröffentlicht: (2024)
Param$Δ$ for Direct Weight Mixing: Post-Train Large Language Model at Zero Cost
von: Cao, Sheng, et al.
Veröffentlicht: (2025)
von: Cao, Sheng, et al.
Veröffentlicht: (2025)
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
von: Ding, Dujian, et al.
Veröffentlicht: (2025)
von: Ding, Dujian, et al.
Veröffentlicht: (2025)
BriLLM: Brain-inspired Large Language Model
von: Zhao, Hai, et al.
Veröffentlicht: (2025)
von: Zhao, Hai, et al.
Veröffentlicht: (2025)
Rethinking Human Preference Evaluation of LLM Rationales
von: Li, Ziang, et al.
Veröffentlicht: (2025)
von: Li, Ziang, et al.
Veröffentlicht: (2025)
Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
SyntheT2C: Generating Synthetic Data for Fine-Tuning Large Language Models on the Text2Cypher Task
von: Zhong, Ziije, et al.
Veröffentlicht: (2024)
von: Zhong, Ziije, et al.
Veröffentlicht: (2024)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adaptive In-conversation Team Building for Language Model Agents
von: Song, Linxin, et al.
Veröffentlicht: (2024) -
Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025) -
StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows
von: Wu, Yiran, et al.
Veröffentlicht: (2024) -
EcoAct: Economic Agent Determines When to Register What Action
von: Zhang, Shaokun, et al.
Veröffentlicht: (2024) -
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)