Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Yifan, Yin, Da, Yue, Xiang, Huang, Jie, Li, Sujian, Lin, Bill Yuchen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MPO: Boosting LLM Agents with Meta Plan Optimization
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023)
von: Yin, Da, et al.
Veröffentlicht: (2023)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
von: Song, Yifan, et al.
Veröffentlicht: (2024)
von: Song, Yifan, et al.
Veröffentlicht: (2024)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
TrajAgent: An LLM-Agent Framework for Trajectory Modeling via Large-and-Small Model Collaboration
von: Du, Yuwei, et al.
Veröffentlicht: (2024)
von: Du, Yuwei, et al.
Veröffentlicht: (2024)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
Temporal Consistency for LLM Reasoning Process Error Identification
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories
von: Song, Yifan, et al.
Veröffentlicht: (2024)
von: Song, Yifan, et al.
Veröffentlicht: (2024)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories
von: Yu, Zhuoyun, et al.
Veröffentlicht: (2026)
von: Yu, Zhuoyun, et al.
Veröffentlicht: (2026)
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
von: Djuhera, Aladin, et al.
Veröffentlicht: (2026)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2026)
ExpeL: LLM Agents Are Experiential Learners
von: Zhao, Andrew, et al.
Veröffentlicht: (2023)
von: Zhao, Andrew, et al.
Veröffentlicht: (2023)
Policy-Invisible Violations in LLM-Based Agents
von: Wu, Jie, et al.
Veröffentlicht: (2026)
von: Wu, Jie, et al.
Veröffentlicht: (2026)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
von: Younesian, Sharareh, et al.
Veröffentlicht: (2026)
von: Younesian, Sharareh, et al.
Veröffentlicht: (2026)
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
von: Wen, Bingbing, et al.
Veröffentlicht: (2026)
von: Wen, Bingbing, et al.
Veröffentlicht: (2026)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
von: Li, Sijia, et al.
Veröffentlicht: (2026)
von: Li, Sijia, et al.
Veröffentlicht: (2026)
DataSciBench: An LLM Agent Benchmark for Data Science
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
von: Wang, Boshi, et al.
Veröffentlicht: (2024)
von: Wang, Boshi, et al.
Veröffentlicht: (2024)
AutoPDL: Automatic Prompt Optimization for LLM Agents
von: Spiess, Claudio, et al.
Veröffentlicht: (2025)
von: Spiess, Claudio, et al.
Veröffentlicht: (2025)
Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling
von: Guo, Yiran, et al.
Veröffentlicht: (2026)
von: Guo, Yiran, et al.
Veröffentlicht: (2026)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
Multi-LLM QA with Embodied Exploration
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
Boosting of Thoughts: Trial-and-Error Problem Solving with Large Language Models
von: Chen, Sijia, et al.
Veröffentlicht: (2024)
von: Chen, Sijia, et al.
Veröffentlicht: (2024)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
von: Xia, Fanzeng, et al.
Veröffentlicht: (2024)
von: Xia, Fanzeng, et al.
Veröffentlicht: (2024)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
Peering Inside the Black Box: Uncovering LLM Errors in Optimization Modelling through Component-Level Evaluation
von: Refai, Dania, et al.
Veröffentlicht: (2025)
von: Refai, Dania, et al.
Veröffentlicht: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker Environments
von: Xia, Boyang, et al.
Veröffentlicht: (2026)
von: Xia, Boyang, et al.
Veröffentlicht: (2026)
Selecting Large Language Model to Fine-tune via Rectified Scaling Law
von: Lin, Haowei, et al.
Veröffentlicht: (2024)
von: Lin, Haowei, et al.
Veröffentlicht: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
Encoding Agent Trajectories as Representations with Sequence Transformers
von: Tsiligkaridis, Athanasios, et al.
Veröffentlicht: (2024)
von: Tsiligkaridis, Athanasios, et al.
Veröffentlicht: (2024)
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
von: Yin, Yueqin, et al.
Veröffentlicht: (2024)
von: Yin, Yueqin, et al.
Veröffentlicht: (2024)
Error Taxonomy-Guided Prompt Optimization
von: Singh, Mayank, et al.
Veröffentlicht: (2026)
von: Singh, Mayank, et al.
Veröffentlicht: (2026)
LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
von: Hoang, Thai, et al.
Veröffentlicht: (2025)
von: Hoang, Thai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MPO: Boosting LLM Agents with Meta Plan Optimization
von: Xiong, Weimin, et al.
Veröffentlicht: (2025) -
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023) -
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
von: Song, Yifan, et al.
Veröffentlicht: (2024) -
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
von: Xiong, Weimin, et al.
Veröffentlicht: (2024) -
TrajAgent: An LLM-Agent Framework for Trajectory Modeling via Large-and-Small Model Collaboration
von: Du, Yuwei, et al.
Veröffentlicht: (2024)