SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jiazheng, Wang, Yawei, Yan, David, Tian, Yijun, Xu, Zhichao, Song, Huan, Xu, Panpan, Cheong, Lin Lee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces
von: Wu, Guande, et al.
Veröffentlicht: (2025)
von: Wu, Guande, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Self-Improving Agent with Skill Library
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2025)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2025)
CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection
von: Liu, Linbo, et al.
Veröffentlicht: (2026)
von: Liu, Linbo, et al.
Veröffentlicht: (2026)
CSPLADE: Learned Sparse Retrieval with Causal Language Models
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling
von: Liu, Yuchen, et al.
Veröffentlicht: (2026)
von: Liu, Yuchen, et al.
Veröffentlicht: (2026)
Graph Neural Prompting with Large Language Models
von: Tian, Yijun, et al.
Veröffentlicht: (2023)
von: Tian, Yijun, et al.
Veröffentlicht: (2023)
A Sober Look at Agentic Misalignment in Automated Workflows
von: Ye, Wenqian, et al.
Veröffentlicht: (2026)
von: Ye, Wenqian, et al.
Veröffentlicht: (2026)
RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
Reinforcement Mid-Training
von: Tian, Yijun, et al.
Veröffentlicht: (2025)
von: Tian, Yijun, et al.
Veröffentlicht: (2025)
Multi-level Advantage Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
von: Zhao, Xutong, et al.
Veröffentlicht: (2025)
von: Zhao, Xutong, et al.
Veröffentlicht: (2025)
RTMC: Step-Level Credit Assignment via Rollout Trees
von: Wang, Tao, et al.
Veröffentlicht: (2026)
von: Wang, Tao, et al.
Veröffentlicht: (2026)
SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning
von: Wang, Jichao, et al.
Veröffentlicht: (2026)
von: Wang, Jichao, et al.
Veröffentlicht: (2026)
Learning to Ideate for Machine Learning Engineering Agents
von: Zhang, Yunxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Yunxiang, et al.
Veröffentlicht: (2026)
WildGEN: Long-horizon Trajectory Generation for Wildlife
von: Al-Lawati, Ali, et al.
Veröffentlicht: (2023)
von: Al-Lawati, Ali, et al.
Veröffentlicht: (2023)
UAV Task Assignment and Path Planning: Problem Definition, Integration, and Challenges
von: Huan Meng, et al.
Veröffentlicht: (2026)
von: Huan Meng, et al.
Veröffentlicht: (2026)
PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement
von: Zhang, Tuo, et al.
Veröffentlicht: (2026)
von: Zhang, Tuo, et al.
Veröffentlicht: (2026)
QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence
von: Lin, Zhichao, et al.
Veröffentlicht: (2026)
von: Lin, Zhichao, et al.
Veröffentlicht: (2026)
Diffusion Trajectory-guided Policy for Long-horizon Robot Manipulation
von: Fan, Shichao, et al.
Veröffentlicht: (2025)
von: Fan, Shichao, et al.
Veröffentlicht: (2025)
Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents
von: Kang, Jiazheng, et al.
Veröffentlicht: (2026)
von: Kang, Jiazheng, et al.
Veröffentlicht: (2026)
LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum
von: Xu, Zhichao, et al.
Veröffentlicht: (2026)
von: Xu, Zhichao, et al.
Veröffentlicht: (2026)
StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
von: Zhang, Yanfei, et al.
Veröffentlicht: (2026)
von: Zhang, Yanfei, et al.
Veröffentlicht: (2026)
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
Ensemble Quadratic Assignment Network for Graph Matching
von: Tan, Haoru, et al.
Veröffentlicht: (2024)
von: Tan, Haoru, et al.
Veröffentlicht: (2024)
Tree of Agents: Improving Long-Context Capabilities of Large Language Models through Multi-Perspective Reasoning
von: Yu, Song, et al.
Veröffentlicht: (2025)
von: Yu, Song, et al.
Veröffentlicht: (2025)
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
LMEB: Long-horizon Memory Embedding Benchmark
von: Zhao, Xinping, et al.
Veröffentlicht: (2026)
von: Zhao, Xinping, et al.
Veröffentlicht: (2026)
SALT4Decompile: Inferring Source-level Abstract Logic Tree for LLM-Based Binary Decompilation
von: Wang, Yongpan, et al.
Veröffentlicht: (2025)
von: Wang, Yongpan, et al.
Veröffentlicht: (2025)
Hindsight Credit Assignment for Long-Horizon LLM Agents
von: Tan, Hui-Ze, et al.
Veröffentlicht: (2026)
von: Tan, Hui-Ze, et al.
Veröffentlicht: (2026)
Collaborative LLM Numerical Reasoning with Local Data Protection
von: Zhang, Min, et al.
Veröffentlicht: (2025)
von: Zhang, Min, et al.
Veröffentlicht: (2025)
The SALT response index: Overcoming limitations of the SALT score for evaluating treatment response
von: Marc Hernández‐Santacana, et al.
Veröffentlicht: (2026)
von: Marc Hernández‐Santacana, et al.
Veröffentlicht: (2026)
Graph is a Substrate Across Data Modalities
von: Li, Ziming, et al.
Veröffentlicht: (2026)
von: Li, Ziming, et al.
Veröffentlicht: (2026)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization
von: Wang, Zhengcheng, et al.
Veröffentlicht: (2025)
von: Wang, Zhengcheng, et al.
Veröffentlicht: (2025)
ACON: Optimizing Context Compression for Long-horizon LLM Agents
von: Kang, Minki, et al.
Veröffentlicht: (2025)
von: Kang, Minki, et al.
Veröffentlicht: (2025)
LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents
von: Rakhsha, Amin, et al.
Veröffentlicht: (2026)
von: Rakhsha, Amin, et al.
Veröffentlicht: (2026)
KLong: Training LLM Agent for Extremely Long-horizon Tasks
von: Liu, Yue, et al.
Veröffentlicht: (2026)
von: Liu, Yue, et al.
Veröffentlicht: (2026)
M$^2$: Dual-Memory Augmentation for Long-Horizon Web Agents via Trajectory Summarization and Insight Retrieval
von: Yan, Dawei, et al.
Veröffentlicht: (2026)
von: Yan, Dawei, et al.
Veröffentlicht: (2026)
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
von: Shi, Yibo, et al.
Veröffentlicht: (2026)
von: Shi, Yibo, et al.
Veröffentlicht: (2026)
Path Learning with Trajectory Advantage Regression
von: Miyaguchi, Kohei
Veröffentlicht: (2025)
von: Miyaguchi, Kohei
Veröffentlicht: (2025)
Cooperative Multi-Agent Assignment over Stochastic Graphs via Constrained Reinforcement Learning
von: Agorio, Leopoldo, et al.
Veröffentlicht: (2025)
von: Agorio, Leopoldo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces
von: Wu, Guande, et al.
Veröffentlicht: (2025) -
Reinforcement Learning for Self-Improving Agent with Skill Library
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2025) -
CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection
von: Liu, Linbo, et al.
Veröffentlicht: (2026) -
CSPLADE: Learned Sparse Retrieval with Causal Language Models
von: Xu, Zhichao, et al.
Veröffentlicht: (2025) -
Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling
von: Liu, Yuchen, et al.
Veröffentlicht: (2026)