On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Sunghwan, Cho, Junhee, Kwak, Beong-woo, Kwon, Taeyoon, Wang, Liang, Yang, Nan, Zhang, Xingxing, Wei, Furu, Yeo, Jinyoung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study
von: Yang, Dongil, et al.
Veröffentlicht: (2025)
von: Yang, Dongil, et al.
Veröffentlicht: (2025)
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
von: Kwon, Taeyoon, et al.
Veröffentlicht: (2025)
von: Kwon, Taeyoon, et al.
Veröffentlicht: (2025)
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
von: Kwak, Beong-woo, et al.
Veröffentlicht: (2025)
von: Kwak, Beong-woo, et al.
Veröffentlicht: (2025)
Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024)
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024)
Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation
von: Kang, Dongjin, et al.
Veröffentlicht: (2024)
von: Kang, Dongjin, et al.
Veröffentlicht: (2024)
Coffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Code
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024)
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
von: Ong, Kai Tzu-iunn, et al.
Veröffentlicht: (2024)
von: Ong, Kai Tzu-iunn, et al.
Veröffentlicht: (2024)
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
von: Choi, Dongwook, et al.
Veröffentlicht: (2025)
von: Choi, Dongwook, et al.
Veröffentlicht: (2025)
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2025)
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2025)
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2025)
Towards Direct Evaluation of Harness Optimizers via Priority Ranking
von: Ong, Kai Tzu-iunn, et al.
Veröffentlicht: (2026)
von: Ong, Kai Tzu-iunn, et al.
Veröffentlicht: (2026)
Pearl: A Review-driven Persona-Knowledge Grounded Conversational Recommendation Dataset
von: Kim, Minjin, et al.
Veröffentlicht: (2024)
von: Kim, Minjin, et al.
Veröffentlicht: (2024)
Bootstrap Your Own Context Length
von: Wang, Liang, et al.
Veröffentlicht: (2024)
von: Wang, Liang, et al.
Veröffentlicht: (2024)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
von: Lee, Seungbeen, et al.
Veröffentlicht: (2024)
von: Lee, Seungbeen, et al.
Veröffentlicht: (2024)
FCRF: Flexible Constructivism Reflection for Long-Horizon Robotic Task Planning with Large Language Models
von: Song, Yufan, et al.
Veröffentlicht: (2025)
von: Song, Yufan, et al.
Veröffentlicht: (2025)
Stop Playing the Guessing Game! Target-free User Simulation for Evaluating Conversational Recommender Systems
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)
STRUCTUREDAGENT: Planning with AND/OR Trees for Long-Horizon Web Tasks
von: Lobo, ELita, et al.
Veröffentlicht: (2026)
von: Lobo, ELita, et al.
Veröffentlicht: (2026)
MLDT: Multi-Level Decomposition for Complex Long-Horizon Robotic Task Planning with Open-Source Large Language Model
von: Wu, Yike, et al.
Veröffentlicht: (2024)
von: Wu, Yike, et al.
Veröffentlicht: (2024)
Spatially Grounded Long-Horizon Task Planning in the Wild
von: Jung, Sehun, et al.
Veröffentlicht: (2026)
von: Jung, Sehun, et al.
Veröffentlicht: (2026)
AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
Learning to Retrieve In-Context Examples for Large Language Models
von: Wang, Liang, et al.
Veröffentlicht: (2023)
von: Wang, Liang, et al.
Veröffentlicht: (2023)
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
von: Park, Minjun, et al.
Veröffentlicht: (2026)
von: Park, Minjun, et al.
Veröffentlicht: (2026)
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
von: Zhou, Zijian, et al.
Veröffentlicht: (2025)
von: Zhou, Zijian, et al.
Veröffentlicht: (2025)
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
von: Kim, Minju, et al.
Veröffentlicht: (2025)
von: Kim, Minju, et al.
Veröffentlicht: (2025)
A Backbone for Long-Horizon Robot Task Understanding
von: Chen, Xiaoshuai, et al.
Veröffentlicht: (2024)
von: Chen, Xiaoshuai, et al.
Veröffentlicht: (2024)
Scaling Optimal LR Across Token Horizons
von: Bjorck, Johan, et al.
Veröffentlicht: (2024)
von: Bjorck, Johan, et al.
Veröffentlicht: (2024)
HiMem: Hierarchical Long-Term Memory for LLM Long-Horizon Agents
von: Zhang, Ningning, et al.
Veröffentlicht: (2026)
von: Zhang, Ningning, et al.
Veröffentlicht: (2026)
LHAW: Controllable Underspecification for Long-Horizon Tasks
von: Pu, George, et al.
Veröffentlicht: (2026)
von: Pu, George, et al.
Veröffentlicht: (2026)
CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
von: He, Shuo, et al.
Veröffentlicht: (2026)
von: He, Shuo, et al.
Veröffentlicht: (2026)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
Guiding Long-Horizon Task and Motion Planning with Vision Language Models
von: Yang, Zhutian, et al.
Veröffentlicht: (2024)
von: Yang, Zhutian, et al.
Veröffentlicht: (2024)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
von: Yeo, Woongyeng, et al.
Veröffentlicht: (2026)
von: Yeo, Woongyeng, et al.
Veröffentlicht: (2026)
EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents
von: Choi, Dongwook, et al.
Veröffentlicht: (2026)
von: Choi, Dongwook, et al.
Veröffentlicht: (2026)
Learning Agent-Compatible Context Management for Long-Horizon Tasks
von: Yi, Lu, et al.
Veröffentlicht: (2026)
von: Yi, Lu, et al.
Veröffentlicht: (2026)
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
von: Lee, Hwiwon, et al.
Veröffentlicht: (2026)
von: Lee, Hwiwon, et al.
Veröffentlicht: (2026)
Generalizable Dense Reward for Long-Horizon Robotic Tasks
von: Yong, Silong, et al.
Veröffentlicht: (2026)
von: Yong, Silong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study
von: Yang, Dongil, et al.
Veröffentlicht: (2025) -
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
von: Kwon, Taeyoon, et al.
Veröffentlicht: (2025) -
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
von: Kwak, Beong-woo, et al.
Veröffentlicht: (2025) -
Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024) -
Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation
von: Kang, Dongjin, et al.
Veröffentlicht: (2024)