KLong: Training LLM Agent for Extremely Long-horizon Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yue, Ma, Yingwei, Miao, Yibo, Li, Yanhao, Xie, Yuchong, Yang, Xinlong, Hu, Zhiyuan, Sung, Flood, Zhang, Jiaheng, Hooi, Bryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Coding Agents via Atomic Skills
von: Ma, Yingwei, et al.
Veröffentlicht: (2026)
von: Ma, Yingwei, et al.
Veröffentlicht: (2026)
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
von: Li, Yibo, et al.
Veröffentlicht: (2025)
von: Li, Yibo, et al.
Veröffentlicht: (2025)
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
von: Yang, Zonghan, et al.
Veröffentlicht: (2025)
von: Yang, Zonghan, et al.
Veröffentlicht: (2025)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Enhancing Multi-Agent Debate System Performance via Confidence Expression
von: Lin, Zijie, et al.
Veröffentlicht: (2025)
von: Lin, Zijie, et al.
Veröffentlicht: (2025)
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
von: He, Yufei, et al.
Veröffentlicht: (2025)
von: He, Yufei, et al.
Veröffentlicht: (2025)
Towards Realistic Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions
von: Guo, Qianyun, et al.
Veröffentlicht: (2026)
von: Guo, Qianyun, et al.
Veröffentlicht: (2026)
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
Safety in Large Reasoning Models: A Survey
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks
von: Wu, Jiaying, et al.
Veröffentlicht: (2023)
von: Wu, Jiaying, et al.
Veröffentlicht: (2023)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
von: Chen, Hui, et al.
Veröffentlicht: (2025)
von: Chen, Hui, et al.
Veröffentlicht: (2025)
ACON: Optimizing Context Compression for Long-horizon LLM Agents
von: Kang, Minki, et al.
Veröffentlicht: (2025)
von: Kang, Minki, et al.
Veröffentlicht: (2025)
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025)
FlipAttack: Jailbreak LLMs via Flipping
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
Efficient Reasoning via Chain of Unconscious Thought
von: Gong, Ruihan, et al.
Veröffentlicht: (2025)
von: Gong, Ruihan, et al.
Veröffentlicht: (2025)
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
von: Sui, Yuan, et al.
Veröffentlicht: (2026)
von: Sui, Yuan, et al.
Veröffentlicht: (2026)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
von: Li, Yibo, et al.
Veröffentlicht: (2026)
von: Li, Yibo, et al.
Veröffentlicht: (2026)
Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks
von: Yu, Zijian, et al.
Veröffentlicht: (2026)
von: Yu, Zijian, et al.
Veröffentlicht: (2026)
Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View
von: Zhang, Jintian, et al.
Veröffentlicht: (2023)
von: Zhang, Jintian, et al.
Veröffentlicht: (2023)
From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
von: Han, Wei, et al.
Veröffentlicht: (2025)
von: Han, Wei, et al.
Veröffentlicht: (2025)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
von: Yang, Shidong, et al.
Veröffentlicht: (2026)
von: Yang, Shidong, et al.
Veröffentlicht: (2026)
LMEB: Long-horizon Memory Embedding Benchmark
von: Zhao, Xinping, et al.
Veröffentlicht: (2026)
von: Zhao, Xinping, et al.
Veröffentlicht: (2026)
Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes
von: Fu, Zihang, et al.
Veröffentlicht: (2026)
von: Fu, Zihang, et al.
Veröffentlicht: (2026)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
How Does Response Length Affect Long-Form Factuality
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
Autonomous Chain-of-Thought Distillation for Graph-Based Fraud Detection
von: Li, Yuan, et al.
Veröffentlicht: (2026)
von: Li, Yuan, et al.
Veröffentlicht: (2026)
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?
von: Zhao, Yibo, et al.
Veröffentlicht: (2026)
von: Zhao, Yibo, et al.
Veröffentlicht: (2026)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Scaling Long-Horizon LLM Agent via Context-Folding
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
von: Ye, Fangda, et al.
Veröffentlicht: (2026)
von: Ye, Fangda, et al.
Veröffentlicht: (2026)
Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question Answering
von: Sui, Yuan, et al.
Veröffentlicht: (2024)
von: Sui, Yuan, et al.
Veröffentlicht: (2024)
LLM$\times$MapReduce-V2: Entropy-Driven Convolutional Test-Time Scaling for Generating Long-Form Articles from Extremely Long Resources
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Coding Agents via Atomic Skills
von: Ma, Yingwei, et al.
Veröffentlicht: (2026) -
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
von: Li, Yibo, et al.
Veröffentlicht: (2025) -
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
von: Yang, Zonghan, et al.
Veröffentlicht: (2025) -
Geneshift: Impact of different scenario shift on Jailbreaking LLM
von: Wu, Tianyi, et al.
Veröffentlicht: (2025) -
Enhancing Multi-Agent Debate System Performance via Confidence Expression
von: Lin, Zijie, et al.
Veröffentlicht: (2025)