Scaling Test-time Compute for LLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, King, Li, Hanhao, Wu, Siwei, Xing, Tianshun, Ma, Dehua, Tang, Xiangru, Liu, Minghao, Yang, Jian, Liu, Jiaheng, Jiang, Yuchen Eleanor, Zhang, Changwang, Lin, Chenghua, Wang, Jun, Zhang, Ge, Zhou, Wangchunshu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OAgents: An Empirical Study of Building Effective Agents
di: Zhu, He, et al.
Pubblicazione: (2025)
di: Zhu, He, et al.
Pubblicazione: (2025)
Efficient Agents: Building Effective Agents While Reducing Cost
di: Wang, Ningning, et al.
Pubblicazione: (2025)
di: Wang, Ningning, et al.
Pubblicazione: (2025)
TaskCraft: Automated Generation of Agentic Tasks
di: Shi, Dingfeng, et al.
Pubblicazione: (2025)
di: Shi, Dingfeng, et al.
Pubblicazione: (2025)
Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
di: Qin, Tianrui, et al.
Pubblicazione: (2025)
di: Qin, Tianrui, et al.
Pubblicazione: (2025)
A$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
di: Chen, Qianben, et al.
Pubblicazione: (2025)
di: Chen, Qianben, et al.
Pubblicazione: (2025)
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
di: Li, Weizhen, et al.
Pubblicazione: (2025)
di: Li, Weizhen, et al.
Pubblicazione: (2025)
ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
di: Gui, Xin, et al.
Pubblicazione: (2025)
di: Gui, Xin, et al.
Pubblicazione: (2025)
OThink-SRR1: Search, Refine and Reasoning with Reinforced Learning for Large Language Models
di: Liang, Haijian, et al.
Pubblicazione: (2026)
di: Liang, Haijian, et al.
Pubblicazione: (2026)
AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
di: Li, Ziming, et al.
Pubblicazione: (2024)
di: Li, Ziming, et al.
Pubblicazione: (2024)
AFSPP: Agent Framework for Shaping Preference and Personality with Large Language Models
di: He, Zihong, et al.
Pubblicazione: (2024)
di: He, Zihong, et al.
Pubblicazione: (2024)
TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
di: Li, Yizhi, et al.
Pubblicazione: (2025)
di: Li, Yizhi, et al.
Pubblicazione: (2025)
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
di: Qu, Xingwei, et al.
Pubblicazione: (2024)
di: Qu, Xingwei, et al.
Pubblicazione: (2024)
A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
di: Wu, Siwei, et al.
Pubblicazione: (2024)
di: Wu, Siwei, et al.
Pubblicazione: (2024)
AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
di: Qiao, Shuofei, et al.
Pubblicazione: (2024)
di: Qiao, Shuofei, et al.
Pubblicazione: (2024)
How Far Are We from Genuinely Useful Deep Research Agents?
di: Zhang, Dingling, et al.
Pubblicazione: (2025)
di: Zhang, Dingling, et al.
Pubblicazione: (2025)
TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddings
di: Wu, Yebo, et al.
Pubblicazione: (2026)
di: Wu, Yebo, et al.
Pubblicazione: (2026)
MMUEChange: A Generalized LLM Agent Framework for Intelligent Multi-Modal Urban Environment Change Analysis
di: Xiao, Zixuan, et al.
Pubblicazione: (2026)
di: Xiao, Zixuan, et al.
Pubblicazione: (2026)
Towards Personalized Deep Research: Benchmarks and Evaluations
di: Liang, Yuan, et al.
Pubblicazione: (2025)
di: Liang, Yuan, et al.
Pubblicazione: (2025)
MIMIR: A Streamlined Platform for Personalized Agent Tuning in Domain Expertise
di: Deng, Chunyuan, et al.
Pubblicazione: (2024)
di: Deng, Chunyuan, et al.
Pubblicazione: (2024)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
di: Tang, Xiangru, et al.
Pubblicazione: (2025)
di: Tang, Xiangru, et al.
Pubblicazione: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
di: Wu, Qinzhuo, et al.
Pubblicazione: (2026)
di: Wu, Qinzhuo, et al.
Pubblicazione: (2026)
Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
di: Chen, Qianben, et al.
Pubblicazione: (2026)
di: Chen, Qianben, et al.
Pubblicazione: (2026)
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
di: Liu, Runze, et al.
Pubblicazione: (2025)
di: Liu, Runze, et al.
Pubblicazione: (2025)
Hybrid Epidemics - A Case Study on Computer Worm Conficker
di: Zhang, Changwang, et al.
Pubblicazione: (2014)
di: Zhang, Changwang, et al.
Pubblicazione: (2014)
COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
di: P Team, et al.
Pubblicazione: (2025)
di: P Team, et al.
Pubblicazione: (2025)
On the Structural Memory of LLM Agents
di: Zeng, Ruihong, et al.
Pubblicazione: (2024)
di: Zeng, Ruihong, et al.
Pubblicazione: (2024)
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
di: Hu, Xavier, et al.
Pubblicazione: (2026)
di: Hu, Xavier, et al.
Pubblicazione: (2026)
Does LLM Focus on the Right Words? Mitigating Context Bias in LLM-based Recommenders
di: Wang, Bohao, et al.
Pubblicazione: (2025)
di: Wang, Bohao, et al.
Pubblicazione: (2025)
OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025)
IFEvalCode: Controlled Code Generation
di: Yang, Jian, et al.
Pubblicazione: (2025)
di: Yang, Jian, et al.
Pubblicazione: (2025)
SpecTran: Spectral-Aware Transformer-based Adapter for LLM-Enhanced Sequential Recommendation
di: Cui, Yu, et al.
Pubblicazione: (2026)
di: Cui, Yu, et al.
Pubblicazione: (2026)
ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
di: Wen, Hao, et al.
Pubblicazione: (2025)
di: Wen, Hao, et al.
Pubblicazione: (2025)
LogiAgent: Automated Logical Testing for REST Systems with LLM-Based Multi-Agents
di: Zhang, Ke, et al.
Pubblicazione: (2025)
di: Zhang, Ke, et al.
Pubblicazione: (2025)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
di: Trae Research Team, et al.
Pubblicazione: (2025)
di: Trae Research Team, et al.
Pubblicazione: (2025)
Human and AI Perceptual Differences in Image Classification Errors
di: Liu, Minghao, et al.
Pubblicazione: (2023)
di: Liu, Minghao, et al.
Pubblicazione: (2023)
MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
di: Ding, Dongyi, et al.
Pubblicazione: (2025)
di: Ding, Dongyi, et al.
Pubblicazione: (2025)
PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
di: Tao, Meiling, et al.
Pubblicazione: (2025)
di: Tao, Meiling, et al.
Pubblicazione: (2025)
Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
di: Zhu, Chenghao, et al.
Pubblicazione: (2025)
di: Zhu, Chenghao, et al.
Pubblicazione: (2025)
OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation
di: Zhang, Shengjia, et al.
Pubblicazione: (2025)
di: Zhang, Shengjia, et al.
Pubblicazione: (2025)
Documenti analoghi
-
OAgents: An Empirical Study of Building Effective Agents
di: Zhu, He, et al.
Pubblicazione: (2025) -
Efficient Agents: Building Effective Agents While Reducing Cost
di: Wang, Ningning, et al.
Pubblicazione: (2025) -
TaskCraft: Automated Generation of Agentic Tasks
di: Shi, Dingfeng, et al.
Pubblicazione: (2025) -
Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
di: Qin, Tianrui, et al.
Pubblicazione: (2025) -
A$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
di: Chen, Qianben, et al.
Pubblicazione: (2025)