Benchmark Test-Time Scaling of General LLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Xiaochuan, Ming, Ryan, Setlur, Pranav, Paladugu, Abhijay, Tang, Andy, Kang, Hao, Shao, Shuai, Jin, Rong, Xiong, Chenyan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
di: Jin, Jiahe, et al.
Pubblicazione: (2025)
di: Jin, Jiahe, et al.
Pubblicazione: (2025)
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
di: Kang, Hao, et al.
Pubblicazione: (2024)
di: Kang, Hao, et al.
Pubblicazione: (2024)
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
di: Li, Xiaochuan, et al.
Pubblicazione: (2024)
di: Li, Xiaochuan, et al.
Pubblicazione: (2024)
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
di: Coelho, João, et al.
Pubblicazione: (2025)
di: Coelho, João, et al.
Pubblicazione: (2025)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
di: Yu, Zichun, et al.
Pubblicazione: (2026)
di: Yu, Zichun, et al.
Pubblicazione: (2026)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
di: Kang, Hao, et al.
Pubblicazione: (2025)
di: Kang, Hao, et al.
Pubblicazione: (2025)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
di: Setlur, Amrith, et al.
Pubblicazione: (2025)
di: Setlur, Amrith, et al.
Pubblicazione: (2025)
Craw4LLM: Efficient Web Crawling for LLM Pretraining
di: Yu, Shi, et al.
Pubblicazione: (2025)
di: Yu, Shi, et al.
Pubblicazione: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
di: Ning, Jingjie, et al.
Pubblicazione: (2026)
di: Ning, Jingjie, et al.
Pubblicazione: (2026)
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
di: Ivanov, Maksim, et al.
Pubblicazione: (2026)
di: Ivanov, Maksim, et al.
Pubblicazione: (2026)
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
di: Wang, Tevin, et al.
Pubblicazione: (2024)
di: Wang, Tevin, et al.
Pubblicazione: (2024)
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
di: Yu, Zichun, et al.
Pubblicazione: (2025)
di: Yu, Zichun, et al.
Pubblicazione: (2025)
Intercept Cancer: Cancer Pre-Screening with Large Scale Healthcare Foundation Models
di: Sun, Liwen, et al.
Pubblicazione: (2025)
di: Sun, Liwen, et al.
Pubblicazione: (2025)
PithTrain: A Compact and Agent-Native MoE Training System
di: Lai, Ruihang, et al.
Pubblicazione: (2026)
di: Lai, Ruihang, et al.
Pubblicazione: (2026)
Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
di: Wang, Shuting, et al.
Pubblicazione: (2025)
di: Wang, Shuting, et al.
Pubblicazione: (2025)
Generate, Not Recommend: Personalized Multimodal Content Generation
di: Liu, Jiongnan, et al.
Pubblicazione: (2025)
di: Liu, Jiongnan, et al.
Pubblicazione: (2025)
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
di: Wang, Tevin, et al.
Pubblicazione: (2025)
di: Wang, Tevin, et al.
Pubblicazione: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
di: Sun, Shengyin, et al.
Pubblicazione: (2025)
di: Sun, Shengyin, et al.
Pubblicazione: (2025)
Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation
di: Sun, Liwen, et al.
Pubblicazione: (2024)
di: Sun, Liwen, et al.
Pubblicazione: (2024)
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search
di: Shi, Wentao, et al.
Pubblicazione: (2025)
di: Shi, Wentao, et al.
Pubblicazione: (2025)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
di: Bi, Zhenni, et al.
Pubblicazione: (2024)
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
di: Zhong, Shanshan, et al.
Pubblicazione: (2026)
di: Zhong, Shanshan, et al.
Pubblicazione: (2026)
Midtraining Bridges Pretraining and Posttraining Distributions
di: Liu, Emmy, et al.
Pubblicazione: (2025)
di: Liu, Emmy, et al.
Pubblicazione: (2025)
Semi-structured LLM Reasoners Can Be Rigorously Audited
di: Leng, Jixuan, et al.
Pubblicazione: (2025)
di: Leng, Jixuan, et al.
Pubblicazione: (2025)
Test-Time Scaling with Reflective Generative Model
di: Wang, Zixiao, et al.
Pubblicazione: (2025)
di: Wang, Zixiao, et al.
Pubblicazione: (2025)
ORBIT -- Open Recommendation Benchmark for Reproducible Research with Hidden Tests
di: He, Jingyuan, et al.
Pubblicazione: (2025)
di: He, Jingyuan, et al.
Pubblicazione: (2025)
Agentic Test-Time Scaling for WebAgents
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
VerifierQ: Enhancing LLM Test Time Compute with Q-Learning-based Verifiers
di: Qi, Jianing, et al.
Pubblicazione: (2024)
di: Qi, Jianing, et al.
Pubblicazione: (2024)
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
di: Yu, Zichun, et al.
Pubblicazione: (2024)
di: Yu, Zichun, et al.
Pubblicazione: (2024)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
di: Wei, Tianxin, et al.
Pubblicazione: (2025)
di: Wei, Tianxin, et al.
Pubblicazione: (2025)
Self-Improving LLM Agents at Test-Time
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
di: Fan, Qihui, et al.
Pubblicazione: (2026)
di: Fan, Qihui, et al.
Pubblicazione: (2026)
Linking Knowledge to Care: Knowledge Graph-Augmented Medical Follow-Up Question Generation
di: Sun, Liwen, et al.
Pubblicazione: (2026)
di: Sun, Liwen, et al.
Pubblicazione: (2026)
Inverse Scaling in Test-Time Compute
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2025)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2025)
Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
di: Kuang, Peng, et al.
Pubblicazione: (2025)
di: Kuang, Peng, et al.
Pubblicazione: (2025)
Select, Read, and Write: A Multi-Agent Framework of Full-Text-based Related Work Generation
di: Liu, Xiaochuan, et al.
Pubblicazione: (2025)
di: Liu, Xiaochuan, et al.
Pubblicazione: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
di: Jin, Jiahe, et al.
Pubblicazione: (2025) -
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
di: Kang, Hao, et al.
Pubblicazione: (2024) -
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
di: Li, Xiaochuan, et al.
Pubblicazione: (2024) -
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
di: Coelho, João, et al.
Pubblicazione: (2025) -
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
di: Yu, Zichun, et al.
Pubblicazione: (2026)