LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zeng, Weihao, Huang, Yuzhen, He, Junxian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
von: Huang, Yuzhen, et al.
Veröffentlicht: (2025)
von: Huang, Yuzhen, et al.
Veröffentlicht: (2025)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
von: Li, Junlong, et al.
Veröffentlicht: (2025)
von: Li, Junlong, et al.
Veröffentlicht: (2025)
FML-bench: Benchmarking Machine Learning Agents for Scientific Research
von: Zou, Qiran, et al.
Veröffentlicht: (2025)
von: Zou, Qiran, et al.
Veröffentlicht: (2025)
Compression Represents Intelligence Linearly
von: Huang, Yuzhen, et al.
Veröffentlicht: (2024)
von: Huang, Yuzhen, et al.
Veröffentlicht: (2024)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context
von: Xiu, Zidi, et al.
Veröffentlicht: (2026)
von: Xiu, Zidi, et al.
Veröffentlicht: (2026)
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol
von: Jobs, Niklas, et al.
Veröffentlicht: (2025)
von: Jobs, Niklas, et al.
Veröffentlicht: (2025)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
von: Zou, Qiran, et al.
Veröffentlicht: (2026)
von: Zou, Qiran, et al.
Veröffentlicht: (2026)
Multi-Agent Transfer Learning via Temporal Contrastive Learning
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
SWE-bench Goes Live!
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
Memory as Asset: From Agent-centric to Human-centric Memory Management
von: Pan, Yanqi, et al.
Veröffentlicht: (2026)
von: Pan, Yanqi, et al.
Veröffentlicht: (2026)
MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
von: Lei, Yuzhen, et al.
Veröffentlicht: (2025)
von: Lei, Yuzhen, et al.
Veröffentlicht: (2025)
MoSLD: An Extremely Parameter-Efficient Mixture-of-Shared LoRAs for Multi-Task Learning
von: Zhao, Lulu, et al.
Veröffentlicht: (2024)
von: Zhao, Lulu, et al.
Veröffentlicht: (2024)
CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
von: Kirmayr, Johannes, et al.
Veröffentlicht: (2026)
von: Kirmayr, Johannes, et al.
Veröffentlicht: (2026)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
von: Toledo, Edan, et al.
Veröffentlicht: (2025)
von: Toledo, Edan, et al.
Veröffentlicht: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Loosely-Structured Software: Engineering Context, Structure, and Evolution Entropy in Runtime-Rewired Multi-Agent Systems
von: Zhang, Weihao, et al.
Veröffentlicht: (2026)
von: Zhang, Weihao, et al.
Veröffentlicht: (2026)
AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool-Environment-Agent(TEA) Protocol
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
von: Kytöniemi, Joona, et al.
Veröffentlicht: (2025)
von: Kytöniemi, Joona, et al.
Veröffentlicht: (2025)
Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism
von: Bu, Tao, et al.
Veröffentlicht: (2025)
von: Bu, Tao, et al.
Veröffentlicht: (2025)
DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
von: Shu, Fan, et al.
Veröffentlicht: (2026)
von: Shu, Fan, et al.
Veröffentlicht: (2026)
LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
von: Wang, Ruipeng, et al.
Veröffentlicht: (2026)
von: Wang, Ruipeng, et al.
Veröffentlicht: (2026)
Harnessing Language for Coordination: A Framework and Benchmark for LLM-Driven Multi-Agent Control
von: Anne, Timothée, et al.
Veröffentlicht: (2024)
von: Anne, Timothée, et al.
Veröffentlicht: (2024)
PredictionMarketBench: A SWE-bench-Style Framework for Backtesting Trading Agents on Prediction Markets
von: Arora, Avi, et al.
Veröffentlicht: (2026)
von: Arora, Avi, et al.
Veröffentlicht: (2026)
CareBot: A Pioneering Full-Process Open-Source Medical Language Model
von: Zhao, Lulu, et al.
Veröffentlicht: (2024)
von: Zhao, Lulu, et al.
Veröffentlicht: (2024)
Non-myopic Generation of Language Models for Reasoning and Planning
von: Ma, Chang, et al.
Veröffentlicht: (2024)
von: Ma, Chang, et al.
Veröffentlicht: (2024)
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
von: Huang, Jiahao, et al.
Veröffentlicht: (2026)
von: Huang, Jiahao, et al.
Veröffentlicht: (2026)
LVBench: An Extreme Long Video Understanding Benchmark
von: Wang, Weihan, et al.
Veröffentlicht: (2024)
von: Wang, Weihan, et al.
Veröffentlicht: (2024)
What Shapes a Creative Machine Mind? Comprehensively Benchmarking Creativity in Foundation Models
von: He, Zicong, et al.
Veröffentlicht: (2025)
von: He, Zicong, et al.
Veröffentlicht: (2025)
Empowering Working Memory for Large Language Model Agents
von: Guo, Jing, et al.
Veröffentlicht: (2023)
von: Guo, Jing, et al.
Veröffentlicht: (2023)
DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
von: Xu, Tianqi, et al.
Veröffentlicht: (2024)
von: Xu, Tianqi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
von: Huang, Yuzhen, et al.
Veröffentlicht: (2025) -
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
von: Zeng, Weihao, et al.
Veröffentlicht: (2024) -
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
von: Zeng, Weihao, et al.
Veröffentlicht: (2025) -
Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
von: Zeng, Weihao, et al.
Veröffentlicht: (2025) -
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
von: Li, Junlong, et al.
Veröffentlicht: (2025)