CocoaBench: Evaluating Unified Digital Agents in the Wild
Fuente:
arXiv
Guardado en:
| Autores principales: | CocoaBench Team, Hao, Shibo, Zhang, Zhining, Liang, Zhiqi, Liu, Tianyang, Zha, Yuheng, Gao, Qiyue, Chen, Jixuan, Wang, Zilong, Cheng, Zhoujun, Zhang, Haoxiang, Wang, Junli, Jin, Hexi, Zheng, Boyuan, Zhou, Kun, Wang, Yu, Yao, Feng, Liu, Licheng, Li, Yijiang, Li, Zhifei, Han, Zhengtao, Promthaw, Pracha, Cerruti, Tommaso, Fu, Xiaohan, Ma, Ziqiao, Shang, Jingbo, Qin, Lianhui, McAuley, Julian, Xing, Eric P., Liu, Zhengzhong, Srivastava, Rupesh Kumar, Hu, Zhiting |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Simulating Task-Oriented Dialogues with State Transition Graphs and Large Language Models
por: Samarinas, Chris, et al.
Publicado: (2024)
por: Samarinas, Chris, et al.
Publicado: (2024)
Choosing Wisely and Learning Deeply: Selective Cross-Modality Distillation via CLIP for Domain Generalization
por: Leng, Jixuan, et al.
Publicado: (2023)
por: Leng, Jixuan, et al.
Publicado: (2023)
DeliveryBench: Can Agents Earn Profit in Real World?
por: Mao, Lingjun, et al.
Publicado: (2025)
por: Mao, Lingjun, et al.
Publicado: (2025)
Three central limit theorems for the unbounded excursion component of a Gaussian field
por: McAuley, Michael
Publicado: (2024)
por: McAuley, Michael
Publicado: (2024)
Children in Custody
por: McAuley, Mary
Publicado: (2022)
por: McAuley, Mary
Publicado: (2022)
Politics and the Soviet Union / Mary McAuley
por: McAuley, Mary
por: McAuley, Mary
Limit theorems for non-local functionals of smooth Gaussian fields via quasi-association
por: McAuley, Michael
Publicado: (2026)
por: McAuley, Michael
Publicado: (2026)
Cacau 2030 strategic guidelines: promotion of decent work and improvement of living conditions in the Brazilian cocoa productive chain
por: International Labour Organization., et al.
Publicado: (2021)
por: International Labour Organization., et al.
Publicado: (2021)
Multi-Behavior Generative Recommendation
por: Liu, Zihan, et al.
Publicado: (2024)
por: Liu, Zihan, et al.
Publicado: (2024)
Simultaneous state‐estimator tuning and parameter estimation for systems with nonstationary disturbances, multi‐rate data, and measurement delays
por: Qiujun A. Liu, et al.
Publicado: (2024)
por: Qiujun A. Liu, et al.
Publicado: (2024)
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
por: Hu, Yuanzhe, et al.
Publicado: (2025)
por: Hu, Yuanzhe, et al.
Publicado: (2025)
Mental-Gen: A Brain-Computer Interface-Based Interactive Method for Interior Space Generative Design
por: Liu, Yijiang, et al.
Publicado: (2024)
por: Liu, Yijiang, et al.
Publicado: (2024)
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
por: Kim, Haven, et al.
Publicado: (2026)
por: Kim, Haven, et al.
Publicado: (2026)
Limit theorems for the number of sign and level-set clusters of the Gaussian free field
por: McAuley, Michael, et al.
Publicado: (2025)
por: McAuley, Michael, et al.
Publicado: (2025)
Educating Young Children: A Structural Approach. Routledge Library Editions: Early Years
por: McAuley, Helen, et al.
Publicado: (2022)
por: McAuley, Helen, et al.
Publicado: (2022)
FASA: Frequency-aware Sparse Attention
por: Wang, Yifei, et al.
Publicado: (2026)
por: Wang, Yifei, et al.
Publicado: (2026)
Trustworthy deep domain adaptation for wearable photoplethysmography signal analysis with decision-theoretic uncertainty quantification
por: Bench, Ciaran
Publicado: (2026)
por: Bench, Ciaran
Publicado: (2026)
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
por: Liu, Yuhan, et al.
Publicado: (2025)
por: Liu, Yuhan, et al.
Publicado: (2025)
LVCHAT: Facilitating Long Video Comprehension
por: Wang, Yu, et al.
Publicado: (2024)
por: Wang, Yu, et al.
Publicado: (2024)
GSPRec: Temporal-Aware Graph Spectral Filtering for Recommendation
por: Rabiah, Ahmad Bin, et al.
Publicado: (2025)
por: Rabiah, Ahmad Bin, et al.
Publicado: (2025)
InstructGraph: Boosting Large Language Models via Graph-centric Instruction Tuning and Preference Alignment
por: Wang, Jianing, et al.
Publicado: (2024)
por: Wang, Jianing, et al.
Publicado: (2024)
Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
por: Wang, Ruoyu, et al.
Publicado: (2025)
por: Wang, Ruoyu, et al.
Publicado: (2025)
Train Once, Deploy Anywhere: Matryoshka Representation Learning for Multimodal Recommendation
por: Wang, Yueqi, et al.
Publicado: (2024)
por: Wang, Yueqi, et al.
Publicado: (2024)
Your Causal Self-Attentive Recommender Hosts a Lonely Neighborhood
por: Wang, Yueqi, et al.
Publicado: (2024)
por: Wang, Yueqi, et al.
Publicado: (2024)
Adaptive Change Point Inference for High Dimensional Time Series with Temporal Dependence
por: Wang, Xiaoyi, et al.
Publicado: (2025)
por: Wang, Xiaoyi, et al.
Publicado: (2025)
New Governing Equations for Fluid Dynamics
por: Liu, Chaoqun, et al.
Publicado: (2021)
por: Liu, Chaoqun, et al.
Publicado: (2021)
Self-Updatable Large Language Models by Integrating Context into Model Parameters
por: Wang, Yu, et al.
Publicado: (2024)
por: Wang, Yu, et al.
Publicado: (2024)
Det-SAM2:Technical Report on the Self-Prompting Segmentation Framework Based on Segment Anything Model 2
por: Wang, Zhiting, et al.
Publicado: (2024)
por: Wang, Zhiting, et al.
Publicado: (2024)
A covariance formula for the number of excursion set components of Gaussian fields and applications
por: Beliaev, Dmitry, et al.
Publicado: (2023)
por: Beliaev, Dmitry, et al.
Publicado: (2023)
A central limit theorem for the number of excursion set components of Gaussian fields
por: Beliaev, Dmitry, et al.
Publicado: (2022)
por: Beliaev, Dmitry, et al.
Publicado: (2022)
Extending Input Contexts of Language Models through Training on Segmented Sequences
por: Karypis, Petros, et al.
Publicado: (2023)
por: Karypis, Petros, et al.
Publicado: (2023)
On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
por: Bhat, Vishvesh, et al.
Publicado: (2025)
por: Bhat, Vishvesh, et al.
Publicado: (2025)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
por: Kim, Haven, et al.
Publicado: (2026)
por: Kim, Haven, et al.
Publicado: (2026)
Conformal welding of independent Gaussian multiplicative chaos measures
por: Kupiainen, Antti, et al.
Publicado: (2023)
por: Kupiainen, Antti, et al.
Publicado: (2023)
Bayesian parameter estimation using truncated normal distributions as priors for parameters in fundamental models of chemical processes
por: Lauren A. Gibson, et al.
Publicado: (2024)
por: Lauren A. Gibson, et al.
Publicado: (2024)
GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis
por: Liu, Haoyang, et al.
Publicado: (2025)
por: Liu, Haoyang, et al.
Publicado: (2025)
Cognitive Bias in Decision-Making with LLMs
por: Echterhoff, Jessica, et al.
Publicado: (2024)
por: Echterhoff, Jessica, et al.
Publicado: (2024)
On SkipGram Word Embedding Models with Negative Sampling: Unified Framework and Impact of Noise Distributions
por: Liu, Dezhi, et al.
Publicado: (2020)
por: Liu, Dezhi, et al.
Publicado: (2020)
Global dynamics of a stage-structured hantavirus infection model with seasonality
por: Liu Junli
Publicado: (2021)
por: Liu Junli
Publicado: (2021)
Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents
por: Ye, Chongrui, et al.
Publicado: (2026)
por: Ye, Chongrui, et al.
Publicado: (2026)
Ejemplares similares
-
Simulating Task-Oriented Dialogues with State Transition Graphs and Large Language Models
por: Samarinas, Chris, et al.
Publicado: (2024) -
Choosing Wisely and Learning Deeply: Selective Cross-Modality Distillation via CLIP for Domain Generalization
por: Leng, Jixuan, et al.
Publicado: (2023) -
DeliveryBench: Can Agents Earn Profit in Real World?
por: Mao, Lingjun, et al.
Publicado: (2025) -
Three central limit theorems for the unbounded excursion component of a Gaussian field
por: McAuley, Michael
Publicado: (2024) -
Children in Custody
por: McAuley, Mary
Publicado: (2022)