UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Darvin, Liu, Teng, Terzolo, Mattie, Hasson, Lance, Sinha, Ayan, Mendes, Pablo, Rabinovich, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GraphMatch: Fusing Language and Graph Representations in a Dynamic Two-Sided Work Marketplace
by: Sacha, Mikołaj, et al.
Published: (2025)
by: Sacha, Mikołaj, et al.
Published: (2025)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026)
by: Yang, Jie, et al.
Published: (2026)
Haemophilus paragallinarum: Etiología de la coriza infecciosa
by: Horacio Raúl Terzolo
Published: (2004)
by: Horacio Raúl Terzolo
Published: (2004)
Epizootiología, prevención y control de la coriza infecciosa
by: Horacio Raúl Terzolo
Published: (2004)
by: Horacio Raúl Terzolo
Published: (2004)
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
by: He, Hang, et al.
Published: (2025)
by: He, Hang, et al.
Published: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Agentic Enterprise: AI-Centric User to User-Centric AI
by: Narechania, Arpit, et al.
Published: (2025)
by: Narechania, Arpit, et al.
Published: (2025)
LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Navigating the SLA/T conceptual landscape and investing in transdisciplinary practices
by: Ron Darvin
Published: (2025)
by: Ron Darvin
Published: (2025)
NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration
by: Twabi, Ahmed, et al.
Published: (2026)
by: Twabi, Ahmed, et al.
Published: (2026)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
Process-Centric Analysis of Agentic Software Systems
by: Liu, Shuyang, et al.
Published: (2025)
by: Liu, Shuyang, et al.
Published: (2025)
Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment
by: Zheng, Jingnan, et al.
Published: (2026)
by: Zheng, Jingnan, et al.
Published: (2026)
On the Robustness of Agentic Function Calling
by: Rabinovich, Ella, et al.
Published: (2025)
by: Rabinovich, Ella, et al.
Published: (2025)
Human-like Navigation in a World Built for Humans
by: Chandaka, Bhargav, et al.
Published: (2025)
by: Chandaka, Bhargav, et al.
Published: (2025)
RealBench: Benchmarking Verilog Generation Models with Real-World IP Designs
by: Jin, Pengwei, et al.
Published: (2025)
by: Jin, Pengwei, et al.
Published: (2025)
MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
by: Lei, Fei, et al.
Published: (2025)
by: Lei, Fei, et al.
Published: (2025)
On the Injectivity of Euler Integral Transforms with Hyperplanes and Quadric Hypersurfaces
by: Ji, Mattie
Published: (2023)
by: Ji, Mattie
Published: (2023)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
CleanUpBench: Embodied Sweeping and Grasping Benchmark
by: Li, Wenbo, et al.
Published: (2025)
by: Li, Wenbo, et al.
Published: (2025)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
RSA-Bench: Benchmarking Audio Large Models in Real-World Acoustic Scenarios
by: Zhang, Yibo, et al.
Published: (2026)
by: Zhang, Yibo, et al.
Published: (2026)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
by: Li, Chenxin, et al.
Published: (2026)
by: Li, Chenxin, et al.
Published: (2026)
Effect of Growing Up Poor on Labor Market Outcomes: Evidence From Indonesia
by: Mayang Rizky, et al.
Published: (2025)
by: Mayang Rizky, et al.
Published: (2025)
Personalized Model-Based Design of Human Centric AI enabled CPS for Long term usage
by: Ngabonziza, Bernard, et al.
Published: (2026)
by: Ngabonziza, Bernard, et al.
Published: (2026)
MobEvolve: An Agentic Self-Evolving Heuristic System for Interpretable Human Mobility Generation
by: He, Junlin, et al.
Published: (2026)
by: He, Junlin, et al.
Published: (2026)
From Clerks to Agentic-AI: How will Technology Change Labor Market in Finance?
by: Yu, Lu, et al.
Published: (2026)
by: Yu, Lu, et al.
Published: (2026)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
by: Wu, Meiqi, et al.
Published: (2026)
by: Wu, Meiqi, et al.
Published: (2026)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets
by: Zhang, Jaden, et al.
Published: (2026)
by: Zhang, Jaden, et al.
Published: (2026)
Adults’ Cognitive and Socioemotional Skills and Their Labor Market Outcomes in Colombia
by: Pablo Acosta
Published: (2020)
by: Pablo Acosta
Published: (2020)
Detection of Deployment Operational Deviations for Safety and Security of AI-Enabled Human-Centric Cyber Physical Systems
by: Ngabonziza, Bernard, et al.
Published: (2026)
by: Ngabonziza, Bernard, et al.
Published: (2026)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
by: Ding, Shuangrui, et al.
Published: (2026)
by: Ding, Shuangrui, et al.
Published: (2026)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
by: Zhang, Qiaohong, et al.
Published: (2026)
by: Zhang, Qiaohong, et al.
Published: (2026)
PPU-Bench:Real World Benchmark for Personalized Partial Unlearning in Vision Language Models
by: Guang, Jiahui, et al.
Published: (2026)
by: Guang, Jiahui, et al.
Published: (2026)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
by: Cui, Fan, et al.
Published: (2026)
by: Cui, Fan, et al.
Published: (2026)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026)
by: Shen, Yuanzhe, et al.
Published: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
Similar Items
-
GraphMatch: Fusing Language and Graph Representations in a Dynamic Two-Sided Work Marketplace
by: Sacha, Mikołaj, et al.
Published: (2025) -
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026) -
Haemophilus paragallinarum: Etiología de la coriza infecciosa
by: Horacio Raúl Terzolo
Published: (2004) -
Epizootiología, prevención y control de la coriza infecciosa
by: Horacio Raúl Terzolo
Published: (2004) -
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
by: He, Hang, et al.
Published: (2025)