CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Luo, Hanjun, Ni, Chiming, Wen, Jiaheng, Huang, Zhimu, Wang, Yiran, Liao, Bingduo, Chung, Sylvia, Jin, Yingbin, Li, Xinfeng, Xu, Wenyuan, Wang, XiaoFeng, Salam, Hanan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters
par: Luo, Hanjun, et autres
Publié: (2026)
par: Luo, Hanjun, et autres
Publié: (2026)
BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models
par: Luo, Hanjun, et autres
Publié: (2026)
par: Luo, Hanjun, et autres
Publié: (2026)
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
par: Luo, Hanjun, et autres
Publié: (2025)
par: Luo, Hanjun, et autres
Publié: (2025)
BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models
par: Luo, Hanjun, et autres
Publié: (2024)
par: Luo, Hanjun, et autres
Publié: (2024)
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction
par: Li, Jialin, et autres
Publié: (2026)
par: Li, Jialin, et autres
Publié: (2026)
CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making
par: Rahman, Hasibur, et autres
Publié: (2025)
par: Rahman, Hasibur, et autres
Publié: (2025)
Supporting Productivity Skill Development in College Students through Social Robot Coaching: A Proof-of-Concept
par: Lalwani, Himanshi, et autres
Publié: (2025)
par: Lalwani, Himanshi, et autres
Publié: (2025)
The Supportiveness-Safety Tradeoff in LLM Well-Being Agents
par: Lalwani, Himanshi, et autres
Publié: (2026)
par: Lalwani, Himanshi, et autres
Publié: (2026)
Ethically-Aware Participatory Design of a Productivity Social Robot for College Students
par: Lalwani, Himanshi, et autres
Publié: (2025)
par: Lalwani, Himanshi, et autres
Publié: (2025)
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
par: Sarangi, Sneheel, et autres
Publié: (2025)
par: Sarangi, Sneheel, et autres
Publié: (2025)
DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity Recognition
par: Luo, Hanjun, et autres
Publié: (2024)
par: Luo, Hanjun, et autres
Publié: (2024)
AC-PKAN: Attention-Enhanced and Chebyshev Polynomial-Based Physics-Informed Kolmogorov-Arnold Networks
par: Zhang, Hangwei, et autres
Publié: (2025)
par: Zhang, Hangwei, et autres
Publié: (2025)
McEval: Massively Multilingual Code Evaluation
par: Chai, Linzheng, et autres
Publié: (2024)
par: Chai, Linzheng, et autres
Publié: (2024)
Improving Personalisation in Valence and Arousal Prediction using Data Augmentation
par: Nwadike, Munachiso, et autres
Publié: (2024)
par: Nwadike, Munachiso, et autres
Publié: (2024)
GROW: A Conversational AI Coach for Goals, Reflection, Optimism, and Well-Being
par: Shah, Keya, et autres
Publié: (2026)
par: Shah, Keya, et autres
Publié: (2026)
Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition
par: Sarangi, Sneheel, et autres
Publié: (2025)
par: Sarangi, Sneheel, et autres
Publié: (2025)
Correlated Errors Challenge Vulnerable Growth
par: Wei Long, et autres
Publié: (2025)
par: Wei Long, et autres
Publié: (2025)
Evolutionary Processes in the Centaur Region
par: Kokotanekova, Rosita, et autres
Publié: (2025)
par: Kokotanekova, Rosita, et autres
Publié: (2025)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
par: Feng, Jia, et autres
Publié: (2024)
par: Feng, Jia, et autres
Publié: (2024)
StackEval: Benchmarking LLMs in Coding Assistance
par: Shah, Nidhish, et autres
Publié: (2024)
par: Shah, Nidhish, et autres
Publié: (2024)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
par: Du, Junjia, et autres
Publié: (2025)
par: Du, Junjia, et autres
Publié: (2025)
On Randomness in Agentic Evals
par: Bjarnason, Bjarni Haukur, et autres
Publié: (2026)
par: Bjarnason, Bjarni Haukur, et autres
Publié: (2026)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
par: Wang, Yixu, et autres
Publié: (2025)
par: Wang, Yixu, et autres
Publié: (2025)
MdEval: Massively Multilingual Code Debugging
par: Liu, Shukai, et autres
Publié: (2024)
par: Liu, Shukai, et autres
Publié: (2024)
Malla: Demystifying Real-world Large Language Model Integrated Malicious Services
par: Lin, Zilong, et autres
Publié: (2024)
par: Lin, Zilong, et autres
Publié: (2024)
Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes
par: Lin, Zilong, et autres
Publié: (2025)
par: Lin, Zilong, et autres
Publié: (2025)
Group of Hercules slaying the Centaur Nessus
par: Scan-the-World
Publié: (2026)
par: Scan-the-World
Publié: (2026)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
par: Zheng, Qinkai, et autres
Publié: (2023)
par: Zheng, Qinkai, et autres
Publié: (2023)
AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
par: Li, Kai, et autres
Publié: (2025)
par: Li, Kai, et autres
Publié: (2025)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
par: Kuhar, Sachit, et autres
Publié: (2024)
par: Kuhar, Sachit, et autres
Publié: (2024)
Reconstruction of the observable universe from the integrated Sachs-Wolfe effect
par: Chung, Julianne, et autres
Publié: (2025)
par: Chung, Julianne, et autres
Publié: (2025)
Comparative Analysis of Open-Source Language Models in Summarizing Medical Text Data
par: Chen, Yuhao, et autres
Publié: (2024)
par: Chen, Yuhao, et autres
Publié: (2024)
Centaurs, Rioting in Thessaly: Memory and the Classical World
par: Hudson, Martyn
Publié: (2019)
par: Hudson, Martyn
Publié: (2019)
High-inclination Centaur reservoirs beyond Neptune
par: Namouni, Fathi
Publié: (2025)
par: Namouni, Fathi
Publié: (2025)
Centaur: a foundation model of human cognition
par: Binz, Marcel, et autres
Publié: (2024)
par: Binz, Marcel, et autres
Publié: (2024)
The Near-Centaur Environment: Satellites, Rings, and Debris
par: Sickafoose, A. A., et autres
Publié: (2025)
par: Sickafoose, A. A., et autres
Publié: (2025)
Centaur Nuclei: Sizes, Shapes, Spins, and Structure
par: Fernandez, Y. R., et autres
Publié: (2025)
par: Fernandez, Y. R., et autres
Publié: (2025)
Effective Generative AI: The Human-Algorithm Centaur
par: Saghafian, Soroush, et autres
Publié: (2024)
par: Saghafian, Soroush, et autres
Publié: (2024)
Value Maximization under Stochastic Quasi-Hyperbolic Discounting
par: Yan, Kaixin, et autres
Publié: (2024)
par: Yan, Kaixin, et autres
Publié: (2024)
Documents similaires
-
AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters
par: Luo, Hanjun, et autres
Publié: (2026) -
BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models
par: Luo, Hanjun, et autres
Publié: (2026) -
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
par: Luo, Hanjun, et autres
Publié: (2025) -
BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models
par: Luo, Hanjun, et autres
Publié: (2024) -
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction
par: Li, Jialin, et autres
Publié: (2026)