Gespeichert in:
| Hauptverfasser: | Cai, Xiaoran, Yang, Wang, Ren, Xiyu, Law, Chekun, Sharma, Rohit, Qi, Peng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.17106 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Human-AI Collaboration: A Review and Methodological Framework
von: Fragiadakis, George, et al.
Veröffentlicht: (2024)
von: Fragiadakis, George, et al.
Veröffentlicht: (2024)
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
von: Ye, Bowen, et al.
Veröffentlicht: (2026)
von: Ye, Bowen, et al.
Veröffentlicht: (2026)
Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation
von: Wu, Zonghan, et al.
Veröffentlicht: (2025)
von: Wu, Zonghan, et al.
Veröffentlicht: (2025)
Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning
von: Linze, Chen, et al.
Veröffentlicht: (2026)
von: Linze, Chen, et al.
Veröffentlicht: (2026)
Towards Trustworthy Legal AI through LLM Agents and Formal Reasoning
von: Chen, Linze, et al.
Veröffentlicht: (2025)
von: Chen, Linze, et al.
Veröffentlicht: (2025)
Improving Health Professionals' Onboarding with AI and XAI for Trustworthy Human-AI Collaborative Decision Making
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
Ethical AI: Towards Defining a Collective Evaluation Framework
von: Sharma, Aasish Kumar, et al.
Veröffentlicht: (2025)
von: Sharma, Aasish Kumar, et al.
Veröffentlicht: (2025)
A Study on the Framework for Evaluating the Ethics and Trustworthiness of Generative AI
von: Jeong, Cheonsu, et al.
Veröffentlicht: (2025)
von: Jeong, Cheonsu, et al.
Veröffentlicht: (2025)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
von: Huang, Jinsheng, et al.
Veröffentlicht: (2024)
von: Huang, Jinsheng, et al.
Veröffentlicht: (2024)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
Explainable AI for Maritime Autonomous Surface Ships (MASS): Adaptive Interfaces and Trustworthy Human-AI Collaboration
von: Zhang, Zhuoyue, et al.
Veröffentlicht: (2025)
von: Zhang, Zhuoyue, et al.
Veröffentlicht: (2025)
Trustworthy Human-AI Collaboration: Reinforcement Learning with Human Feedback and Physics Knowledge for Safe Autonomous Driving
von: Huang, Zilin, et al.
Veröffentlicht: (2024)
von: Huang, Zilin, et al.
Veröffentlicht: (2024)
Designing The Internet of Agents: A Framework for Trustworthy, Transparent, and Collaborative Human-Agent Interaction (HAX)
von: Scibelli, Marc, et al.
Veröffentlicht: (2025)
von: Scibelli, Marc, et al.
Veröffentlicht: (2025)
Trust the AI, Doubt Yourself: The Effect of Urgency on Self-Confidence in Human-AI Interaction
von: Shajari, Baran, et al.
Veröffentlicht: (2026)
von: Shajari, Baran, et al.
Veröffentlicht: (2026)
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering
von: Dong, Tao, et al.
Veröffentlicht: (2025)
von: Dong, Tao, et al.
Veröffentlicht: (2025)
AI Benchmarks and Datasets for LLM Evaluation
von: Ivanov, Todor, et al.
Veröffentlicht: (2024)
von: Ivanov, Todor, et al.
Veröffentlicht: (2024)
Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
von: Machlovi, Naseem, et al.
Veröffentlicht: (2025)
von: Machlovi, Naseem, et al.
Veröffentlicht: (2025)
Human-Centered Human-AI Collaboration (HCHAC)
von: Gao, Qi, et al.
Veröffentlicht: (2025)
von: Gao, Qi, et al.
Veröffentlicht: (2025)
Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs
von: Xiao, Changrong, et al.
Veröffentlicht: (2024)
von: Xiao, Changrong, et al.
Veröffentlicht: (2024)
Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks
von: Syed, Toqeer Ali, et al.
Veröffentlicht: (2025)
von: Syed, Toqeer Ali, et al.
Veröffentlicht: (2025)
Decidable By Construction: Design-Time Verification for Trustworthy AI
von: Haynes, Houston
Veröffentlicht: (2026)
von: Haynes, Houston
Veröffentlicht: (2026)
Causal Responsibility Attribution for Human-AI Collaboration
von: Qi, Yahang, et al.
Veröffentlicht: (2024)
von: Qi, Yahang, et al.
Veröffentlicht: (2024)
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
von: Qi, Jinhu, et al.
Veröffentlicht: (2026)
von: Qi, Jinhu, et al.
Veröffentlicht: (2026)
AI Simulation by Digital Twins: Systematic Survey, Reference Framework, and Mapping to a Standardized Architecture
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration
von: Calboreanu, Elias
Veröffentlicht: (2026)
von: Calboreanu, Elias
Veröffentlicht: (2026)
NavTrust: Benchmarking Trustworthiness for Embodied Navigation
von: Jiang, Huaide, et al.
Veröffentlicht: (2026)
von: Jiang, Huaide, et al.
Veröffentlicht: (2026)
HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark
von: Go, Seonghyeon, et al.
Veröffentlicht: (2026)
von: Go, Seonghyeon, et al.
Veröffentlicht: (2026)
Towards a Comprehensive Human-Centred Evaluation Framework for Explainable AI
von: Donoso-Guzmán, Ivania, et al.
Veröffentlicht: (2023)
von: Donoso-Guzmán, Ivania, et al.
Veröffentlicht: (2023)
UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
von: Mao, Qi, et al.
Veröffentlicht: (2025)
von: Mao, Qi, et al.
Veröffentlicht: (2025)
Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems
von: Cai, Yicheng, et al.
Veröffentlicht: (2026)
von: Cai, Yicheng, et al.
Veröffentlicht: (2026)
Advancing Trustworthy AI for Sustainable Development: Recommendations for Standardising AI Incident Reporting
von: Agarwal, Avinash, et al.
Veröffentlicht: (2025)
von: Agarwal, Avinash, et al.
Veröffentlicht: (2025)
The Journey to Trustworthy AI: Pursuit of Pragmatic Frameworks
von: Nasr-Azadani, Mohamad M, et al.
Veröffentlicht: (2024)
von: Nasr-Azadani, Mohamad M, et al.
Veröffentlicht: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
von: Yang, Chao, et al.
Veröffentlicht: (2024)
von: Yang, Chao, et al.
Veröffentlicht: (2024)
Reliability, Resilience and Human Factors Engineering for Trustworthy AI Systems
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)
Trustworthy and Responsible AI for Human-Centric Autonomous Decision-Making Systems
von: Dehghani, Farzaneh, et al.
Veröffentlicht: (2024)
von: Dehghani, Farzaneh, et al.
Veröffentlicht: (2024)
Bridging the Communication Gap: Evaluating AI Labeling Practices for Trustworthy AI Development
von: Fischer, Raphael, et al.
Veröffentlicht: (2025)
von: Fischer, Raphael, et al.
Veröffentlicht: (2025)
Towards Responsible AI Music: an Investigation of Trustworthy Features for Creative Systems
von: de Berardinis, Jacopo, et al.
Veröffentlicht: (2025)
von: de Berardinis, Jacopo, et al.
Veröffentlicht: (2025)
A No Free Lunch Theorem for Human-AI Collaboration
von: Peng, Kenny, et al.
Veröffentlicht: (2024)
von: Peng, Kenny, et al.
Veröffentlicht: (2024)
Agentic AI as Undercover Teammates: Argumentative Knowledge Construction in Hybrid Human-AI Collaborative Learning
von: Yan, Lixiang, et al.
Veröffentlicht: (2025)
von: Yan, Lixiang, et al.
Veröffentlicht: (2025)
A Knowledge-Component-Based Methodology for Evaluating AI Assistants
von: Qi, Laryn, et al.
Veröffentlicht: (2024)
von: Qi, Laryn, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating Human-AI Collaboration: A Review and Methodological Framework
von: Fragiadakis, George, et al.
Veröffentlicht: (2024) -
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
von: Ye, Bowen, et al.
Veröffentlicht: (2026) -
Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation
von: Wu, Zonghan, et al.
Veröffentlicht: (2025) -
Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning
von: Linze, Chen, et al.
Veröffentlicht: (2026) -
Towards Trustworthy Legal AI through LLM Agents and Formal Reasoning
von: Chen, Linze, et al.
Veröffentlicht: (2025)