Making AI Evaluation Deployment Relevant Through Context Specification
Fuente:
arXiv
Salvato in:
| Autori principali: | Holmes, Matthew, Lacerda, Thiago, Schwartz, Reva |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma
di: Schwartz, Reva, et al.
Pubblicazione: (2026)
di: Schwartz, Reva, et al.
Pubblicazione: (2026)
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
di: Schwartz, Reva, et al.
Pubblicazione: (2026)
di: Schwartz, Reva, et al.
Pubblicazione: (2026)
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
di: Schwartz, Reva, et al.
Pubblicazione: (2025)
di: Schwartz, Reva, et al.
Pubblicazione: (2025)
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
di: Kryshtal, Andrii
Pubblicazione: (2026)
di: Kryshtal, Andrii
Pubblicazione: (2026)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
di: Griffin, Charlie, et al.
Pubblicazione: (2024)
di: Griffin, Charlie, et al.
Pubblicazione: (2024)
Dynamic Context-Aware Prompt Recommendation for Domain-Specific AI Applications
di: Tang, Xinye, et al.
Pubblicazione: (2025)
di: Tang, Xinye, et al.
Pubblicazione: (2025)
12 Angry AI Agents: Evaluating Multi-Agent LLM Decision-Making Through Cinematic Jury Deliberation
di: Ersoz, Ahmet Bahaddin
Pubblicazione: (2026)
di: Ersoz, Ahmet Bahaddin
Pubblicazione: (2026)
Behavioral Determinants of Deployed AI Agents in Social Networks: A Multi-Factor Study of Personality, Model, and Guardrail Specification
di: Wilson, Sarah, et al.
Pubblicazione: (2026)
di: Wilson, Sarah, et al.
Pubblicazione: (2026)
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
di: Vishwarupe, Varad, et al.
Pubblicazione: (2026)
di: Vishwarupe, Varad, et al.
Pubblicazione: (2026)
Generation, Evaluation, and Explanation of Novelists' Styles with Single-Token Prompts
di: Rezaei, Mosab, et al.
Pubblicazione: (2025)
di: Rezaei, Mosab, et al.
Pubblicazione: (2025)
Internal Deployment Gaps in AI Regulation
di: Kwon, Joe, et al.
Pubblicazione: (2026)
di: Kwon, Joe, et al.
Pubblicazione: (2026)
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
di: Gallego, Víctor
Pubblicazione: (2025)
di: Gallego, Víctor
Pubblicazione: (2025)
A Field Guide to Deploying AI Agents in Clinical Practice
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
DAO-AI: Evaluating Collective Decision-Making through Agentic AI in Decentralized Governance
di: Capponi, Agostino, et al.
Pubblicazione: (2025)
di: Capponi, Agostino, et al.
Pubblicazione: (2025)
Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
di: Zhang, Hongxiang, et al.
Pubblicazione: (2026)
di: Zhang, Hongxiang, et al.
Pubblicazione: (2026)
RAN Cortex: Memory-Augmented Intelligence for Context-Aware Decision-Making in AI-Native Networks
di: Barros, Sebastian
Pubblicazione: (2025)
di: Barros, Sebastian
Pubblicazione: (2025)
Bridging Protocol and Production: Design Patterns for Deploying AI Agents with Model Context Protocol
di: Srinivasan, Vasundra
Pubblicazione: (2026)
di: Srinivasan, Vasundra
Pubblicazione: (2026)
PATHWAYS: Evaluating Investigation and Context Discovery in AI Web Agents
di: Arman, Shifat E., et al.
Pubblicazione: (2026)
di: Arman, Shifat E., et al.
Pubblicazione: (2026)
Safety Must Precede the Deployment of Open-Ended AI
di: Sheth, Ivaxi, et al.
Pubblicazione: (2025)
di: Sheth, Ivaxi, et al.
Pubblicazione: (2025)
Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning
di: Linze, Chen, et al.
Pubblicazione: (2026)
di: Linze, Chen, et al.
Pubblicazione: (2026)
Relevance-driven Decision Making for Safer and More Efficient Human Robot Collaboration
di: Zhang, Xiaotong, et al.
Pubblicazione: (2024)
di: Zhang, Xiaotong, et al.
Pubblicazione: (2024)
Can AI Make Energy Retrofit Decisions? An Evaluation of Large Language Models
di: Shu, Lei, et al.
Pubblicazione: (2025)
di: Shu, Lei, et al.
Pubblicazione: (2025)
Monitoring Deployed AI Systems in Health Care
di: Keyes, Timothy, et al.
Pubblicazione: (2025)
di: Keyes, Timothy, et al.
Pubblicazione: (2025)
Real-world Deployment and Evaluation of PErioperative AI CHatbot (PEACH) -- a Large Language Model Chatbot for Perioperative Medicine
di: Ke, Yu He, et al.
Pubblicazione: (2024)
di: Ke, Yu He, et al.
Pubblicazione: (2024)
Solving Context Window Overflow in AI Agents
di: Labate, Anton Bulle, et al.
Pubblicazione: (2025)
di: Labate, Anton Bulle, et al.
Pubblicazione: (2025)
AI2Agent: An End-to-End Framework for Deploying AI Projects as Autonomous Agents
di: Chen, Jiaxiang, et al.
Pubblicazione: (2025)
di: Chen, Jiaxiang, et al.
Pubblicazione: (2025)
Interactive AI Alignment: Specification, Process, and Evaluation Alignment
di: Terry, Michael, et al.
Pubblicazione: (2023)
di: Terry, Michael, et al.
Pubblicazione: (2023)
Responsible Evaluation of AI for Mental Health
di: Arnaout, Hiba, et al.
Pubblicazione: (2026)
di: Arnaout, Hiba, et al.
Pubblicazione: (2026)
CATCODER: Repository-Level Code Generation with Relevant Code and Type Context
di: Pan, Zhiyuan, et al.
Pubblicazione: (2024)
di: Pan, Zhiyuan, et al.
Pubblicazione: (2024)
XChoice: Explainable Evaluation of AI-Human Alignment in LLM-based Constrained Choice Decision Making
di: Qi, Weihong, et al.
Pubblicazione: (2026)
di: Qi, Weihong, et al.
Pubblicazione: (2026)
CRANE: Causal Relevance Analysis of Language-Specific Neurons in Multilingual Large Language Models
di: Le, Yifan, et al.
Pubblicazione: (2026)
di: Le, Yifan, et al.
Pubblicazione: (2026)
Failure-Centered Runtime Evaluation for Deployed Trilingual Public-Space Agents
di: Meng, M.
Pubblicazione: (2026)
di: Meng, M.
Pubblicazione: (2026)
The Ethics of AI in Education
di: Porayska-Pomsta, Kaska, et al.
Pubblicazione: (2024)
di: Porayska-Pomsta, Kaska, et al.
Pubblicazione: (2024)
The Deployment Gap in AI Media Detection: Platform-Aware and Visually Constrained Adversarial Evaluation
di: Budhkar, Aishwarya, et al.
Pubblicazione: (2026)
di: Budhkar, Aishwarya, et al.
Pubblicazione: (2026)
In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
di: Liu, Sheng, et al.
Pubblicazione: (2023)
di: Liu, Sheng, et al.
Pubblicazione: (2023)
An Agentic Framework for Rapid Deployment of Edge AI Solutions in Industry 5.0
di: Martinez-Gil, Jorge, et al.
Pubblicazione: (2025)
di: Martinez-Gil, Jorge, et al.
Pubblicazione: (2025)
Making AI Intelligible: Philosophical Foundations
di: Cappelen, Herman, et al.
Pubblicazione: (2024)
di: Cappelen, Herman, et al.
Pubblicazione: (2024)
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
di: Arrieta, Aitor, et al.
Pubblicazione: (2025)
di: Arrieta, Aitor, et al.
Pubblicazione: (2025)
Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs
di: Tian, Runchu, et al.
Pubblicazione: (2024)
di: Tian, Runchu, et al.
Pubblicazione: (2024)
Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
di: Garg, Kartik, et al.
Pubblicazione: (2025)
di: Garg, Kartik, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma
di: Schwartz, Reva, et al.
Pubblicazione: (2026) -
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
di: Schwartz, Reva, et al.
Pubblicazione: (2026) -
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
di: Schwartz, Reva, et al.
Pubblicazione: (2025) -
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
di: Kryshtal, Andrii
Pubblicazione: (2026) -
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
di: Griffin, Charlie, et al.
Pubblicazione: (2024)