Similar Items
Modeling Clinical Concern Trajectories in Language Model Agents
by: Subaharan, Sukesh, et al.
Published: (2026)
by: Subaharan, Sukesh, et al.
Published: (2026)
humancompatible.detect: a Python Toolkit for Detecting Bias in AI Models
by: Matilla, German M., et al.
Published: (2025)
by: Matilla, German M., et al.
Published: (2025)
Tape: A Cellular Automata Benchmark for Evaluating Rule-Shift Generalization in Reinforcement Learning
by: Pan, Enze
Published: (2026)
by: Pan, Enze
Published: (2026)
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
by: Sadjoli, Nicholas, et al.
Published: (2026)
by: Sadjoli, Nicholas, et al.
Published: (2026)
Creativity in the Age of AI: Rethinking the Role of Intentional Agency
by: Pearson, James S., et al.
Published: (2026)
by: Pearson, James S., et al.
Published: (2026)
Mathematical reasoning and the computer
by: Buzzard, Kevin
Published: (2025)
by: Buzzard, Kevin
Published: (2025)
How well can a large language model explain business processes as perceived by users?
by: Fahland, Dirk, et al.
Published: (2024)
by: Fahland, Dirk, et al.
Published: (2024)
Toward a Dynamic Stackelberg Game-Theoretic Framework for Agentic AI Defense Against LLM Jailbreaking
by: Han, Zhengye, et al.
Published: (2025)
by: Han, Zhengye, et al.
Published: (2025)
BernGraph: Probabilistic Graph Neural Networks for EHR-based Medication Recommendations
by: Piao, Xihao, et al.
Published: (2024)
by: Piao, Xihao, et al.
Published: (2024)
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
by: Waggoner, Philip
Published: (2026)
by: Waggoner, Philip
Published: (2026)
The Hidden Costs of AI: A Review of Energy, E-Waste, and Inequality in Model Development
by: Winsta, Jenis
Published: (2025)
by: Winsta, Jenis
Published: (2025)
Rethinking Foundation Models for Medical Image Classification through a Benchmark Study on MedMNIST
by: Wu, Fuping, et al.
Published: (2025)
by: Wu, Fuping, et al.
Published: (2025)
Benchmarking Energy Efficiency of Large Language Models Using vLLM
by: Pronk, K., et al.
Published: (2025)
by: Pronk, K., et al.
Published: (2025)
From Language Models to Practical Self-Improving Computer Agents
by: Sheng, Alex
Published: (2024)
by: Sheng, Alex
Published: (2024)
RASP-Tuner: Retrieval-Augmented Soft Prompts for Context-Aware Black-Box Optimization in Non-Stationary Environments
by: Pan, Enze
Published: (2026)
by: Pan, Enze
Published: (2026)
Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning
by: García-Márquez, Mario, et al.
Published: (2026)
by: García-Márquez, Mario, et al.
Published: (2026)
ConSensus: Multi-Agent Collaboration for Multimodal Sensing
by: Yoon, Hyungjun, et al.
Published: (2026)
by: Yoon, Hyungjun, et al.
Published: (2026)
seqme: a Python library for evaluating biological sequence design
by: Møller-Larsen, Rasmus, et al.
Published: (2025)
by: Møller-Larsen, Rasmus, et al.
Published: (2025)
A Theoretical Analysis of Soft-Label vs Hard-Label Training in Neural Networks
by: Mandal, Saptarshi, et al.
Published: (2024)
by: Mandal, Saptarshi, et al.
Published: (2024)
Leveraging Diversity in Online Interactions
by: Osman, Nardine, et al.
Published: (2023)
by: Osman, Nardine, et al.
Published: (2023)
Adaptive Orchestration for Large-Scale Inference on Heterogeneous Accelerator Systems Balancing Cost, Performance, and Resilience
by: Biran, Yahav, et al.
Published: (2025)
by: Biran, Yahav, et al.
Published: (2025)
Script-Based Dialog Policy Planning for LLM-Powered Conversational Agents: A Basic Architecture for an "AI Therapist"
by: Wasenmüller, Robert, et al.
Published: (2024)
by: Wasenmüller, Robert, et al.
Published: (2024)
SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions
by: Xu, Weijie, et al.
Published: (2025)
by: Xu, Weijie, et al.
Published: (2025)
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
by: Lim, Soohan, et al.
Published: (2025)
by: Lim, Soohan, et al.
Published: (2025)
Fast, close, non-singular and property-preserving approximations of entropic measures
by: Horenko, Illia, et al.
Published: (2025)
by: Horenko, Illia, et al.
Published: (2025)
The PacifAIst Benchmark:Would an Artificial Intelligence Choose to Sacrifice Itself for Human Safety?
by: Herrador, Manuel
Published: (2025)
by: Herrador, Manuel
Published: (2025)
Beyond the Org Chart: AI and the Transformation of Invisible Work
by: Rosenthal, Stephanie, et al.
Published: (2026)
by: Rosenthal, Stephanie, et al.
Published: (2026)
Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
by: Capdevielle, Tomás, et al.
Published: (2025)
by: Capdevielle, Tomás, et al.
Published: (2025)
Understanding Knowledge Transferability for Transfer Learning: A Survey
by: Wang, Haohua, et al.
Published: (2025)
by: Wang, Haohua, et al.
Published: (2025)
A Taxonomy of Omnicidal Futures Involving Artificial Intelligence
by: Critch, Andrew, et al.
Published: (2025)
by: Critch, Andrew, et al.
Published: (2025)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
by: Aksoy, Sinan G., et al.
Published: (2026)
by: Aksoy, Sinan G., et al.
Published: (2026)
Abductive explanations of classifiers under constraints: Complexity and properties
by: Cooper, Martin, et al.
Published: (2024)
by: Cooper, Martin, et al.
Published: (2024)
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models
by: Jiang, Yifan, et al.
Published: (2025)
by: Jiang, Yifan, et al.
Published: (2025)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
by: Štefánik, Michal, et al.
Published: (2025)
by: Štefánik, Michal, et al.
Published: (2025)
What is the $\textit{intrinsic}$ dimension of your binary data? -- and how to compute it quickly
by: Hanika, Tom, et al.
Published: (2024)
by: Hanika, Tom, et al.
Published: (2024)
MedMNIST-NAS-Bench: A Tabular Neural Architecture Search Benchmark on MedMNIST v2
by: Wang, Wei (William)
Published: (2026)
by: Wang, Wei (William)
Published: (2026)
FedUNet: A Lightweight Additive U-Net Module for Federated Learning with Heterogeneous Models
by: Seo, Beomseok, et al.
Published: (2025)
by: Seo, Beomseok, et al.
Published: (2025)
Modelling Human Values for AI Reasoning
by: Osman, Nardine, et al.
Published: (2024)
by: Osman, Nardine, et al.
Published: (2024)
Authenticated Delegation and Authorized AI Agents
by: South, Tobin, et al.
Published: (2025)
by: South, Tobin, et al.
Published: (2025)
Challenges and Future Directions in Agentic Reverse Engineering Systems
by: Radey, Salem, et al.
Published: (2026)
by: Radey, Salem, et al.
Published: (2026)
Similar Items
-
Modeling Clinical Concern Trajectories in Language Model Agents
by: Subaharan, Sukesh, et al.
Published: (2026) -
humancompatible.detect: a Python Toolkit for Detecting Bias in AI Models
by: Matilla, German M., et al.
Published: (2025) -
Tape: A Cellular Automata Benchmark for Evaluating Rule-Shift Generalization in Reinforcement Learning
by: Pan, Enze
Published: (2026) -
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
by: Sadjoli, Nicholas, et al.
Published: (2026) -
Creativity in the Age of AI: Rethinking the Role of Intentional Agency
by: Pearson, James S., et al.
Published: (2026)