M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Jeong, Jihoon |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Model Medicine: A Clinical Framework for Understanding, Diagnosing, and Treating AI Models
par: Jeong, Jihoon
Publié: (2026)
par: Jeong, Jihoon
Publié: (2026)
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
par: Sargeant, Holli, et autres
Publié: (2024)
par: Sargeant, Holli, et autres
Publié: (2024)
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
par: Imperial, Joseph Marvin, et autres
Publié: (2025)
par: Imperial, Joseph Marvin, et autres
Publié: (2025)
ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
par: Heakl, Ahmed, et autres
Publié: (2024)
par: Heakl, Ahmed, et autres
Publié: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
par: Goethals, Sofie, et autres
Publié: (2026)
par: Goethals, Sofie, et autres
Publié: (2026)
DAIC-WOZ: On the Validity of Using the Therapist's prompts in Automatic Depression Detection from Clinical Interviews
par: Burdisso, Sergio, et autres
Publié: (2024)
par: Burdisso, Sergio, et autres
Publié: (2024)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
par: Han, Pengrui, et autres
Publié: (2025)
par: Han, Pengrui, et autres
Publié: (2025)
Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism
par: Andric, Sandro
Publié: (2025)
par: Andric, Sandro
Publié: (2025)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
par: Bouchard, Dylan, et autres
Publié: (2025)
par: Bouchard, Dylan, et autres
Publié: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
par: AlDahoul, Nouar, et autres
Publié: (2025)
par: AlDahoul, Nouar, et autres
Publié: (2025)
Large Language Models for Automating Clinical Data Standardization: HL7 FHIR Use Case
par: Riquelme, Alvaro, et autres
Publié: (2025)
par: Riquelme, Alvaro, et autres
Publié: (2025)
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
par: Heddaya, Mourad, et autres
Publié: (2024)
par: Heddaya, Mourad, et autres
Publié: (2024)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
par: Hu, Tiancheng, et autres
Publié: (2025)
par: Hu, Tiancheng, et autres
Publié: (2025)
MTI: A Behavior-Based Temperament Profiling System for AI Agents
par: Jeong, Jihoon
Publié: (2026)
par: Jeong, Jihoon
Publié: (2026)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
par: van der Weij, Teun, et autres
Publié: (2024)
par: van der Weij, Teun, et autres
Publié: (2024)
IoT-Based Preventive Mental Health Using Knowledge Graphs and Standards for Better Well-Being
par: Gyrard, Amelie, et autres
Publié: (2024)
par: Gyrard, Amelie, et autres
Publié: (2024)
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
par: Zhang, Yue, et autres
Publié: (2023)
par: Zhang, Yue, et autres
Publié: (2023)
EigenBench: A Comparative Behavioral Measure of Value Alignment
par: Chang, Jonathn, et autres
Publié: (2025)
par: Chang, Jonathn, et autres
Publié: (2025)
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
par: Era, Jalisha Jashim, et autres
Publié: (2025)
par: Era, Jalisha Jashim, et autres
Publié: (2025)
DualAlign: Generating Clinically Grounded Synthetic Data
par: Li, Rumeng, et autres
Publié: (2025)
par: Li, Rumeng, et autres
Publié: (2025)
From Data to Behavior: Predicting Unintended Model Behaviors Before Training
par: Wang, Mengru, et autres
Publié: (2026)
par: Wang, Mengru, et autres
Publié: (2026)
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
par: Schroeder, Daniel Thilo, et autres
Publié: (2025)
par: Schroeder, Daniel Thilo, et autres
Publié: (2025)
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
par: Chen, Ruxiao, et autres
Publié: (2025)
par: Chen, Ruxiao, et autres
Publié: (2025)
AI-AI Bias: large language models favor communications generated by large language models
par: Laurito, Walter, et autres
Publié: (2024)
par: Laurito, Walter, et autres
Publié: (2024)
Transparent AI: The Case for Interpretability and Explainability
par: Ramachandram, Dhanesh, et autres
Publié: (2025)
par: Ramachandram, Dhanesh, et autres
Publié: (2025)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
par: Xiao, Boyu, et autres
Publié: (2026)
par: Xiao, Boyu, et autres
Publié: (2026)
The Case for ESM3 as a General-Purpose AI Model with Systemic Risk Under the EU AI Act
par: Qureshi, Taro, et autres
Publié: (2026)
par: Qureshi, Taro, et autres
Publié: (2026)
Questionnaire Responses Do not Capture the Safety of AI Agents
par: Hellrigel-Holderbaum, Max, et autres
Publié: (2026)
par: Hellrigel-Holderbaum, Max, et autres
Publié: (2026)
Understanding and Mitigating Risks of Generative AI in Financial Services
par: Gehrmann, Sebastian, et autres
Publié: (2025)
par: Gehrmann, Sebastian, et autres
Publié: (2025)
Managing extreme AI risks amid rapid progress
par: Bengio, Yoshua, et autres
Publié: (2023)
par: Bengio, Yoshua, et autres
Publié: (2023)
Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
par: Tang, Xiangru, et autres
Publié: (2024)
par: Tang, Xiangru, et autres
Publié: (2024)
Know Thyself? On the Incapability and Implications of AI Self-Recognition
par: Bai, Xiaoyan, et autres
Publié: (2025)
par: Bai, Xiaoyan, et autres
Publié: (2025)
Transfer Learning for the Prediction of Entity Modifiers in Clinical Text: Application to Opioid Use Disorder Case Detection
par: Almudaifer, Abdullateef I., et autres
Publié: (2024)
par: Almudaifer, Abdullateef I., et autres
Publié: (2024)
Case Studies of AI Policy Development in Africa
par: Diallo, Kadijatou, et autres
Publié: (2024)
par: Diallo, Kadijatou, et autres
Publié: (2024)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
par: Ren, Richard, et autres
Publié: (2025)
par: Ren, Richard, et autres
Publié: (2025)
Impacts of Racial Bias in Historical Training Data for News AI
par: Bhargava, Rahul, et autres
Publié: (2025)
par: Bhargava, Rahul, et autres
Publié: (2025)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
par: Shojaei, Mostafa Faghih, et autres
Publié: (2025)
par: Shojaei, Mostafa Faghih, et autres
Publié: (2025)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
par: Schoenegger, Philipp, et autres
Publié: (2024)
par: Schoenegger, Philipp, et autres
Publié: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
par: Ren, Richard, et autres
Publié: (2024)
par: Ren, Richard, et autres
Publié: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
par: Yang, Chao, et autres
Publié: (2024)
par: Yang, Chao, et autres
Publié: (2024)
Documents similaires
-
Model Medicine: A Clinical Framework for Understanding, Diagnosing, and Treating AI Models
par: Jeong, Jihoon
Publié: (2026) -
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
par: Sargeant, Holli, et autres
Publié: (2024) -
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
par: Imperial, Joseph Marvin, et autres
Publié: (2025) -
ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models
par: Heakl, Ahmed, et autres
Publié: (2024) -
Prompt-Counterfactual Explanations for Generative AI System Behavior
par: Goethals, Sofie, et autres
Publié: (2026)