MIRA: A Bilingual Benchmark for Medical Information Response Audit
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Mengyu, Yang, Qiaoxin, Wang, Qianqian, Dai, Xiwei, Wu, Weiyi, Gao, Chongyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AuditWen:An Open-Source Large Language Model for Audit
von: Huang, Jiajia, et al.
Veröffentlicht: (2024)
von: Huang, Jiajia, et al.
Veröffentlicht: (2024)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
Right to be Forgotten in the Era of Large Language Models: Implications, Challenges, and Solutions
von: Zhang, Dawen, et al.
Veröffentlicht: (2023)
von: Zhang, Dawen, et al.
Veröffentlicht: (2023)
PRISM: A Methodology for Auditing Biases in Large Language Models
von: Azzopardi, Leif, et al.
Veröffentlicht: (2024)
von: Azzopardi, Leif, et al.
Veröffentlicht: (2024)
AuditGPT: Auditing Smart Contracts with ChatGPT
von: Xia, Shihao, et al.
Veröffentlicht: (2024)
von: Xia, Shihao, et al.
Veröffentlicht: (2024)
From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship
von: Xu, Yue, et al.
Veröffentlicht: (2025)
von: Xu, Yue, et al.
Veröffentlicht: (2025)
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
von: Gringras, David, et al.
Veröffentlicht: (2026)
von: Gringras, David, et al.
Veröffentlicht: (2026)
The Algorithmic Caricature: Auditing LLM-Generated Political Discourse Across Crisis Events
von: Gunjan, et al.
Veröffentlicht: (2026)
von: Gunjan, et al.
Veröffentlicht: (2026)
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2025)
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2025)
Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios
von: Xu, Shaochen, et al.
Veröffentlicht: (2024)
von: Xu, Shaochen, et al.
Veröffentlicht: (2024)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review
von: Weng, Yixuan, et al.
Veröffentlicht: (2026)
von: Weng, Yixuan, et al.
Veröffentlicht: (2026)
EulerESG: Automating ESG Disclosure Analysis with LLMs
von: Ding, Yi, et al.
Veröffentlicht: (2025)
von: Ding, Yi, et al.
Veröffentlicht: (2025)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
von: Yang, Shujian, et al.
Veröffentlicht: (2025)
von: Yang, Shujian, et al.
Veröffentlicht: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
von: Shi, Yuzhen, et al.
Veröffentlicht: (2026)
von: Shi, Yuzhen, et al.
Veröffentlicht: (2026)
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
von: Kim, Yubin, et al.
Veröffentlicht: (2025)
von: Kim, Yubin, et al.
Veröffentlicht: (2025)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
von: Wang, Ze, et al.
Veröffentlicht: (2024)
von: Wang, Ze, et al.
Veröffentlicht: (2024)
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
von: Guo, Ruohao, et al.
Veröffentlicht: (2025)
von: Guo, Ruohao, et al.
Veröffentlicht: (2025)
When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models
von: Venkata, Pruthvinath Jeripity
Veröffentlicht: (2026)
von: Venkata, Pruthvinath Jeripity
Veröffentlicht: (2026)
Auditing Gender Presentation Differences in Text-to-Image Models
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
von: Wang, Lichao, et al.
Veröffentlicht: (2026)
von: Wang, Lichao, et al.
Veröffentlicht: (2026)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
Human-Centred LLM Privacy Audits: Findings and Frictions
von: Staufer, Dimitri, et al.
Veröffentlicht: (2026)
von: Staufer, Dimitri, et al.
Veröffentlicht: (2026)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
von: Li, Yuchong, et al.
Veröffentlicht: (2025)
von: Li, Yuchong, et al.
Veröffentlicht: (2025)
From Text to Multimodality: Exploring the Evolution and Impact of Large Language Models in Medical Practice
von: Niu, Qian, et al.
Veröffentlicht: (2024)
von: Niu, Qian, et al.
Veröffentlicht: (2024)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
von: Wang, Tianyu, et al.
Veröffentlicht: (2026)
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
von: Wang, Xing, et al.
Veröffentlicht: (2025)
von: Wang, Xing, et al.
Veröffentlicht: (2025)
Social Bias in Popular Question-Answering Benchmarks
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
von: Lim, Kyung Ho, et al.
Veröffentlicht: (2025)
von: Lim, Kyung Ho, et al.
Veröffentlicht: (2025)
Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias
von: Wu, Sirui, et al.
Veröffentlicht: (2025)
von: Wu, Sirui, et al.
Veröffentlicht: (2025)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
von: Wang, Huandong, et al.
Veröffentlicht: (2025)
von: Wang, Huandong, et al.
Veröffentlicht: (2025)
Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
von: Feng, Duanyu, et al.
Veröffentlicht: (2023)
von: Feng, Duanyu, et al.
Veröffentlicht: (2023)
Can Large Language Models Simulate Human Responses? A Case Study of Stated Preference Experiments in the Context of Heating-related Choices
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
AccessEval: Benchmarking Disability Bias in Large Language Models
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
DarkBench: Benchmarking Dark Patterns in Large Language Models
von: Kran, Esben, et al.
Veröffentlicht: (2025)
von: Kran, Esben, et al.
Veröffentlicht: (2025)
The Responsible Development of Automated Student Feedback with Generative AI
von: Lindsay, Euan D, et al.
Veröffentlicht: (2023)
von: Lindsay, Euan D, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
AuditWen:An Open-Source Large Language Model for Audit
von: Huang, Jiajia, et al.
Veröffentlicht: (2024) -
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
von: Qiu, Peiran, et al.
Veröffentlicht: (2025) -
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
von: Gao, Zihan, et al.
Veröffentlicht: (2025) -
Right to be Forgotten in the Era of Large Language Models: Implications, Challenges, and Solutions
von: Zhang, Dawen, et al.
Veröffentlicht: (2023) -
PRISM: A Methodology for Auditing Biases in Large Language Models
von: Azzopardi, Leif, et al.
Veröffentlicht: (2024)