Gespeichert in:
| Hauptverfasser: | Li, Xiaomin, Gao, Mingye, Hao, Yuexing, Li, Taoran, Wan, Guangya, Wang, Zihan, Wang, Yijun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2505.11613 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data-adaptive Safety Rules for Training Reward Models
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
Derailer-Rerailer: Adaptive Verification for Efficient and Reliable Language Model Reasoning
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
Large Language Models for Causal Discovery: Current Landscape and Future Directions
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
Self-Adaptive Cognitive Debiasing for Large Language Models in Decision-Making
von: Lyu, Yougang, et al.
Veröffentlicht: (2025)
von: Lyu, Yougang, et al.
Veröffentlicht: (2025)
CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios
von: Ouyang, Zetian, et al.
Veröffentlicht: (2024)
von: Ouyang, Zetian, et al.
Veröffentlicht: (2024)
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
von: Khoshnoodi, Mahsa, et al.
Veröffentlicht: (2024)
von: Khoshnoodi, Mahsa, et al.
Veröffentlicht: (2024)
Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making
von: Huang, Jen-tse, et al.
Veröffentlicht: (2026)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2026)
ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory
von: Ge, Zhuohan, et al.
Veröffentlicht: (2026)
von: Ge, Zhuohan, et al.
Veröffentlicht: (2026)
CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2024)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2024)
EEG-MedRAG: Enhancing EEG-based Clinical Decision-Making via Hierarchical Hypergraph Retrieval-Augmented Generation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
von: Xiao, Yunpeng, et al.
Veröffentlicht: (2025)
von: Xiao, Yunpeng, et al.
Veröffentlicht: (2025)
DynaThink: Fast or Slow? A Dynamic Decision-Making Framework for Large Language Models
von: Pan, Jiabao, et al.
Veröffentlicht: (2024)
von: Pan, Jiabao, et al.
Veröffentlicht: (2024)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
High-Fidelity Pruning for Large Language Models
von: Zhu, Yijun, et al.
Veröffentlicht: (2026)
von: Zhu, Yijun, et al.
Veröffentlicht: (2026)
CMQCIC-Bench: A Chinese Benchmark for Evaluating Large Language Models in Medical Quality Control Indicator Calculation
von: Yu, Guangya, et al.
Veröffentlicht: (2025)
von: Yu, Guangya, et al.
Veröffentlicht: (2025)
Knowledge-Augmented Large Language Model Agents for Explainable Financial Decision-Making
von: Zhang, Qingyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Qingyuan, et al.
Veröffentlicht: (2025)
MedBench: A Comprehensive, Standardized, and Reliable Benchmarking System for Evaluating Chinese Medical Large Language Models
von: Liu, Mianxin, et al.
Veröffentlicht: (2024)
von: Liu, Mianxin, et al.
Veröffentlicht: (2024)
Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making
von: M3 Team, et al.
Veröffentlicht: (2026)
von: M3 Team, et al.
Veröffentlicht: (2026)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
GAIN: A Benchmark for Goal-Aligned Decision-Making of Large Language Models under Imperfect Norms
von: Kawarada, Masayuki, et al.
Veröffentlicht: (2026)
von: Kawarada, Masayuki, et al.
Veröffentlicht: (2026)
A Large-Scale Simulation on Large Language Models for Decision-Making in Political Science
von: Yu, Chenxiao, et al.
Veröffentlicht: (2024)
von: Yu, Chenxiao, et al.
Veröffentlicht: (2024)
MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
von: Tam, Zhi Rui, et al.
Veröffentlicht: (2025)
von: Tam, Zhi Rui, et al.
Veröffentlicht: (2025)
Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
von: Gallifant, Jack, et al.
Veröffentlicht: (2024)
von: Gallifant, Jack, et al.
Veröffentlicht: (2024)
UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024)
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024)
Debt Collection Negotiations with Large Language Models: An Evaluation System and Optimizing Decision Making with Multi-Agent
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2025)
Efficient Sequential Decision Making with Large Language Models
von: Chen, Dingyang, et al.
Veröffentlicht: (2024)
von: Chen, Dingyang, et al.
Veröffentlicht: (2024)
CLIMB: A Benchmark of Clinical Bias in Large Language Models
von: Zhang, Yubo, et al.
Veröffentlicht: (2024)
von: Zhang, Yubo, et al.
Veröffentlicht: (2024)
A Survey on Large Language Model Benchmarks
von: Ni, Shiwen, et al.
Veröffentlicht: (2025)
von: Ni, Shiwen, et al.
Veröffentlicht: (2025)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
Large Language Models in the Clinic: A Comprehensive Benchmark
von: Liu, Fenglin, et al.
Veröffentlicht: (2024)
von: Liu, Fenglin, et al.
Veröffentlicht: (2024)
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
von: Alonso, Iñigo, et al.
Veröffentlicht: (2024)
von: Alonso, Iñigo, et al.
Veröffentlicht: (2024)
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
von: Daoud, Mouath Abu, et al.
Veröffentlicht: (2025)
von: Daoud, Mouath Abu, et al.
Veröffentlicht: (2025)
Enhancing Large Language Models for Clinical Decision Support by Incorporating Clinical Practice Guidelines
von: Oniani, David, et al.
Veröffentlicht: (2024)
von: Oniani, David, et al.
Veröffentlicht: (2024)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
von: Li, Manling, et al.
Veröffentlicht: (2024)
von: Li, Manling, et al.
Veröffentlicht: (2024)
Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models
von: Palnitkar, Aadi, et al.
Veröffentlicht: (2026)
von: Palnitkar, Aadi, et al.
Veröffentlicht: (2026)
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
von: Feng, Yijun
Veröffentlicht: (2025)
von: Feng, Yijun
Veröffentlicht: (2025)
Ähnliche Einträge
-
Data-adaptive Safety Rules for Training Reward Models
von: Li, Xiaomin, et al.
Veröffentlicht: (2025) -
Derailer-Rerailer: Adaptive Verification for Efficient and Reliable Language Model Reasoning
von: Wan, Guangya, et al.
Veröffentlicht: (2024) -
Large Language Models for Causal Discovery: Current Landscape and Future Directions
von: Wan, Guangya, et al.
Veröffentlicht: (2024) -
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
von: Li, Xiaomin, et al.
Veröffentlicht: (2025) -
Self-Adaptive Cognitive Debiasing for Large Language Models in Decision-Making
von: Lyu, Yougang, et al.
Veröffentlicht: (2025)