Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liao, Yusheng, Meng, Yutong, Wang, Yuhao, Liu, Hongcheng, Wang, Yanfeng, Wang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
MING-MOE: Enhancing Medical Multi-Task Learning in Large Language Models with Sparse Mixture of Low-Rank Adapter Experts
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
TAIA: Large Language Models are Out-of-Distribution Data Learners
von: Jiang, Shuyang, et al.
Veröffentlicht: (2024)
von: Jiang, Shuyang, et al.
Veröffentlicht: (2024)
Leveraging Diverse Modeling Contexts with Collaborating Learning for Neural Machine Translation
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2025)
von: Chen, Zhe, et al.
Veröffentlicht: (2025)
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications
von: Chen, Zhe, et al.
Veröffentlicht: (2025)
von: Chen, Zhe, et al.
Veröffentlicht: (2025)
Decoding Linguistic Representations of Human Brain
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024)
Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
von: Wang, Pingjie, et al.
Veröffentlicht: (2025)
von: Wang, Pingjie, et al.
Veröffentlicht: (2025)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models
von: Guo, YiQiu, et al.
Veröffentlicht: (2025)
von: Guo, YiQiu, et al.
Veröffentlicht: (2025)
AgentEHR: Advancing Autonomous Clinical Decision-Making via Retrospective Summarization
von: Liao, Yusheng, et al.
Veröffentlicht: (2026)
von: Liao, Yusheng, et al.
Veröffentlicht: (2026)
MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
von: Liu, Hongcheng, et al.
Veröffentlicht: (2026)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2026)
DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
von: Li, Yiqi, et al.
Veröffentlicht: (2025)
von: Li, Yiqi, et al.
Veröffentlicht: (2025)
Overthinking Reduction with Decoupled Rewards and Curriculum Data Scheduling
von: Jiang, Shuyang, et al.
Veröffentlicht: (2025)
von: Jiang, Shuyang, et al.
Veröffentlicht: (2025)
Miner:Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models
von: Jiang, Shuyang, et al.
Veröffentlicht: (2026)
von: Jiang, Shuyang, et al.
Veröffentlicht: (2026)
H2HTalk: Evaluating Large Language Models as Emotional Companion
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
MedS$^3$: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process Supervision
von: Jiang, Shuyang, et al.
Veröffentlicht: (2025)
von: Jiang, Shuyang, et al.
Veröffentlicht: (2025)
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
von: Zhang, Xiechi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiechi, et al.
Veröffentlicht: (2025)
LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models
von: Zhao, Zihan, et al.
Veröffentlicht: (2023)
von: Zhao, Zihan, et al.
Veröffentlicht: (2023)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
Towards Evaluating and Building Versatile Large Language Models for Medicine
von: Wu, Chaoyi, et al.
Veröffentlicht: (2024)
von: Wu, Chaoyi, et al.
Veröffentlicht: (2024)
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
von: Feng, Sheng, et al.
Veröffentlicht: (2024)
von: Feng, Sheng, et al.
Veröffentlicht: (2024)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
DictLLM: Harnessing Key-Value Data Structures with Large Language Models for Enhanced Medical Diagnostics
von: Guo, YiQiu, et al.
Veröffentlicht: (2024)
von: Guo, YiQiu, et al.
Veröffentlicht: (2024)
CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios
von: Ouyang, Zetian, et al.
Veröffentlicht: (2024)
von: Ouyang, Zetian, et al.
Veröffentlicht: (2024)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026)
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026)
EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
von: Liao, Yusheng, et al.
Veröffentlicht: (2025)
von: Liao, Yusheng, et al.
Veröffentlicht: (2025)
Hallucination Detection and Evaluation of Large Language Model
von: Zhang, Chenggong, et al.
Veröffentlicht: (2025)
von: Zhang, Chenggong, et al.
Veröffentlicht: (2025)
TasTe: Teaching Large Language Models to Translate through Self-Reflection
von: Wang, Yutong, et al.
Veröffentlicht: (2024)
von: Wang, Yutong, et al.
Veröffentlicht: (2024)
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
von: Li, Ruosen, et al.
Veröffentlicht: (2024)
von: Li, Ruosen, et al.
Veröffentlicht: (2024)
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
von: Song, Yusheng, et al.
Veröffentlicht: (2025)
von: Song, Yusheng, et al.
Veröffentlicht: (2025)
M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
von: Chen, Zhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024) -
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024) -
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
von: Wang, Yuhao, et al.
Veröffentlicht: (2024) -
MING-MOE: Enhancing Medical Multi-Task Learning in Large Language Models with Sparse Mixture of Low-Rank Adapter Experts
von: Liao, Yusheng, et al.
Veröffentlicht: (2024) -
TAIA: Large Language Models are Out-of-Distribution Data Learners
von: Jiang, Shuyang, et al.
Veröffentlicht: (2024)