Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Oh, Gyutaek, Kim, Seoyeon, Park, Sangjoon, Kim, Byung-Hoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
di: Oh, Gyutaek, et al.
Pubblicazione: (2025)
di: Oh, Gyutaek, et al.
Pubblicazione: (2025)
Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
di: Lim, Kyung Ho, et al.
Pubblicazione: (2025)
di: Lim, Kyung Ho, et al.
Pubblicazione: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
di: Shin, Jisu, et al.
Pubblicazione: (2025)
di: Shin, Jisu, et al.
Pubblicazione: (2025)
LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory
di: Kim, Kyung-Hoon
Pubblicazione: (2025)
di: Kim, Kyung-Hoon
Pubblicazione: (2025)
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
di: Kim, Yubin, et al.
Pubblicazione: (2025)
di: Kim, Yubin, et al.
Pubblicazione: (2025)
Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs
di: Lee, Sungjae, et al.
Pubblicazione: (2025)
di: Lee, Sungjae, et al.
Pubblicazione: (2025)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
di: Cho, Jay Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Jay Hyeon, et al.
Pubblicazione: (2025)
Classroom AI: Large Language Models as Grade-Specific Teachers
di: Oh, Jio, et al.
Pubblicazione: (2026)
di: Oh, Jio, et al.
Pubblicazione: (2026)
AI Awareness
di: Li, Xiaojian, et al.
Pubblicazione: (2025)
di: Li, Xiaojian, et al.
Pubblicazione: (2025)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
di: D'addario, Andrew Maranhão Ventura
Pubblicazione: (2025)
di: D'addario, Andrew Maranhão Ventura
Pubblicazione: (2025)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
di: Subedi, Krishna
Pubblicazione: (2025)
di: Subedi, Krishna
Pubblicazione: (2025)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
di: Singh, Shrutika, et al.
Pubblicazione: (2025)
di: Singh, Shrutika, et al.
Pubblicazione: (2025)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
di: Dawson, Fiifi, et al.
Pubblicazione: (2024)
di: Dawson, Fiifi, et al.
Pubblicazione: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
di: Lee, Kyeongryul, et al.
Pubblicazione: (2025)
di: Lee, Kyeongryul, et al.
Pubblicazione: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
How Did We Get Here? Summarizing Conversation Dynamics
di: Hua, Yilun, et al.
Pubblicazione: (2024)
di: Hua, Yilun, et al.
Pubblicazione: (2024)
Responsible AI for Test Equity and Quality: The Duolingo English Test as a Case Study
di: Burstein, Jill, et al.
Pubblicazione: (2024)
di: Burstein, Jill, et al.
Pubblicazione: (2024)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
di: Shin, Jisu, et al.
Pubblicazione: (2024)
di: Shin, Jisu, et al.
Pubblicazione: (2024)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
di: Kim, Kyuhee, et al.
Pubblicazione: (2025)
di: Kim, Kyuhee, et al.
Pubblicazione: (2025)
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
di: Sil, Pritam, et al.
Pubblicazione: (2026)
di: Sil, Pritam, et al.
Pubblicazione: (2026)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
di: Rios-Sialer, Ian
Pubblicazione: (2026)
di: Rios-Sialer, Ian
Pubblicazione: (2026)
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
di: Marius, Dumitran Adrian, et al.
Pubblicazione: (2025)
di: Marius, Dumitran Adrian, et al.
Pubblicazione: (2025)
QueerGen: How LLMs Reflect Societal Norms on Gender and Sexuality in Sentence Completion Tasks
di: Sosto, Mae, et al.
Pubblicazione: (2026)
di: Sosto, Mae, et al.
Pubblicazione: (2026)
LAPIS: Language Model-Augmented Police Investigation System
di: Kim, Heedou, et al.
Pubblicazione: (2024)
di: Kim, Heedou, et al.
Pubblicazione: (2024)
AI Act and Large Language Models (LLMs): When critical issues and privacy impact require human and ethical oversight
di: Fabiano, Nicola
Pubblicazione: (2024)
di: Fabiano, Nicola
Pubblicazione: (2024)
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
di: Park, Eunkyu, et al.
Pubblicazione: (2025)
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
di: Kim, Seoyeon, et al.
Pubblicazione: (2026)
di: Kim, Seoyeon, et al.
Pubblicazione: (2026)
Teaching at Scale: Leveraging AI to Evaluate and Elevate Engineering Education
di: Chamberland, Jean-Francois, et al.
Pubblicazione: (2025)
di: Chamberland, Jean-Francois, et al.
Pubblicazione: (2025)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
di: Zhou, Lexin, et al.
Pubblicazione: (2025)
di: Zhou, Lexin, et al.
Pubblicazione: (2025)
Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
di: Ihugba, Benign John, et al.
Pubblicazione: (2025)
di: Ihugba, Benign John, et al.
Pubblicazione: (2025)
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
di: Kochmar, Ekaterina, et al.
Pubblicazione: (2025)
di: Kochmar, Ekaterina, et al.
Pubblicazione: (2025)
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
di: Shen, Hanwen, et al.
Pubblicazione: (2026)
di: Shen, Hanwen, et al.
Pubblicazione: (2026)
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use
di: Chandra, Mohit, et al.
Pubblicazione: (2024)
di: Chandra, Mohit, et al.
Pubblicazione: (2024)
MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical Education
di: Hicke, Yann, et al.
Pubblicazione: (2025)
di: Hicke, Yann, et al.
Pubblicazione: (2025)
The ADAIO System at the BEA-2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues
di: Adigwe, Adaeze, et al.
Pubblicazione: (2023)
di: Adigwe, Adaeze, et al.
Pubblicazione: (2023)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
di: Ko, Changgeon, et al.
Pubblicazione: (2024)
di: Ko, Changgeon, et al.
Pubblicazione: (2024)
Punctuated Equilibria in Artificial Intelligence: The Institutional Scaling Law and the Speciation of Sovereign AI
di: Baciak, Mark, et al.
Pubblicazione: (2026)
di: Baciak, Mark, et al.
Pubblicazione: (2026)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
di: Mavi, John, et al.
Pubblicazione: (2025)
di: Mavi, John, et al.
Pubblicazione: (2025)
Mapping the Methodological Space of Classroom Interaction Research: Scale, Duration, and Modality in an Age of AI
di: Demszky, Dorottya, et al.
Pubblicazione: (2026)
di: Demszky, Dorottya, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
di: Oh, Gyutaek, et al.
Pubblicazione: (2025) -
Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
di: Lim, Kyung Ho, et al.
Pubblicazione: (2025) -
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
di: Shin, Jisu, et al.
Pubblicazione: (2025) -
LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory
di: Kim, Kyung-Hoon
Pubblicazione: (2025) -
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
di: Kim, Yubin, et al.
Pubblicazione: (2025)