Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Oh, Gyutaek, Kim, Seoyeon, Park, Sangjoon, Kim, Byung-Hoon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
por: Oh, Gyutaek, et al.
Publicado: (2025)
por: Oh, Gyutaek, et al.
Publicado: (2025)
Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
por: Lim, Kyung Ho, et al.
Publicado: (2025)
por: Lim, Kyung Ho, et al.
Publicado: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
por: Shin, Jisu, et al.
Publicado: (2025)
por: Shin, Jisu, et al.
Publicado: (2025)
LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory
por: Kim, Kyung-Hoon
Publicado: (2025)
por: Kim, Kyung-Hoon
Publicado: (2025)
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
por: Kim, Yubin, et al.
Publicado: (2025)
por: Kim, Yubin, et al.
Publicado: (2025)
Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs
por: Lee, Sungjae, et al.
Publicado: (2025)
por: Lee, Sungjae, et al.
Publicado: (2025)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
por: Cho, Jay Hyeon, et al.
Publicado: (2025)
por: Cho, Jay Hyeon, et al.
Publicado: (2025)
Classroom AI: Large Language Models as Grade-Specific Teachers
por: Oh, Jio, et al.
Publicado: (2026)
por: Oh, Jio, et al.
Publicado: (2026)
AI Awareness
por: Li, Xiaojian, et al.
Publicado: (2025)
por: Li, Xiaojian, et al.
Publicado: (2025)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
por: D'addario, Andrew Maranhão Ventura
Publicado: (2025)
por: D'addario, Andrew Maranhão Ventura
Publicado: (2025)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
por: Subedi, Krishna
Publicado: (2025)
por: Subedi, Krishna
Publicado: (2025)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
por: Singh, Shrutika, et al.
Publicado: (2025)
por: Singh, Shrutika, et al.
Publicado: (2025)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
por: Dawson, Fiifi, et al.
Publicado: (2024)
por: Dawson, Fiifi, et al.
Publicado: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
por: Kim, Jiseon, et al.
Publicado: (2025)
por: Kim, Jiseon, et al.
Publicado: (2025)
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
por: Lee, Kyeongryul, et al.
Publicado: (2025)
por: Lee, Kyeongryul, et al.
Publicado: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
por: Kabir, Mohsinul, et al.
Publicado: (2026)
por: Kabir, Mohsinul, et al.
Publicado: (2026)
How Did We Get Here? Summarizing Conversation Dynamics
por: Hua, Yilun, et al.
Publicado: (2024)
por: Hua, Yilun, et al.
Publicado: (2024)
Responsible AI for Test Equity and Quality: The Duolingo English Test as a Case Study
por: Burstein, Jill, et al.
Publicado: (2024)
por: Burstein, Jill, et al.
Publicado: (2024)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
por: Shin, Jisu, et al.
Publicado: (2024)
por: Shin, Jisu, et al.
Publicado: (2024)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
por: Kim, Kyuhee, et al.
Publicado: (2025)
por: Kim, Kyuhee, et al.
Publicado: (2025)
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
por: Sil, Pritam, et al.
Publicado: (2026)
por: Sil, Pritam, et al.
Publicado: (2026)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
por: Rios-Sialer, Ian
Publicado: (2026)
por: Rios-Sialer, Ian
Publicado: (2026)
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
por: Marius, Dumitran Adrian, et al.
Publicado: (2025)
por: Marius, Dumitran Adrian, et al.
Publicado: (2025)
QueerGen: How LLMs Reflect Societal Norms on Gender and Sexuality in Sentence Completion Tasks
por: Sosto, Mae, et al.
Publicado: (2026)
por: Sosto, Mae, et al.
Publicado: (2026)
LAPIS: Language Model-Augmented Police Investigation System
por: Kim, Heedou, et al.
Publicado: (2024)
por: Kim, Heedou, et al.
Publicado: (2024)
AI Act and Large Language Models (LLMs): When critical issues and privacy impact require human and ethical oversight
por: Fabiano, Nicola
Publicado: (2024)
por: Fabiano, Nicola
Publicado: (2024)
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
por: Park, Eunkyu, et al.
Publicado: (2025)
por: Park, Eunkyu, et al.
Publicado: (2025)
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
por: Kim, Seoyeon, et al.
Publicado: (2026)
por: Kim, Seoyeon, et al.
Publicado: (2026)
Teaching at Scale: Leveraging AI to Evaluate and Elevate Engineering Education
por: Chamberland, Jean-Francois, et al.
Publicado: (2025)
por: Chamberland, Jean-Francois, et al.
Publicado: (2025)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
por: Zhou, Lexin, et al.
Publicado: (2025)
por: Zhou, Lexin, et al.
Publicado: (2025)
Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
por: Ihugba, Benign John, et al.
Publicado: (2025)
por: Ihugba, Benign John, et al.
Publicado: (2025)
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
por: Kochmar, Ekaterina, et al.
Publicado: (2025)
por: Kochmar, Ekaterina, et al.
Publicado: (2025)
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
por: Shen, Hanwen, et al.
Publicado: (2026)
por: Shen, Hanwen, et al.
Publicado: (2026)
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use
por: Chandra, Mohit, et al.
Publicado: (2024)
por: Chandra, Mohit, et al.
Publicado: (2024)
MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical Education
por: Hicke, Yann, et al.
Publicado: (2025)
por: Hicke, Yann, et al.
Publicado: (2025)
The ADAIO System at the BEA-2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues
por: Adigwe, Adaeze, et al.
Publicado: (2023)
por: Adigwe, Adaeze, et al.
Publicado: (2023)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
por: Ko, Changgeon, et al.
Publicado: (2024)
por: Ko, Changgeon, et al.
Publicado: (2024)
Punctuated Equilibria in Artificial Intelligence: The Institutional Scaling Law and the Speciation of Sovereign AI
por: Baciak, Mark, et al.
Publicado: (2026)
por: Baciak, Mark, et al.
Publicado: (2026)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
por: Mavi, John, et al.
Publicado: (2025)
por: Mavi, John, et al.
Publicado: (2025)
Mapping the Methodological Space of Classroom Interaction Research: Scale, Duration, and Modality in an Age of AI
por: Demszky, Dorottya, et al.
Publicado: (2026)
por: Demszky, Dorottya, et al.
Publicado: (2026)
Ejemplares similares
-
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
por: Oh, Gyutaek, et al.
Publicado: (2025) -
Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
por: Lim, Kyung Ho, et al.
Publicado: (2025) -
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
por: Shin, Jisu, et al.
Publicado: (2025) -
LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory
por: Kim, Kyung-Hoon
Publicado: (2025) -
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
por: Kim, Yubin, et al.
Publicado: (2025)