When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Karen Jia-Hui, Balloccu, Simone, Dusek, Ondrej, Reiter, Ehud |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Machine Can't Replace the Human Heart
por: Lin, Baihan
Publicado: (2024)
por: Lin, Baihan
Publicado: (2024)
From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications
por: Steigerwald, Philipp, et al.
Publicado: (2026)
por: Steigerwald, Philipp, et al.
Publicado: (2026)
You Can't Get There From Here: Redefining Information Science to address our sociotechnical futures
por: Humr, Scott, et al.
Publicado: (2025)
por: Humr, Scott, et al.
Publicado: (2025)
Facilitating Holistic Evaluations with LLMs: Insights from Scenario-Based Experiments
por: Ishida, Toru, et al.
Publicado: (2024)
por: Ishida, Toru, et al.
Publicado: (2024)
Understanding Help-Seeking Behavior of Students Using LLMs vs. Web Search for Writing SQL Queries
por: Kumar, Harsh, et al.
Publicado: (2024)
por: Kumar, Harsh, et al.
Publicado: (2024)
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
por: Choong, Yee-Yin, et al.
Publicado: (2026)
por: Choong, Yee-Yin, et al.
Publicado: (2026)
Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
por: El, Batu, et al.
Publicado: (2025)
por: El, Batu, et al.
Publicado: (2025)
MediTools -- Medical Education Powered by LLMs
por: Alshatnawi, Amr, et al.
Publicado: (2025)
por: Alshatnawi, Amr, et al.
Publicado: (2025)
A Comparative Study of Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics
por: Liu, Suqing, et al.
Publicado: (2025)
por: Liu, Suqing, et al.
Publicado: (2025)
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
por: Chen, Baiyu, et al.
Publicado: (2025)
por: Chen, Baiyu, et al.
Publicado: (2025)
Hope, Aspirations, and the Impact of LLMs on Female Programming Learners in Afghanistan
por: Behmanush, Hamayoon, et al.
Publicado: (2025)
por: Behmanush, Hamayoon, et al.
Publicado: (2025)
Between Myths and Metaphors: Rethinking LLMs for SRH in Conservative Contexts
por: Humayun, Ameemah, et al.
Publicado: (2025)
por: Humayun, Ameemah, et al.
Publicado: (2025)
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
por: Najera, Aisha, et al.
Publicado: (2026)
por: Najera, Aisha, et al.
Publicado: (2026)
Documenting Deployment with Fabric: A Repository of Real-World AI Governance
por: Jorgensen, Mackenzie, et al.
Publicado: (2025)
por: Jorgensen, Mackenzie, et al.
Publicado: (2025)
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
por: Garcia, Adriana Alvarado, et al.
Publicado: (2026)
por: Garcia, Adriana Alvarado, et al.
Publicado: (2026)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
por: Balloccu, Simone, et al.
Publicado: (2024)
por: Balloccu, Simone, et al.
Publicado: (2024)
Disability Across Cultures: A Human-Centered Audit of Ableism in Western and Indic LLMs
por: Phutane, Mahika, et al.
Publicado: (2025)
por: Phutane, Mahika, et al.
Publicado: (2025)
Harnessing LLMs for Automated Video Content Analysis: An Exploratory Workflow of Short Videos on Depression
por: Liu, Jiaying Lizzy, et al.
Publicado: (2024)
por: Liu, Jiaying Lizzy, et al.
Publicado: (2024)
A Conditional Companion: Lived Experiences of People with Mental Health Disorders Using LLMs
por: Purohit, Aditya Kumar, et al.
Publicado: (2026)
por: Purohit, Aditya Kumar, et al.
Publicado: (2026)
ROBOPSY PL[AI]: Using Role-Play to Investigate how LLMs Present Collective Memory
por: Jahrmann, Margarete, et al.
Publicado: (2025)
por: Jahrmann, Margarete, et al.
Publicado: (2025)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
por: Pang, Rock Yuren, et al.
Publicado: (2025)
por: Pang, Rock Yuren, et al.
Publicado: (2025)
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
por: Kumar, Harsh, et al.
Publicado: (2025)
por: Kumar, Harsh, et al.
Publicado: (2025)
Can LLMs Help Improve Analogical Reasoning For Strategic Decisions? Experimental Evidence from Humans and GPT-4
por: Puranam, Phanish, et al.
Publicado: (2025)
por: Puranam, Phanish, et al.
Publicado: (2025)
Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context
por: Jia, Jingru, et al.
Publicado: (2024)
por: Jia, Jingru, et al.
Publicado: (2024)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
por: Zheng, Mingqian, et al.
Publicado: (2023)
por: Zheng, Mingqian, et al.
Publicado: (2023)
Can AI be Auditable?
por: Verma, Himanshu, et al.
Publicado: (2025)
por: Verma, Himanshu, et al.
Publicado: (2025)
Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives
por: Drożdż, Karolina, et al.
Publicado: (2025)
por: Drożdż, Karolina, et al.
Publicado: (2025)
OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs
por: Li, Jiahao Nick, et al.
Publicado: (2024)
por: Li, Jiahao Nick, et al.
Publicado: (2024)
Evaluation of Human-Understandability of Global Model Explanations using Decision Tree
por: Sivaprasad, Adarsa, et al.
Publicado: (2023)
por: Sivaprasad, Adarsa, et al.
Publicado: (2023)
Wearable Device-Based Real-Time Monitoring of Physiological Signals: Evaluating Cognitive Load Across Different Tasks
por: He, Ling, et al.
Publicado: (2024)
por: He, Ling, et al.
Publicado: (2024)
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
por: Sharma, Mrinank, et al.
Publicado: (2026)
por: Sharma, Mrinank, et al.
Publicado: (2026)
REALM: A Dataset of Real-World LLM Use Cases
por: Cheng, Jingwen, et al.
Publicado: (2025)
por: Cheng, Jingwen, et al.
Publicado: (2025)
When Autonomy Breaks: The Hidden Existential Risk of AI
por: Krook, Joshua
Publicado: (2025)
por: Krook, Joshua
Publicado: (2025)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
por: Subedi, Krishna
Publicado: (2025)
por: Subedi, Krishna
Publicado: (2025)
A Comprehensive Survey of Bias in LLMs: Current Landscape and Future Directions
por: Ranjan, Rajesh, et al.
Publicado: (2024)
por: Ranjan, Rajesh, et al.
Publicado: (2024)
Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
por: Ashkinaze, Joshua, et al.
Publicado: (2024)
por: Ashkinaze, Joshua, et al.
Publicado: (2024)
When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines
por: Pendse, Sachin R., et al.
Publicado: (2025)
por: Pendse, Sachin R., et al.
Publicado: (2025)
When combinations of humans and AI are useful: A systematic review and meta-analysis
por: Vaccaro, Michelle, et al.
Publicado: (2024)
por: Vaccaro, Michelle, et al.
Publicado: (2024)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
por: Badawi, Abeer, et al.
Publicado: (2025)
por: Badawi, Abeer, et al.
Publicado: (2025)
When Visibility Outpaces Verification: Delayed Verification and Narrative Lock-in in Agentic AI Discourse
por: Shi, Hanjing, et al.
Publicado: (2026)
por: Shi, Hanjing, et al.
Publicado: (2026)
Ejemplares similares
-
The Machine Can't Replace the Human Heart
por: Lin, Baihan
Publicado: (2024) -
From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications
por: Steigerwald, Philipp, et al.
Publicado: (2026) -
You Can't Get There From Here: Redefining Information Science to address our sociotechnical futures
por: Humr, Scott, et al.
Publicado: (2025) -
Facilitating Holistic Evaluations with LLMs: Insights from Scenario-Based Experiments
por: Ishida, Toru, et al.
Publicado: (2024) -
Understanding Help-Seeking Behavior of Students Using LLMs vs. Web Search for Writing SQL Queries
por: Kumar, Harsh, et al.
Publicado: (2024)