Reliable Decision Support with LLMs: A Framework for Evaluating Consistency in Binary Text Classification Applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Megahed, Fadel M., Chen, Ying-Ju, Jones-Farmer, L. Allision, Lee, Younghwa, Wang, Jiawei Brooke, Zwetsloot, Inez M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Introducing ChatSQC: Enhancing Statistical Quality Control with Augmented AI
von: Megahed, Fadel M., et al.
Veröffentlicht: (2023)
von: Megahed, Fadel M., et al.
Veröffentlicht: (2023)
Adapting OpenAI's CLIP Model for Few-Shot Image Inspection in Manufacturing Quality Control: An Expository Case Study with Multiple Application Examples
von: Megahed, Fadel M., et al.
Veröffentlicht: (2025)
von: Megahed, Fadel M., et al.
Veröffentlicht: (2025)
ChatISA: A Prompt-Engineered, In-House Multi-Modal Generative AI Chatbot for Information Systems Education
von: Megahed, Fadel M., et al.
Veröffentlicht: (2024)
von: Megahed, Fadel M., et al.
Veröffentlicht: (2024)
Social Network Datasets on Reddit Financial Discussion
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
von: Wang, Zezhong, et al.
Veröffentlicht: (2024)
Domain-Informed Negative Sampling Strategies for Dynamic Graph Embedding in Meme Stock-Related Social Networks
von: Hui, Yunming, et al.
Veröffentlicht: (2024)
von: Hui, Yunming, et al.
Veröffentlicht: (2024)
Non-Progressive Influence Maximization in Dynamic Social Networks
von: Hui, Yunming, et al.
Veröffentlicht: (2024)
von: Hui, Yunming, et al.
Veröffentlicht: (2024)
Decision-Oriented Text Evaluation
von: Huang, Yu-Shiang, et al.
Veröffentlicht: (2025)
von: Huang, Yu-Shiang, et al.
Veröffentlicht: (2025)
Benchmarking Self-Supervised Models for Cardiac Ultrasound View Classification
von: Megahed, Youssef, et al.
Veröffentlicht: (2026)
von: Megahed, Youssef, et al.
Veröffentlicht: (2026)
Area under the ROC Curve has the Most Consistent Evaluation for Binary Classification
von: Li, Jing
Veröffentlicht: (2024)
von: Li, Jing
Veröffentlicht: (2024)
Balancing Benefits and Risks of AI Adoption in Nursing Practice in Saudi Arabia
von: Ateya Megahed Ibrahim
Veröffentlicht: (2025)
von: Ateya Megahed Ibrahim
Veröffentlicht: (2025)
The Adversarial Consistency of Surrogate Risks for Binary Classification
von: Frank, Natalie, et al.
Veröffentlicht: (2023)
von: Frank, Natalie, et al.
Veröffentlicht: (2023)
Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification
von: Peng, Le, et al.
Veröffentlicht: (2025)
von: Peng, Le, et al.
Veröffentlicht: (2025)
A Systematic Evaluation of Imbalance Handling Methods in Biomedical Binary Classification
von: Chen, Jiandong, et al.
Veröffentlicht: (2026)
von: Chen, Jiandong, et al.
Veröffentlicht: (2026)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
von: Lee, Sangyub, et al.
Veröffentlicht: (2026)
von: Lee, Sangyub, et al.
Veröffentlicht: (2026)
Context-Aware Hospitalization Forecasting Evaluations for Decision Support using LLMs
von: Makkuni, Rhea, et al.
Veröffentlicht: (2026)
von: Makkuni, Rhea, et al.
Veröffentlicht: (2026)
MRD‐GAN: Multi‐representation discrimination GAN for enhancing the diversity of the generated data
von: Mohammed Megahed, et al.
Veröffentlicht: (2024)
von: Mohammed Megahed, et al.
Veröffentlicht: (2024)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
von: Wang, Guanghui, et al.
Veröffentlicht: (2025)
von: Wang, Guanghui, et al.
Veröffentlicht: (2025)
Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
von: Yoo, Seungju, et al.
Veröffentlicht: (2025)
von: Yoo, Seungju, et al.
Veröffentlicht: (2025)
Computing Time-varying Network Reliability using Binary Decision Diagrams
von: Nakahata, Yu, et al.
Veröffentlicht: (2025)
von: Nakahata, Yu, et al.
Veröffentlicht: (2025)
Association of Hypertension With Telomere Length, Considering Non‐Genetic and Genetic Factors, in Middle‐Aged Koreans
von: Younghwa Baek, et al.
Veröffentlicht: (2025)
von: Younghwa Baek, et al.
Veröffentlicht: (2025)
Leveraging Expert Consistency to Improve Algorithmic Decision Support
von: De-Arteaga, Maria, et al.
Veröffentlicht: (2021)
von: De-Arteaga, Maria, et al.
Veröffentlicht: (2021)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
Making Reliable and Flexible Decisions in Long-tailed Classification
von: Li, Bolian, et al.
Veröffentlicht: (2025)
von: Li, Bolian, et al.
Veröffentlicht: (2025)
Enhancing Text-Based Hierarchical Multilabel Classification for Mobile Applications via Contrastive Learning
von: Guo, Jiawei, et al.
Veröffentlicht: (2025)
von: Guo, Jiawei, et al.
Veröffentlicht: (2025)
Ontwikkelingsprincipes voor de Inrichting van de Informatievoorziening over de Curatieve zorg
von: de Vries Robbé, P.F., et al.
Veröffentlicht: (2013)
von: de Vries Robbé, P.F., et al.
Veröffentlicht: (2013)
From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications
von: Ma, Yongqiang, et al.
Veröffentlicht: (2024)
von: Ma, Yongqiang, et al.
Veröffentlicht: (2024)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
von: Bail, Mathis Le, et al.
Veröffentlicht: (2025)
von: Bail, Mathis Le, et al.
Veröffentlicht: (2025)
Text Clustering as Classification with LLMs
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
A Multimodal Manufacturing Safety Chatbot: Knowledge Base Design, Benchmark Development, and Evaluation of Multiple RAG Approaches
von: Singh, Ryan, et al.
Veröffentlicht: (2025)
von: Singh, Ryan, et al.
Veröffentlicht: (2025)
Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection
von: Lau, Matthew, et al.
Veröffentlicht: (2023)
von: Lau, Matthew, et al.
Veröffentlicht: (2023)
Evaluating LLMs for Text-to-SQL Generation With Complex SQL Workload
von: Ma, Limin, et al.
Veröffentlicht: (2024)
von: Ma, Limin, et al.
Veröffentlicht: (2024)
EvaluateXAI: A Framework to Evaluate the Reliability and Consistency of Rule-based XAI Techniques for Software Analytics Tasks
von: Awal, Md Abdul, et al.
Veröffentlicht: (2024)
von: Awal, Md Abdul, et al.
Veröffentlicht: (2024)
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
von: Subedi, Krishna
Veröffentlicht: (2025)
von: Subedi, Krishna
Veröffentlicht: (2025)
Remaining Useful Life Modelling With an Escalator Health Condition Analytic System
von: Inez M. Zwetsloot, et al.
Veröffentlicht: (2025)
von: Inez M. Zwetsloot, et al.
Veröffentlicht: (2025)
RADIANT-LLM: an Agentic Retrieval Augmented Generation Framework for Reliable Decision Support in Safety-Critical Nuclear Engineering
von: Ndum, Zavier Ndum, et al.
Veröffentlicht: (2026)
von: Ndum, Zavier Ndum, et al.
Veröffentlicht: (2026)
Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
RAmBLA: A Framework for Evaluating the Reliability of LLMs as Assistants in the Biomedical Domain
von: Bolton, William James, et al.
Veröffentlicht: (2024)
von: Bolton, William James, et al.
Veröffentlicht: (2024)
Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context
von: Jia, Jingru, et al.
Veröffentlicht: (2024)
von: Jia, Jingru, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Introducing ChatSQC: Enhancing Statistical Quality Control with Augmented AI
von: Megahed, Fadel M., et al.
Veröffentlicht: (2023) -
Adapting OpenAI's CLIP Model for Few-Shot Image Inspection in Manufacturing Quality Control: An Expository Case Study with Multiple Application Examples
von: Megahed, Fadel M., et al.
Veröffentlicht: (2025) -
ChatISA: A Prompt-Engineered, In-House Multi-Modal Generative AI Chatbot for Information Systems Education
von: Megahed, Fadel M., et al.
Veröffentlicht: (2024) -
Social Network Datasets on Reddit Financial Discussion
von: Wang, Zezhong, et al.
Veröffentlicht: (2024) -
Domain-Informed Negative Sampling Strategies for Dynamic Graph Embedding in Meme Stock-Related Social Networks
von: Hui, Yunming, et al.
Veröffentlicht: (2024)