Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Ke, Kejriwal, Mayank |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Evaluation of Estimative Uncertainty in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
SelECT-SQL: Self-correcting ensemble Chain-of-Thought for Text-to-SQL
von: Shen, Ke, et al.
Veröffentlicht: (2024)
von: Shen, Ke, et al.
Veröffentlicht: (2024)
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
HALO: An Ontology for Representing and Categorizing Hallucinations in Large Language Models
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023)
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023)
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
Navigating Semantic Relations: Challenges for Language Models in Abstract Common-Sense Reasoning
von: Gawin, Cole, et al.
Veröffentlicht: (2025)
von: Gawin, Cole, et al.
Veröffentlicht: (2025)
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
Is persona enough for personality? Using ChatGPT to reconstruct an agent's latent personality from simple descriptions
von: Ji, Yongyi, et al.
Veröffentlicht: (2024)
von: Ji, Yongyi, et al.
Veröffentlicht: (2024)
Evaluating Ill-Defined Tasks in Large Language Models
von: Zhou, Yi, et al.
Veröffentlicht: (2026)
von: Zhou, Yi, et al.
Veröffentlicht: (2026)
Pushing the boundary on Natural Language Inference
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
Cross-lingual Editing in Multilingual Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
Evaluating Morphological Compositional Generalization in Large Language Models
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
Building Efficient Universal Classifiers with Natural Language Inference
von: Laurer, Moritz, et al.
Veröffentlicht: (2023)
von: Laurer, Moritz, et al.
Veröffentlicht: (2023)
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
von: Madusanka, Tharindu, et al.
Veröffentlicht: (2025)
von: Madusanka, Tharindu, et al.
Veröffentlicht: (2025)
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
von: Hwang, Yeonjun, et al.
Veröffentlicht: (2026)
von: Hwang, Yeonjun, et al.
Veröffentlicht: (2026)
On the Decision-Making Abilities in Role-Playing using Large Language Models
von: Shen, Chenglei, et al.
Veröffentlicht: (2024)
von: Shen, Chenglei, et al.
Veröffentlicht: (2024)
WRDScore: New Metric for Evaluation of Natural Language Generation Models
von: Mussabayev, Ravil
Veröffentlicht: (2024)
von: Mussabayev, Ravil
Veröffentlicht: (2024)
TinyHelen's First Curriculum: Training and Evaluating Tiny Language Models in a Simpler Language Environment
von: Yang, Ke, et al.
Veröffentlicht: (2024)
von: Yang, Ke, et al.
Veröffentlicht: (2024)
Towards Coarse-to-Fine Evaluation of Inference Efficiency for Large Language Models
von: Chen, Yushuo, et al.
Veröffentlicht: (2024)
von: Chen, Yushuo, et al.
Veröffentlicht: (2024)
A Survey on Efficient Inference for Large Language Models
von: Zhou, Zixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Zixuan, et al.
Veröffentlicht: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
Relation-based Counterfactual Data Augmentation and Contrastive Learning for Robustifying Natural Language Inference Models
von: Yang, Heerin, et al.
Veröffentlicht: (2024)
von: Yang, Heerin, et al.
Veröffentlicht: (2024)
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
von: Weir, Nathaniel, et al.
Veröffentlicht: (2024)
von: Weir, Nathaniel, et al.
Veröffentlicht: (2024)
Filling the Gap: Is Commonsense Knowledge Generation useful for Natural Language Inference?
von: Jayaweera, Chathuri, et al.
Veröffentlicht: (2025)
von: Jayaweera, Chathuri, et al.
Veröffentlicht: (2025)
POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
von: Chen, Yizhuo, et al.
Veröffentlicht: (2025)
von: Chen, Yizhuo, et al.
Veröffentlicht: (2025)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
von: Wu, Yulong, et al.
Veröffentlicht: (2025)
von: Wu, Yulong, et al.
Veröffentlicht: (2025)
Code-Driven Planning in Grid Worlds with Large Language Models
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2025)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2025)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
An Analysis of Artificial Intelligence Adoption in NIH-Funded Research
von: Nananukul, Navapat, et al.
Veröffentlicht: (2026)
von: Nananukul, Navapat, et al.
Veröffentlicht: (2026)
Compositional Causal Reasoning Evaluation in Language Models
von: Maasch, Jacqueline R. M. A., et al.
Veröffentlicht: (2025)
von: Maasch, Jacqueline R. M. A., et al.
Veröffentlicht: (2025)
An analysis of AI Decision under Risk: Prospect theory emerges in Large Language Models
von: Payne, Kenneth
Veröffentlicht: (2025)
von: Payne, Kenneth
Veröffentlicht: (2025)
A Novel Cartography-Based Curriculum Learning Method Applied on RoNLI: The First Romanian Natural Language Inference Corpus
von: Poesina, Eduard, et al.
Veröffentlicht: (2024)
von: Poesina, Eduard, et al.
Veröffentlicht: (2024)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
An Evaluation of Estimative Uncertainty in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024) -
SelECT-SQL: Self-correcting ensemble Chain-of-Thought for Text-to-SQL
von: Shen, Ke, et al.
Veröffentlicht: (2024) -
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024) -
HALO: An Ontology for Representing and Categorizing Hallucinations in Large Language Models
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023) -
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)