Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yadav, Srishti, Gajcin, Jasmina, Miehling, Erik, Daly, Elizabeth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
Localizing Persona Representations in LLMs
von: Cintas, Celia, et al.
Veröffentlicht: (2025)
von: Cintas, Celia, et al.
Veröffentlicht: (2025)
FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
von: Carnerero-Cano, Javier, et al.
Veröffentlicht: (2026)
von: Carnerero-Cano, Javier, et al.
Veröffentlicht: (2026)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
ACTER: Diverse and Actionable Counterfactual Sequences for Explaining and Diagnosing RL Policies
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
Redefining Counterfactual Explanations for Reinforcement Learning: Overview, Challenges and Opportunities
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2022)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2022)
CELL your Model: Contrastive Explanations for Large Language Models
von: Luss, Ronny, et al.
Veröffentlicht: (2024)
von: Luss, Ronny, et al.
Veröffentlicht: (2024)
Evaluating the Prompt Steerability of Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
Semifactual Explanations for Reinforcement Learning
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
Multimodal Analysis of State-Funded News Coverage of the Israel-Hamas War on YouTube Shorts
von: Miehling, Daniel, et al.
Veröffentlicht: (2026)
von: Miehling, Daniel, et al.
Veröffentlicht: (2026)
AI Steerability 360: A Toolkit for Steering Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2026)
von: Miehling, Erik, et al.
Veröffentlicht: (2026)
Who's Who: Large Language Models Meet Knowledge Conflicts in Practice
von: Pham, Quang Hieu, et al.
Veröffentlicht: (2024)
von: Pham, Quang Hieu, et al.
Veröffentlicht: (2024)
A Brief Summary of Explanatory Virtues
von: Zukerman, Ingrid
Veröffentlicht: (2024)
von: Zukerman, Ingrid
Veröffentlicht: (2024)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
von: Al-Tawaha, Ahmad, et al.
Veröffentlicht: (2026)
von: Al-Tawaha, Ahmad, et al.
Veröffentlicht: (2026)
Generative LLM Powered Conversational AI Application for Personalized Risk Assessment: A Case Study in COVID-19
von: Roshani, Mohammad Amin, et al.
Veröffentlicht: (2024)
von: Roshani, Mohammad Amin, et al.
Veröffentlicht: (2024)
Cross-modal Information Flow in Multimodal Large Language Models
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
Who can we trust? LLM-as-a-jury for Comparative Assessment
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
von: Fang, Tianqing, et al.
Veröffentlicht: (2023)
von: Fang, Tianqing, et al.
Veröffentlicht: (2023)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
Depression Risk Assessment in Social Media via Large Language Models
von: Gulino, Giorgia, et al.
Veröffentlicht: (2026)
von: Gulino, Giorgia, et al.
Veröffentlicht: (2026)
Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
von: Cui, Tianyu, et al.
Veröffentlicht: (2024)
von: Cui, Tianyu, et al.
Veröffentlicht: (2024)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
von: Hou, Yufang, et al.
Veröffentlicht: (2024)
von: Hou, Yufang, et al.
Veröffentlicht: (2024)
FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models
von: Marinescu, Radu, et al.
Veröffentlicht: (2025)
von: Marinescu, Radu, et al.
Veröffentlicht: (2025)
Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
von: Xu, Derong, et al.
Veröffentlicht: (2024)
von: Xu, Derong, et al.
Veröffentlicht: (2024)
Explanatory Summarization with Discourse-Driven Planning
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
Micro-Act: Mitigating Knowledge Conflict in LLM-based RAG via Actionable Self-Reasoning
von: Huo, Nan, et al.
Veröffentlicht: (2025)
von: Huo, Nan, et al.
Veröffentlicht: (2025)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
von: Ginjala, Srishti, et al.
Veröffentlicht: (2026)
von: Ginjala, Srishti, et al.
Veröffentlicht: (2026)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
von: Zhou, Lexin, et al.
Veröffentlicht: (2025)
von: Zhou, Lexin, et al.
Veröffentlicht: (2025)
"Who Am I, and Who Else Is Here?" Behavioral Differentiation Without Role Assignment in Multi-Agent LLM Systems
von: Kandoussi, Houssam EL
Veröffentlicht: (2026)
von: Kandoussi, Houssam EL
Veröffentlicht: (2026)
Towards Robust ESG Analysis Against Greenwashing Risks: Aspect-Action Analysis with Cross-Category Generalization
von: Ong, Keane, et al.
Veröffentlicht: (2025)
von: Ong, Keane, et al.
Veröffentlicht: (2025)
Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2026)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2026)
LLM-Augmented Symptom Analysis for Cardiovascular Disease Risk Prediction: A Clinical NLP
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
von: Sun, Chenkai, et al.
Veröffentlicht: (2025)
von: Sun, Chenkai, et al.
Veröffentlicht: (2025)
Verification Limits Code LLM Training
von: Gureja, Srishti, et al.
Veröffentlicht: (2025)
von: Gureja, Srishti, et al.
Veröffentlicht: (2025)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
von: Ingimundarson, Finnur Ágúst, et al.
Veröffentlicht: (2026)
von: Ingimundarson, Finnur Ágúst, et al.
Veröffentlicht: (2026)
Agentic AI Needs a Systems Theory
von: Miehling, Erik, et al.
Veröffentlicht: (2025)
von: Miehling, Erik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025) -
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025) -
Localizing Persona Representations in LLMs
von: Cintas, Celia, et al.
Veröffentlicht: (2025) -
FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
von: Carnerero-Cano, Javier, et al.
Veröffentlicht: (2026) -
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
von: Miehling, Erik, et al.
Veröffentlicht: (2024)