Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Son, Yejin, Kim, Minseo, Kim, Sungwoong, Han, Seungju, Kim, Jian, Jang, Dongju, Yu, Youngjae, Park, Chanyoung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Latent Preference Modeling for Cross-Session Personalized Tool Calling
von: Yoon, Yejin, et al.
Veröffentlicht: (2026)
von: Yoon, Yejin, et al.
Veröffentlicht: (2026)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
Ion‐Regulating SPEEK–BNNS Hybrid Interfaces Enabling Low‐Barrier Zn‐Ion Transport and Dendrite‐Free Zinc Anodes
von: Minseo Kim, et al.
Veröffentlicht: (2026)
von: Minseo Kim, et al.
Veröffentlicht: (2026)
Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
von: Kim, Shubin, et al.
Veröffentlicht: (2026)
von: Kim, Shubin, et al.
Veröffentlicht: (2026)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
von: Choi, Suhwan, et al.
Veröffentlicht: (2024)
von: Choi, Suhwan, et al.
Veröffentlicht: (2024)
Towards Continuous Sign Language Conversation from Isolated Signs
von: Kim, Youngmin, et al.
Veröffentlicht: (2026)
von: Kim, Youngmin, et al.
Veröffentlicht: (2026)
Reasoning Structure Matters for Safety Alignment of Reasoning Models
von: In, Yeonjun, et al.
Veröffentlicht: (2026)
von: In, Yeonjun, et al.
Veröffentlicht: (2026)
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
von: Kim, Yoonsu, et al.
Veröffentlicht: (2025)
von: Kim, Yoonsu, et al.
Veröffentlicht: (2025)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
von: Lee, Seungbeen, et al.
Veröffentlicht: (2025)
von: Lee, Seungbeen, et al.
Veröffentlicht: (2025)
Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
von: Yoon, Yejin, et al.
Veröffentlicht: (2025)
von: Yoon, Yejin, et al.
Veröffentlicht: (2025)
SLIP & ETHICS: Graduated Intervention for AI Emotional Companions
von: Kim, Minseo
Veröffentlicht: (2026)
von: Kim, Minseo
Veröffentlicht: (2026)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
von: Lee, Sangyub, et al.
Veröffentlicht: (2026)
von: Lee, Sangyub, et al.
Veröffentlicht: (2026)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
von: Han, Seungju, et al.
Veröffentlicht: (2024)
von: Han, Seungju, et al.
Veröffentlicht: (2024)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
von: Park, Jaewoo, et al.
Veröffentlicht: (2025)
von: Park, Jaewoo, et al.
Veröffentlicht: (2025)
On the size of universal graphs for spanning trees
von: Kim, Jaehoon, et al.
Veröffentlicht: (2025)
von: Kim, Jaehoon, et al.
Veröffentlicht: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
von: Kim, Junhyeok, et al.
Veröffentlicht: (2025)
von: Kim, Junhyeok, et al.
Veröffentlicht: (2025)
Floquet Chern Insulators and Radiation-Induced Zero Resistance in Irradiated Graphene
von: Kim, Youngjae, et al.
Veröffentlicht: (2025)
von: Kim, Youngjae, et al.
Veröffentlicht: (2025)
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
von: Hyun, Lee, et al.
Veröffentlicht: (2023)
von: Hyun, Lee, et al.
Veröffentlicht: (2023)
Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
von: Kim, Sewon, et al.
Veröffentlicht: (2025)
von: Kim, Sewon, et al.
Veröffentlicht: (2025)
A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
von: Aoyagui, Paula Akemi, et al.
Veröffentlicht: (2025)
von: Aoyagui, Paula Akemi, et al.
Veröffentlicht: (2025)
Coupling-Robust Accuracy in Multiphysics Physics Informed Neural Networks via Kronecker-Preconditioned Optimization
von: Park, Youngjae, et al.
Veröffentlicht: (2026)
von: Park, Youngjae, et al.
Veröffentlicht: (2026)
Light-Wave Engineering for Selective Polarization of a Single $\mathbf{Q}$ Valley in Transition Metal Dichalcogenides
von: Kim, Youngjae
Veröffentlicht: (2025)
von: Kim, Youngjae
Veröffentlicht: (2025)
Pseudospins revealed through the giant dynamical Franz-Keldysh effect in massless Dirac materials
von: Kim, Youngjae
Veröffentlicht: (2024)
von: Kim, Youngjae
Veröffentlicht: (2024)
Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding
von: Jeon, Jaehyun, et al.
Veröffentlicht: (2025)
von: Jeon, Jaehyun, et al.
Veröffentlicht: (2025)
DUSK: Do Not Unlearn Shared Knowledge
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
GraphCliff: Short-Long Range Gating for Subtle Differences but Critical Changes
von: Kim, Hajung, et al.
Veröffentlicht: (2025)
von: Kim, Hajung, et al.
Veröffentlicht: (2025)
Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
von: Jeon, Yejin, et al.
Veröffentlicht: (2025)
von: Jeon, Yejin, et al.
Veröffentlicht: (2025)
Is GPT-4 Alone Sufficient for Automated Essay Scoring?: A Comparative Judgment Approach Based on Rater Cognition
von: Kim, Seungju, et al.
Veröffentlicht: (2024)
von: Kim, Seungju, et al.
Veröffentlicht: (2024)
Representation Bending for Large Language Model Safety
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
Memorize Early, Then Query: Inlier-Memorization-Guided Active Outlier Detection
von: Kang, Minseo, et al.
Veröffentlicht: (2026)
von: Kang, Minseo, et al.
Veröffentlicht: (2026)
X-ray studies of PSR J1838$-$0655 and its wind nebula associated with HESS J1837$-$069 and 1LHAASO J1837$-$0654u
von: Park, Minseo, et al.
Veröffentlicht: (2025)
von: Park, Minseo, et al.
Veröffentlicht: (2025)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
von: Li, Manling, et al.
Veröffentlicht: (2024)
von: Li, Manling, et al.
Veröffentlicht: (2024)
Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
von: Ham, Seokil, et al.
Veröffentlicht: (2025)
von: Ham, Seokil, et al.
Veröffentlicht: (2025)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
von: Kim, Kibum, et al.
Veröffentlicht: (2024)
von: Kim, Kibum, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Latent Preference Modeling for Cross-Session Personalized Tool Calling
von: Yoon, Yejin, et al.
Veröffentlicht: (2026) -
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
von: Chung, Jiwan, et al.
Veröffentlicht: (2024) -
Ion‐Regulating SPEEK–BNNS Hybrid Interfaces Enabling Low‐Barrier Zn‐Ion Transport and Dendrite‐Free Zinc Anodes
von: Minseo Kim, et al.
Veröffentlicht: (2026) -
Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
von: Kim, Shubin, et al.
Veröffentlicht: (2026) -
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
von: In, Yeonjun, et al.
Veröffentlicht: (2025)