Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
Fuente:
arXiv
Saved in:
| Main Authors: | Son, Yejin, Kim, Minseo, Kim, Sungwoong, Han, Seungju, Kim, Jian, Jang, Dongju, Yu, Youngjae, Park, Chanyoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Preference Modeling for Cross-Session Personalized Tool Calling
by: Yoon, Yejin, et al.
Published: (2026)
by: Yoon, Yejin, et al.
Published: (2026)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
Ion‐Regulating SPEEK–BNNS Hybrid Interfaces Enabling Low‐Barrier Zn‐Ion Transport and Dendrite‐Free Zinc Anodes
by: Minseo Kim, et al.
Published: (2026)
by: Minseo Kim, et al.
Published: (2026)
Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
by: Kim, Shubin, et al.
Published: (2026)
by: Kim, Shubin, et al.
Published: (2026)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
by: In, Yeonjun, et al.
Published: (2025)
by: In, Yeonjun, et al.
Published: (2025)
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms
by: Lim, Seungwon, et al.
Published: (2025)
by: Lim, Seungwon, et al.
Published: (2025)
CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
by: Choi, Suhwan, et al.
Published: (2024)
by: Choi, Suhwan, et al.
Published: (2024)
Towards Continuous Sign Language Conversation from Isolated Signs
by: Kim, Youngmin, et al.
Published: (2026)
by: Kim, Youngmin, et al.
Published: (2026)
Reasoning Structure Matters for Safety Alignment of Reasoning Models
by: In, Yeonjun, et al.
Published: (2026)
by: In, Yeonjun, et al.
Published: (2026)
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
by: Kim, Yoonsu, et al.
Published: (2025)
by: Kim, Yoonsu, et al.
Published: (2025)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
by: Kim, Jiwan, et al.
Published: (2025)
by: Kim, Jiwan, et al.
Published: (2025)
Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
by: Lee, Seungbeen, et al.
Published: (2025)
by: Lee, Seungbeen, et al.
Published: (2025)
Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
by: Yoon, Yejin, et al.
Published: (2025)
by: Yoon, Yejin, et al.
Published: (2025)
SLIP & ETHICS: Graduated Intervention for AI Emotional Companions
by: Kim, Minseo
Published: (2026)
by: Kim, Minseo
Published: (2026)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
by: In, Yeonjun, et al.
Published: (2025)
by: In, Yeonjun, et al.
Published: (2025)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
by: Lee, Sangyub, et al.
Published: (2026)
by: Lee, Sangyub, et al.
Published: (2026)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
by: Kim, Youngjae, et al.
Published: (2024)
by: Kim, Youngjae, et al.
Published: (2024)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
by: Park, Jaewoo, et al.
Published: (2025)
by: Park, Jaewoo, et al.
Published: (2025)
On the size of universal graphs for spanning trees
by: Kim, Jaehoon, et al.
Published: (2025)
by: Kim, Jaehoon, et al.
Published: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
Floquet Chern Insulators and Radiation-Induced Zero Resistance in Irradiated Graphene
by: Kim, Youngjae, et al.
Published: (2025)
by: Kim, Youngjae, et al.
Published: (2025)
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
by: Hyun, Lee, et al.
Published: (2023)
by: Hyun, Lee, et al.
Published: (2023)
Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
by: Kim, Sewon, et al.
Published: (2025)
by: Kim, Sewon, et al.
Published: (2025)
A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
by: Aoyagui, Paula Akemi, et al.
Published: (2025)
by: Aoyagui, Paula Akemi, et al.
Published: (2025)
Coupling-Robust Accuracy in Multiphysics Physics Informed Neural Networks via Kronecker-Preconditioned Optimization
by: Park, Youngjae, et al.
Published: (2026)
by: Park, Youngjae, et al.
Published: (2026)
Light-Wave Engineering for Selective Polarization of a Single $\mathbf{Q}$ Valley in Transition Metal Dichalcogenides
by: Kim, Youngjae
Published: (2025)
by: Kim, Youngjae
Published: (2025)
Pseudospins revealed through the giant dynamical Franz-Keldysh effect in massless Dirac materials
by: Kim, Youngjae
Published: (2024)
by: Kim, Youngjae
Published: (2024)
Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding
by: Jeon, Jaehyun, et al.
Published: (2025)
by: Jeon, Jaehyun, et al.
Published: (2025)
DUSK: Do Not Unlearn Shared Knowledge
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
GraphCliff: Short-Long Range Gating for Subtle Differences but Critical Changes
by: Kim, Hajung, et al.
Published: (2025)
by: Kim, Hajung, et al.
Published: (2025)
Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
by: Jeon, Yejin, et al.
Published: (2025)
by: Jeon, Yejin, et al.
Published: (2025)
Is GPT-4 Alone Sufficient for Automated Essay Scoring?: A Comparative Judgment Approach Based on Rater Cognition
by: Kim, Seungju, et al.
Published: (2024)
by: Kim, Seungju, et al.
Published: (2024)
Representation Bending for Large Language Model Safety
by: Yousefpour, Ashkan, et al.
Published: (2025)
by: Yousefpour, Ashkan, et al.
Published: (2025)
Memorize Early, Then Query: Inlier-Memorization-Guided Active Outlier Detection
by: Kang, Minseo, et al.
Published: (2026)
by: Kang, Minseo, et al.
Published: (2026)
X-ray studies of PSR J1838$-$0655 and its wind nebula associated with HESS J1837$-$069 and 1LHAASO J1837$-$0654u
by: Park, Minseo, et al.
Published: (2025)
by: Park, Minseo, et al.
Published: (2025)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
by: Ham, Seokil, et al.
Published: (2025)
by: Ham, Seokil, et al.
Published: (2025)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2024)
by: Kim, Kibum, et al.
Published: (2024)
Similar Items
-
Latent Preference Modeling for Cross-Session Personalized Tool Calling
by: Yoon, Yejin, et al.
Published: (2026) -
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
by: Chung, Jiwan, et al.
Published: (2024) -
Ion‐Regulating SPEEK–BNNS Hybrid Interfaces Enabling Low‐Barrier Zn‐Ion Transport and Dendrite‐Free Zinc Anodes
by: Minseo Kim, et al.
Published: (2026) -
Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
by: Kim, Shubin, et al.
Published: (2026) -
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
by: In, Yeonjun, et al.
Published: (2025)