History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions
Fuente:
arXiv
Guardado en:
| Autor principal: | Salgado, Alberto G. Rodríguez |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Understanding Unsafe Video Generation
por: Pang, Yan, et al.
Publicado: (2024)
por: Pang, Yan, et al.
Publicado: (2024)
PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs
por: Gutiérrez, Juan, et al.
Publicado: (2026)
por: Gutiérrez, Juan, et al.
Publicado: (2026)
Safe Vision-Language Models via Unsafe Weights Manipulation
por: D'Incà, Moreno, et al.
Publicado: (2025)
por: D'Incà, Moreno, et al.
Publicado: (2025)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
por: He, Jingtao, et al.
Publicado: (2026)
por: He, Jingtao, et al.
Publicado: (2026)
HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios
por: Pu, Jiayue, et al.
Publicado: (2026)
por: Pu, Jiayue, et al.
Publicado: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
por: Zhang, Junyang, et al.
Publicado: (2025)
por: Zhang, Junyang, et al.
Publicado: (2025)
Context Matters: Peer-Aware Student Behavioral Engagement Measurement via VLM Action Parsing and LLM Sequence Classification
por: Abdelkawy, Ahmed, et al.
Publicado: (2026)
por: Abdelkawy, Ahmed, et al.
Publicado: (2026)
AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation
por: Zhang, Jian, et al.
Publicado: (2026)
por: Zhang, Jian, et al.
Publicado: (2026)
Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection
por: Fu, Jinhu, et al.
Publicado: (2026)
por: Fu, Jinhu, et al.
Publicado: (2026)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
por: Pham, Trong-Thang, et al.
Publicado: (2026)
por: Pham, Trong-Thang, et al.
Publicado: (2026)
iPay: Integrated Payment Action Recognition via Multimodal Networks and Adaptive Spatial Prior Learning
por: Huang, Kaicong, et al.
Publicado: (2026)
por: Huang, Kaicong, et al.
Publicado: (2026)
Bi-Anchor Interpolation Solver for Accelerating Generative Modeling
por: Chen, Hongxu, et al.
Publicado: (2026)
por: Chen, Hongxu, et al.
Publicado: (2026)
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
por: Zhu, Hongguang, et al.
Publicado: (2025)
por: Zhu, Hongguang, et al.
Publicado: (2025)
DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models
por: Hao, Yuhan, et al.
Publicado: (2025)
por: Hao, Yuhan, et al.
Publicado: (2025)
RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving
por: Wang, Yujin, et al.
Publicado: (2025)
por: Wang, Yujin, et al.
Publicado: (2025)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
por: Yang, Zhuoyi, et al.
Publicado: (2026)
por: Yang, Zhuoyi, et al.
Publicado: (2026)
ReGenNet: Towards Human Action-Reaction Synthesis
por: Xu, Liang, et al.
Publicado: (2024)
por: Xu, Liang, et al.
Publicado: (2024)
Inline Critic Steers Image Editing
por: Kang, Weitai, et al.
Publicado: (2026)
por: Kang, Weitai, et al.
Publicado: (2026)
VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
por: Yang, Xindi, et al.
Publicado: (2025)
por: Yang, Xindi, et al.
Publicado: (2025)
Beyond Fixed Anchors: Precisely Erasing Concepts with Sibling Exclusive Counterparts
por: Zhang, Tong, et al.
Publicado: (2025)
por: Zhang, Tong, et al.
Publicado: (2025)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
por: Guo, Jun, et al.
Publicado: (2026)
por: Guo, Jun, et al.
Publicado: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
por: Yuan, Lingzhi, et al.
Publicado: (2025)
por: Yuan, Lingzhi, et al.
Publicado: (2025)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
por: Xiang, Yuxiao, et al.
Publicado: (2025)
por: Xiang, Yuxiao, et al.
Publicado: (2025)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
por: Benavent-Lledo, Manuel, et al.
Publicado: (2024)
por: Benavent-Lledo, Manuel, et al.
Publicado: (2024)
Decoding Vision Transformers: the Diffusion Steering Lens
por: Takatsuki, Ryota, et al.
Publicado: (2025)
por: Takatsuki, Ryota, et al.
Publicado: (2025)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
por: Wang, Zun, et al.
Publicado: (2026)
por: Wang, Zun, et al.
Publicado: (2026)
ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance
por: Yang, Yang, et al.
Publicado: (2026)
por: Yang, Yang, et al.
Publicado: (2026)
PBADet: A One-Stage Anchor-Free Approach for Part-Body Association
por: Gao, Zhongpai, et al.
Publicado: (2024)
por: Gao, Zhongpai, et al.
Publicado: (2024)
Proxy-Anchor and EVT-Driven Continual Learning Method for Generalized Category Discovery
por: Fathalizadeh, Alireza, et al.
Publicado: (2025)
por: Fathalizadeh, Alireza, et al.
Publicado: (2025)
PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning
por: Xi, Yingjie, et al.
Publicado: (2025)
por: Xi, Yingjie, et al.
Publicado: (2025)
Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
por: Nadeem, Numair, et al.
Publicado: (2025)
por: Nadeem, Numair, et al.
Publicado: (2025)
IVAC-P2L: Leveraging Irregular Repetition Priors for Improving Video Action Counting
por: Wang, Hang, et al.
Publicado: (2024)
por: Wang, Hang, et al.
Publicado: (2024)
How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
por: Zha, Jirong, et al.
Publicado: (2025)
por: Zha, Jirong, et al.
Publicado: (2025)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
por: Li, Yuan-Ming, et al.
Publicado: (2024)
por: Li, Yuan-Ming, et al.
Publicado: (2024)
Towards End-to-End Neuromorphic Event-based 3D Object Reconstruction Without Physical Priors
por: Xu, Chuanzhi, et al.
Publicado: (2025)
por: Xu, Chuanzhi, et al.
Publicado: (2025)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
por: Meng, Xiangtao, et al.
Publicado: (2025)
por: Meng, Xiangtao, et al.
Publicado: (2025)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
por: Wang, Zun, et al.
Publicado: (2025)
por: Wang, Zun, et al.
Publicado: (2025)
A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition
por: Shen, Yaomin, et al.
Publicado: (2025)
por: Shen, Yaomin, et al.
Publicado: (2025)
Learning with Instance-Dependent Noisy Labels by Anchor Hallucination and Hard Sample Label Correction
por: Huang, Po-Hsuan, et al.
Publicado: (2024)
por: Huang, Po-Hsuan, et al.
Publicado: (2024)
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
por: Guo, Chaohong, et al.
Publicado: (2025)
por: Guo, Chaohong, et al.
Publicado: (2025)
Ejemplares similares
-
Towards Understanding Unsafe Video Generation
por: Pang, Yan, et al.
Publicado: (2024) -
PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs
por: Gutiérrez, Juan, et al.
Publicado: (2026) -
Safe Vision-Language Models via Unsafe Weights Manipulation
por: D'Incà, Moreno, et al.
Publicado: (2025) -
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
por: He, Jingtao, et al.
Publicado: (2026) -
HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios
por: Pu, Jiayue, et al.
Publicado: (2026)