Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Chenkai, Zhang, Denghui, Zhai, ChengXiang, Ji, Heng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TinyHelen's First Curriculum: Training and Evaluating Tiny Language Models in a Simpler Language Environment
by: Yang, Ke, et al.
Published: (2024)
by: Yang, Ke, et al.
Published: (2024)
Persona-DB: Efficient Large Language Model Personalization for Response Prediction with Collaborative Data Refinement
by: Sun, Chenkai, et al.
Published: (2024)
by: Sun, Chenkai, et al.
Published: (2024)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
User Simulation for Evaluating Information Access Systems
by: Balog, Krisztian, et al.
Published: (2023)
by: Balog, Krisztian, et al.
Published: (2023)
The Indispensable Role of User Simulation in the Pursuit of AGI
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction
by: Hao, Yuren, et al.
Published: (2026)
by: Hao, Yuren, et al.
Published: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
by: Yuan, Tongxin, et al.
Published: (2024)
by: Yuan, Tongxin, et al.
Published: (2024)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
by: Yang, Ke, et al.
Published: (2026)
by: Yang, Ke, et al.
Published: (2026)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
by: Ge, Tao, et al.
Published: (2026)
by: Ge, Tao, et al.
Published: (2026)
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
by: Brito, Iago Alves, et al.
Published: (2026)
by: Brito, Iago Alves, et al.
Published: (2026)
Sensitivity Meets Sparsity: The Impact of Extremely Sparse Parameter Patterns on Theory-of-Mind of Large Language Models
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
Preference-Aware Memory Update for Long-Term LLM Agents
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Provable Defense Framework for LLM Jailbreaks via Noise-Augumented Alignment
by: Cheng, Zehua, et al.
Published: (2026)
by: Cheng, Zehua, et al.
Published: (2026)
Globally Optimal Training of Spiking Neural Networks via Parameter Reconstruction
by: Udupi, Himanshu, et al.
Published: (2026)
by: Udupi, Himanshu, et al.
Published: (2026)
Efficient Knowledge Infusion via KG-LLM Alignment
by: Jiang, Zhouyu, et al.
Published: (2024)
by: Jiang, Zhouyu, et al.
Published: (2024)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
by: Li, Shuyue Stella, et al.
Published: (2026)
by: Li, Shuyue Stella, et al.
Published: (2026)
Searching for Privacy Risks in LLM Agents via Simulation
by: Zhang, Yanzhe, et al.
Published: (2025)
by: Zhang, Yanzhe, et al.
Published: (2025)
Causal Intervention-Based Memory Selection for Long-Horizon LLM Agents
by: Srivastava, Saksham Sahai
Published: (2026)
by: Srivastava, Saksham Sahai
Published: (2026)
Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations
by: Liu, Ziyang
Published: (2026)
by: Liu, Ziyang
Published: (2026)
Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts
by: Wang, Rushi, et al.
Published: (2025)
by: Wang, Rushi, et al.
Published: (2025)
Measuring Copyright Risks of Large Language Model via Partial Information Probing
by: Zhao, Weijie, et al.
Published: (2024)
by: Zhao, Weijie, et al.
Published: (2024)
Ten Principles of AI Agent Economics
by: Yang, Ke, et al.
Published: (2025)
by: Yang, Ke, et al.
Published: (2025)
Computation Mechanism Behind LLM Position Generalization
by: Han, Chi, et al.
Published: (2025)
by: Han, Chi, et al.
Published: (2025)
LongSafety: Evaluating Long-Context Safety of Large Language Models
by: Lu, Yida, et al.
Published: (2025)
by: Lu, Yida, et al.
Published: (2025)
PMMT: Preference Alignment in Multilingual Machine Translation via LLM Distillation
by: Sun, Shuqiao, et al.
Published: (2024)
by: Sun, Shuqiao, et al.
Published: (2024)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
by: Chen, Luyu, et al.
Published: (2025)
by: Chen, Luyu, et al.
Published: (2025)
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
by: Luo, Haotian, et al.
Published: (2025)
by: Luo, Haotian, et al.
Published: (2025)
ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning
by: Huang, Fan
Published: (2026)
by: Huang, Fan
Published: (2026)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
Safety Alignment via Constrained Knowledge Unlearning
by: Shi, Zesheng, et al.
Published: (2025)
by: Shi, Zesheng, et al.
Published: (2025)
Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
by: Gao, Jiaxuan, et al.
Published: (2025)
by: Gao, Jiaxuan, et al.
Published: (2025)
DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
by: Yang, Junxiao, et al.
Published: (2026)
by: Yang, Junxiao, et al.
Published: (2026)
LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages
by: Lu, Yinquan, et al.
Published: (2024)
by: Lu, Yinquan, et al.
Published: (2024)
MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
by: Wang, Yuhui, et al.
Published: (2026)
by: Wang, Yuhui, et al.
Published: (2026)
REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
by: Chu, Zheng, et al.
Published: (2026)
by: Chu, Zheng, et al.
Published: (2026)
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
by: Hao, Haochang, et al.
Published: (2026)
by: Hao, Haochang, et al.
Published: (2026)
Similar Items
-
TinyHelen's First Curriculum: Training and Evaluating Tiny Language Models in a Simpler Language Environment
by: Yang, Ke, et al.
Published: (2024) -
Persona-DB: Efficient Large Language Model Personalization for Response Prediction with Collaborative Data Refinement
by: Sun, Chenkai, et al.
Published: (2024) -
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025) -
User Simulation for Evaluating Information Access Systems
by: Balog, Krisztian, et al.
Published: (2023) -
The Indispensable Role of User Simulation in the Pursuit of AGI
by: Balog, Krisztian, et al.
Published: (2025)