Saved in:
| Main Authors: | Wang, Lichao, Ren, Zhaoxing, Yang, Tianzhuo, Ji, Jiaming, Liu, Chi Harold, Yang, Yaodong, Dai, Juntao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2606.01991 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Game-Theoretic Negotiation Framework for Cross-Cultural Consensus in LLMs
by: Zhang, Guoxi, et al.
Published: (2025)
by: Zhang, Guoxi, et al.
Published: (2025)
VISA: Value Injection via Shielded Adaptation for Personalized LLM Alignment
by: Chen, Jiawei, et al.
Published: (2026)
by: Chen, Jiawei, et al.
Published: (2026)
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
by: Zhou, Jiayi, et al.
Published: (2024)
by: Zhou, Jiayi, et al.
Published: (2024)
SafeLawBench: Towards Safe Alignment of Large Language Models
by: Cao, Chuxue, et al.
Published: (2025)
by: Cao, Chuxue, et al.
Published: (2025)
SafeMT: Multi-turn Safety for Multimodal Language Models
by: Zhu, Han, et al.
Published: (2025)
by: Zhu, Han, et al.
Published: (2025)
SafeDreamer: Safe Reinforcement Learning with World Models
by: Huang, Weidong, et al.
Published: (2023)
by: Huang, Weidong, et al.
Published: (2023)
ThinkPatterns-21k: A Systematic Study on the Impact of Thinking Patterns in LLMs
by: Wen, Pengcheng, et al.
Published: (2025)
by: Wen, Pengcheng, et al.
Published: (2025)
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
by: Dai, Josef, et al.
Published: (2024)
by: Dai, Josef, et al.
Published: (2024)
Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry
by: Zhang, Guoxi, et al.
Published: (2026)
by: Zhang, Guoxi, et al.
Published: (2026)
SafeEditor: Unified MLLM for Efficient Post-hoc T2I Safety Editing
by: Zhang, Ruiyang, et al.
Published: (2025)
by: Zhang, Ruiyang, et al.
Published: (2025)
Look-Ahead Reasoning on Learning Platforms
by: Zhu, Haiqing, et al.
Published: (2025)
by: Zhu, Haiqing, et al.
Published: (2025)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
by: Bu, Yuyan, et al.
Published: (2026)
by: Bu, Yuyan, et al.
Published: (2026)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
by: Zhang, Borong, et al.
Published: (2025)
by: Zhang, Borong, et al.
Published: (2025)
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
by: Bu, Yuyan, et al.
Published: (2026)
by: Bu, Yuyan, et al.
Published: (2026)
Aligner: Efficient Alignment by Learning to Correct
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
Look One Step Ahead: Forward-Looking Incentive Design with Strategic Privacy for Proactive Service Provisioning over Air-Ground Integrated Edge Networks
by: Wu, Sicheng, et al.
Published: (2026)
by: Wu, Sicheng, et al.
Published: (2026)
MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
by: Yang, Tianzhuo, et al.
Published: (2026)
by: Yang, Tianzhuo, et al.
Published: (2026)
SAE-V: Interpreting Multimodal Models for Enhanced Alignment
by: Lou, Hantao, et al.
Published: (2025)
by: Lou, Hantao, et al.
Published: (2025)
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction
by: Lou, Hantao, et al.
Published: (2025)
by: Lou, Hantao, et al.
Published: (2025)
Language Models Resist Alignment: Evidence From Data Compression
by: Ji, Jiaming, et al.
Published: (2024)
by: Ji, Jiaming, et al.
Published: (2024)
ProgressGym: Alignment with a Millennium of Moral Progress
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
What, Whether and How? Unveiling Process Reward Models for Thinking with Images Reasoning
by: Zhou, Yujin, et al.
Published: (2026)
by: Zhou, Yujin, et al.
Published: (2026)
No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors
by: Liu, Tao, et al.
Published: (2026)
by: Liu, Tao, et al.
Published: (2026)
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation
by: Dai, Juntao, et al.
Published: (2024)
by: Dai, Juntao, et al.
Published: (2024)
Look Ahead Text Understanding and LLM Stitching
by: Jiang, Junlin Julian, et al.
Published: (2024)
by: Jiang, Junlin Julian, et al.
Published: (2024)
Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
by: Xiang, Yang, et al.
Published: (2025)
by: Xiang, Yang, et al.
Published: (2025)
Look-Ahead Screening Rules for the Lasso
by: Larsson, Johan
Published: (2021)
by: Larsson, Johan
Published: (2021)
Discovering Self-Regulated Learning Patterns in Chatbot-Powered Education Environment
by: Lyu, Yilin, et al.
Published: (2025)
by: Lyu, Yilin, et al.
Published: (2025)
Look-Ahead and Look-Back Flows: Training-Free Image Generation with Trajectory Smoothing
by: Luo, Yan, et al.
Published: (2026)
by: Luo, Yan, et al.
Published: (2026)
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance
by: Benhenda, Mostapha
Published: (2026)
by: Benhenda, Mostapha
Published: (2026)
A Design Trajectory Map of Human-AI Collaborative Reinforcement Learning Systems: Survey and Taxonomy
by: Li, Zhaoxing
Published: (2024)
by: Li, Zhaoxing
Published: (2024)
Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
by: Ji, Jiaming, et al.
Published: (2023)
by: Ji, Jiaming, et al.
Published: (2023)
Towards Proactive Defense Against Cyber Cognitive Attacks
by: Rushing, Bonnie, et al.
Published: (2025)
by: Rushing, Bonnie, et al.
Published: (2025)
Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving
by: Rajaee, Sara, et al.
Published: (2025)
by: Rajaee, Sara, et al.
Published: (2025)
DeContext as Defense: Safe Image Editing in Diffusion Transformers
by: Shen, Linghui, et al.
Published: (2025)
by: Shen, Linghui, et al.
Published: (2025)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
by: Sun, Bowen, et al.
Published: (2026)
by: Sun, Bowen, et al.
Published: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
LookAhead Tuning: Safer Language Models via Partial Answer Previews
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Similar Items
-
A Game-Theoretic Negotiation Framework for Cross-Cultural Consensus in LLMs
by: Zhang, Guoxi, et al.
Published: (2025) -
VISA: Value Injection via Shielded Adaptation for Personalized LLM Alignment
by: Chen, Jiawei, et al.
Published: (2026) -
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
by: Zhou, Jiayi, et al.
Published: (2024) -
SafeLawBench: Towards Safe Alignment of Large Language Models
by: Cao, Chuxue, et al.
Published: (2025) -
SafeMT: Multi-turn Safety for Multimodal Language Models
by: Zhu, Han, et al.
Published: (2025)