DRAFT: Task Decoupled Latent Reasoning for Agent Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Lin, Fang, Junfeng, Zhang, Dan, Shen, Fei, Wang, Xiang, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
by: Fang, Junfeng, et al.
Published: (2026)
by: Fang, Junfeng, et al.
Published: (2026)
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
by: Wang, Zhaoxin, et al.
Published: (2026)
by: Wang, Zhaoxin, et al.
Published: (2026)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
by: Sheng, Leheng, et al.
Published: (2026)
by: Sheng, Leheng, et al.
Published: (2026)
NextMem: Towards Latent Factual Memory for LLM-based Agents
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
by: He, Jinghan, et al.
Published: (2026)
by: He, Jinghan, et al.
Published: (2026)
Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling
by: Luo, Yanchen, et al.
Published: (2025)
by: Luo, Yanchen, et al.
Published: (2025)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
by: Sheng, Leheng, et al.
Published: (2025)
by: Sheng, Leheng, et al.
Published: (2025)
Are Reasoning Models More Prone to Hallucination?
by: Yao, Zijun, et al.
Published: (2025)
by: Yao, Zijun, et al.
Published: (2025)
Rubric-based On-policy Distillation
by: Fang, Junfeng, et al.
Published: (2026)
by: Fang, Junfeng, et al.
Published: (2026)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
The Missing Half: Unveiling Training-time Implicit Safety Risks Beyond Deployment
by: Zhang, Zhexin, et al.
Published: (2026)
by: Zhang, Zhexin, et al.
Published: (2026)
Rethinking Tokenizer and Decoder in Masked Graph Modeling for Molecules
by: Liu, Zhiyuan, et al.
Published: (2023)
by: Liu, Zhiyuan, et al.
Published: (2023)
NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule Generation
by: Liu, Zhiyuan, et al.
Published: (2025)
by: Liu, Zhiyuan, et al.
Published: (2025)
Enhancing Spectral Graph Neural Networks with LLM-Predicted Homophily
by: Lu, Kangkang, et al.
Published: (2025)
by: Lu, Kangkang, et al.
Published: (2025)
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
by: Yang, Yonghui, et al.
Published: (2026)
by: Yang, Yonghui, et al.
Published: (2026)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
by: Ye, Hao, et al.
Published: (2026)
by: Ye, Hao, et al.
Published: (2026)
NExT-GPT: Any-to-Any Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2023)
by: Wu, Shengqiong, et al.
Published: (2023)
Continual Multimodal Contrastive Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Less Approximates More: Harmonizing Performance and Confidence Faithfulness via Hybrid Post-Training for High-Stakes Tasks
by: Ma, Haokai, et al.
Published: (2026)
by: Ma, Haokai, et al.
Published: (2026)
CARL: Criticality-Aware Agentic Reinforcement Learning
by: Shen, Leyang, et al.
Published: (2025)
by: Shen, Leyang, et al.
Published: (2025)
Reasoning on Time-Series for Financial Technical Analysis
by: Koa, Kelvin J. L., et al.
Published: (2025)
by: Koa, Kelvin J. L., et al.
Published: (2025)
Bridging Jensen Gap for Max-Min Group Fairness Optimization in Recommendation
by: Xu, Chen, et al.
Published: (2025)
by: Xu, Chen, et al.
Published: (2025)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
by: Li, Moxin, et al.
Published: (2025)
by: Li, Moxin, et al.
Published: (2025)
Principled Multimodal Representation Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
GOODAT: Towards Test-time Graph Out-of-Distribution Detection
by: Wang, Luzhi, et al.
Published: (2024)
by: Wang, Luzhi, et al.
Published: (2024)
Towards 3D Molecule-Text Interpretation in Language Models
by: Li, Sihang, et al.
Published: (2024)
by: Li, Sihang, et al.
Published: (2024)
Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
by: Deng, Yang, et al.
Published: (2024)
by: Deng, Yang, et al.
Published: (2024)
ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning
by: Zhang, Jinyang, et al.
Published: (2025)
by: Zhang, Jinyang, et al.
Published: (2025)
Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models
by: Koa, Kelvin J. L., et al.
Published: (2024)
by: Koa, Kelvin J. L., et al.
Published: (2024)
Addressing Graph Heterogeneity and Heterophily from A Spectral Perspective
by: Lu, Kangkang, et al.
Published: (2024)
by: Lu, Kangkang, et al.
Published: (2024)
DRAFT-ing Architectural Design Decisions using LLMs
by: Dhar, Rudra, et al.
Published: (2025)
by: Dhar, Rudra, et al.
Published: (2025)
Temporal Relational Reasoning of Large Language Models for Detecting Stock Portfolio Crashes
by: Koa, Kelvin J. L., et al.
Published: (2024)
by: Koa, Kelvin J. L., et al.
Published: (2024)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Towards Modality Generalization: A Benchmark and Prospective Analysis
by: Liu, Xiaohao, et al.
Published: (2024)
by: Liu, Xiaohao, et al.
Published: (2024)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
by: Qu, Leigang, et al.
Published: (2025)
by: Qu, Leigang, et al.
Published: (2025)
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
by: Qiu, Haiyi, et al.
Published: (2024)
by: Qiu, Haiyi, et al.
Published: (2024)
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
by: Zhou, Zhenglin, et al.
Published: (2025)
by: Zhou, Zhenglin, et al.
Published: (2025)
Optimize Incompatible Parameters through Compatibility-aware Knowledge Integration
by: Lv, Zheqi, et al.
Published: (2025)
by: Lv, Zheqi, et al.
Published: (2025)
Similar Items
-
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
by: Fang, Junfeng, et al.
Published: (2025) -
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025) -
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
by: Fang, Junfeng, et al.
Published: (2026) -
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
by: Wang, Zhaoxin, et al.
Published: (2026) -
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
by: Sheng, Leheng, et al.
Published: (2026)