SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Cuadron, Alejandro, Yu, Pengfei, Liu, Yang, Gupta, Arpit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents
por: Zhen, Shuai, et al.
Publicado: (2026)
por: Zhen, Shuai, et al.
Publicado: (2026)
HashAttention: Semantic Sparsity for Faster Inference
por: Desai, Aditya, et al.
Publicado: (2024)
por: Desai, Aditya, et al.
Publicado: (2024)
Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment
por: Wang, Haozhong, et al.
Publicado: (2026)
por: Wang, Haozhong, et al.
Publicado: (2026)
CORA: Conformal Risk-Controlled Agents for Safeguarded Mobile GUI Automation
por: Feng, Yushi, et al.
Publicado: (2026)
por: Feng, Yushi, et al.
Publicado: (2026)
Learnware of Language Models: Specialized Small Language Models Can Do Big
por: Tan, Zhi-Hao, et al.
Publicado: (2025)
por: Tan, Zhi-Hao, et al.
Publicado: (2025)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
GuardReasoner: Towards Reasoning-based LLM Safeguards
por: Liu, Yue, et al.
Publicado: (2025)
por: Liu, Yue, et al.
Publicado: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
vAttention: Verified Sparse Attention
por: Desai, Aditya, et al.
Publicado: (2025)
por: Desai, Aditya, et al.
Publicado: (2025)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
por: Zhao, Kai, et al.
Publicado: (2025)
por: Zhao, Kai, et al.
Publicado: (2025)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
por: Lewis, Ashley, et al.
Publicado: (2025)
por: Lewis, Ashley, et al.
Publicado: (2025)
Verifiability-First Agents: Provable Observability and Lightweight Audit Agents for Controlling Autonomous LLM Systems
por: Gupta, Abhivansh
Publicado: (2025)
por: Gupta, Abhivansh
Publicado: (2025)
Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
por: Gupta, Prakhar, et al.
Publicado: (2025)
por: Gupta, Prakhar, et al.
Publicado: (2025)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
por: Xiong, Weimin, et al.
Publicado: (2024)
por: Xiong, Weimin, et al.
Publicado: (2024)
One-Step Graph-Structured Neural Flows for Irregular Multivariate Time Series Classification
por: Gao, Mengzhou, et al.
Publicado: (2026)
por: Gao, Mengzhou, et al.
Publicado: (2026)
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
por: Ye, Wen, et al.
Publicado: (2025)
por: Ye, Wen, et al.
Publicado: (2025)
QDeepGR4J: Quantile-based ensemble of deep learning and GR4J hybrid rainfall-runoff models for extreme flow prediction with uncertainty quantification
por: Kapoor, Arpit, et al.
Publicado: (2025)
por: Kapoor, Arpit, et al.
Publicado: (2025)
Uncertainty in Action: Confidence Elicitation in Embodied Agents
por: Yu, Tianjiao, et al.
Publicado: (2025)
por: Yu, Tianjiao, et al.
Publicado: (2025)
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
por: Song, Yifan, et al.
Publicado: (2024)
por: Song, Yifan, et al.
Publicado: (2024)
SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
por: Rahman, Salman, et al.
Publicado: (2025)
por: Rahman, Salman, et al.
Publicado: (2025)
An Example Safety Case for Safeguards Against Misuse
por: Clymer, Joshua, et al.
Publicado: (2025)
por: Clymer, Joshua, et al.
Publicado: (2025)
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
por: Wang, Hanlin, et al.
Publicado: (2025)
por: Wang, Hanlin, et al.
Publicado: (2025)
Two-Step Offline Preference-Based Reinforcement Learning with Constrained Actions
por: Xu, Yinglun, et al.
Publicado: (2023)
por: Xu, Yinglun, et al.
Publicado: (2023)
GoAgent: Group-of-Agents Communication Topology Generation for LLM-based Multi-Agent Systems
por: Chen, Hongjiang, et al.
Publicado: (2026)
por: Chen, Hongjiang, et al.
Publicado: (2026)
PoPE: Legendre Orthogonal Polynomials Based Position Encoding for Large Language Models
por: Aggarwal, Arpit
Publicado: (2024)
por: Aggarwal, Arpit
Publicado: (2024)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
por: Yan, Bingyu, et al.
Publicado: (2026)
por: Yan, Bingyu, et al.
Publicado: (2026)
Bellman Error Centering
por: Chen, Xingguo, et al.
Publicado: (2025)
por: Chen, Xingguo, et al.
Publicado: (2025)
PFGuard: A Generative Framework with Privacy and Fairness Safeguards
por: Kim, Soyeon, et al.
Publicado: (2024)
por: Kim, Soyeon, et al.
Publicado: (2024)
Multi-Agent Decision Transformers for Dynamic Dispatching in Material Handling Systems Leveraging Enterprise Big Data
por: Lee, Xian Yeow, et al.
Publicado: (2024)
por: Lee, Xian Yeow, et al.
Publicado: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
por: Halawi, Danny, et al.
Publicado: (2024)
por: Halawi, Danny, et al.
Publicado: (2024)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
por: Rezaei, Mohammad, et al.
Publicado: (2026)
por: Rezaei, Mohammad, et al.
Publicado: (2026)
Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning
por: Yang, Zhicheng, et al.
Publicado: (2026)
por: Yang, Zhicheng, et al.
Publicado: (2026)
FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents
por: Li, Qizheng, et al.
Publicado: (2026)
por: Li, Qizheng, et al.
Publicado: (2026)
Step-by-Step Causality: Transparent Causal Discovery with Multi-Agent Tree-Query and Adversarial Confidence Estimation
por: Ding, Ziyi, et al.
Publicado: (2026)
por: Ding, Ziyi, et al.
Publicado: (2026)
Agent-Omit: Adaptive Context Omission for Efficient LLM Agents
por: Ning, Yansong, et al.
Publicado: (2026)
por: Ning, Yansong, et al.
Publicado: (2026)
GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
por: Wu, Xiongbin, et al.
Publicado: (2026)
por: Wu, Xiongbin, et al.
Publicado: (2026)
Think like a Scientist: Physics-guided LLM Agent for Equation Discovery
por: Yang, Jianke, et al.
Publicado: (2026)
por: Yang, Jianke, et al.
Publicado: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
por: Wang, Bin, et al.
Publicado: (2026)
por: Wang, Bin, et al.
Publicado: (2026)
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
por: Pang, Jing-Cheng, et al.
Publicado: (2021)
por: Pang, Jing-Cheng, et al.
Publicado: (2021)
Ejemplares similares
-
Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents
por: Zhen, Shuai, et al.
Publicado: (2026) -
HashAttention: Semantic Sparsity for Faster Inference
por: Desai, Aditya, et al.
Publicado: (2024) -
Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment
por: Wang, Haozhong, et al.
Publicado: (2026) -
CORA: Conformal Risk-Controlled Agents for Safeguarded Mobile GUI Automation
por: Feng, Yushi, et al.
Publicado: (2026) -
Learnware of Language Models: Specialized Small Language Models Can Do Big
por: Tan, Zhi-Hao, et al.
Publicado: (2025)