PEAR: Planner-Executor Agent Robustness Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Shen, Zhang, Mingxuan, He, Pengfei, Ma, Li, Thuraisingham, Bhavani, Liu, Hui, Xing, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MUBox: A Critical Evaluation Framework of Deep Machine Unlearning
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
by: Berrayana, Lina, et al.
Published: (2025)
by: Berrayana, Lina, et al.
Published: (2025)
Robust and Explainable Divide-and-Conquer Learning for Intrusion Detection
by: Zhou, Yan, et al.
Published: (2026)
by: Zhou, Yan, et al.
Published: (2026)
DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
by: Qiu, Han, et al.
Published: (2020)
by: Qiu, Han, et al.
Published: (2020)
Memory Injection Attacks on LLM Agents via Query-Only Interaction
by: Dong, Shen, et al.
Published: (2025)
by: Dong, Shen, et al.
Published: (2025)
PEAR: Equal Area Weather Forecasting on the Sphere
by: Linander, Hampus, et al.
Published: (2025)
by: Linander, Hampus, et al.
Published: (2025)
PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement Learning
by: Singh, Utsav, et al.
Published: (2023)
by: Singh, Utsav, et al.
Published: (2023)
RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks
by: Yan, Mingxuan, et al.
Published: (2025)
by: Yan, Mingxuan, et al.
Published: (2025)
Towards the Effect of Examples on In-Context Learning: A Theoretical Case Study
by: He, Pengfei, et al.
Published: (2024)
by: He, Pengfei, et al.
Published: (2024)
Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor Attacks
by: He, Bowei, et al.
Published: (2025)
by: He, Bowei, et al.
Published: (2025)
LaMMA-P: Generalizable Multi-Agent Long-Horizon Task Allocation and Planning with LM-Driven PDDL Planner
by: Zhang, Xiaopan, et al.
Published: (2024)
by: Zhang, Xiaopan, et al.
Published: (2024)
Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AI
by: Wu, Feiyu, et al.
Published: (2026)
by: Wu, Feiyu, et al.
Published: (2026)
Less is more -- the Dispatcher/ Executor principle for multi-task Reinforcement Learning
by: Riedmiller, Martin, et al.
Published: (2023)
by: Riedmiller, Martin, et al.
Published: (2023)
SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors
by: Lyu, Bohan, et al.
Published: (2025)
by: Lyu, Bohan, et al.
Published: (2025)
Make LLMs better zero-shot reasoners: Structure-orientated autonomous reasoning
by: He, Pengfei, et al.
Published: (2024)
by: He, Pengfei, et al.
Published: (2024)
Causal Flow Q-Learning for Robust Offline Reinforcement Learning
by: Li, Mingxuan, et al.
Published: (2026)
by: Li, Mingxuan, et al.
Published: (2026)
Crafting Reversible SFT Behaviors in Large Language Models
by: Lin, Yuping, et al.
Published: (2026)
by: Lin, Yuping, et al.
Published: (2026)
PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs
by: Uysal, Erdem, et al.
Published: (2026)
by: Uysal, Erdem, et al.
Published: (2026)
What Makes a Good Diffusion Planner for Decision Making?
by: Lu, Haofei, et al.
Published: (2025)
by: Lu, Haofei, et al.
Published: (2025)
Ensuring Calibration Robustness in Split Conformal Prediction Under Adversarial Attacks
by: Qian, Xunlei, et al.
Published: (2025)
by: Qian, Xunlei, et al.
Published: (2025)
A Comprehensive Graph Pooling Benchmark: Effectiveness, Robustness and Generalizability
by: Wang, Pengyun, et al.
Published: (2024)
by: Wang, Pengyun, et al.
Published: (2024)
PRIMER: Perception-Aware Robust Learning-based Multiagent Trajectory Planner
by: Kondo, Kota, et al.
Published: (2024)
by: Kondo, Kota, et al.
Published: (2024)
FTimeXer: Frequency-aware Time-series Transformer with Exogenous variables for Robust Carbon Footprint Forecasting
by: Li, Qingzhong, et al.
Published: (2026)
by: Li, Qingzhong, et al.
Published: (2026)
Efficient Agent Training for Computer Use
by: He, Yanheng, et al.
Published: (2025)
by: He, Yanheng, et al.
Published: (2025)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
by: Xing, Zeyu, et al.
Published: (2026)
by: Xing, Zeyu, et al.
Published: (2026)
Can We Rely on LLM Agents to Draft Long-Horizon Plans? Let's Take TravelPlanner as an Example
by: Chen, Yanan, et al.
Published: (2024)
by: Chen, Yanan, et al.
Published: (2024)
Confounding Robust Continuous Control via Automatic Reward Shaping
by: Juliani, Mateo, et al.
Published: (2026)
by: Juliani, Mateo, et al.
Published: (2026)
Circuit Transformer: A Transformer That Preserves Logical Equivalence
by: Li, Xihan, et al.
Published: (2024)
by: Li, Xihan, et al.
Published: (2024)
A Simple Plug-in for Improving Eviction-Based KV Cache Compression
by: Lin, Yuping, et al.
Published: (2026)
by: Lin, Yuping, et al.
Published: (2026)
When Planners Meet Reality: How Learned, Reactive Traffic Agents Shift nuPlan Benchmarks
by: Hagedorn, Steffen, et al.
Published: (2025)
by: Hagedorn, Steffen, et al.
Published: (2025)
Multi-Faceted Studies on Data Poisoning can Advance LLM Development
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
Impact of Positional Encoding: Clean and Adversarial Rademacher Complexity for Transformers under In-Context Regression
by: He, Weiyi, et al.
Published: (2025)
by: He, Weiyi, et al.
Published: (2025)
Off-dynamics Conditional Diffusion Planners
by: Ng, Wen Zheng Terence, et al.
Published: (2024)
by: Ng, Wen Zheng Terence, et al.
Published: (2024)
ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
by: Yu, Han, et al.
Published: (2025)
by: Yu, Han, et al.
Published: (2025)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
by: Younesian, Sharareh, et al.
Published: (2026)
by: Younesian, Sharareh, et al.
Published: (2026)
Bidirectional Distillation: A Mixed-Play Framework for Multi-Agent Generalizable Behaviors
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
Knowledge Discovery in Surveys using Machine Learning: A Case Study of Women in Entrepreneurship in UAE
by: Ahmad, Syed Farhan, et al.
Published: (2021)
by: Ahmad, Syed Farhan, et al.
Published: (2021)
How to Enhance Downstream Adversarial Robustness (almost) without Touching the Pre-Trained Foundation Model?
by: Liu, Meiqi, et al.
Published: (2025)
by: Liu, Meiqi, et al.
Published: (2025)
SmartPlay: A Benchmark for LLMs as Intelligent Agents
by: Wu, Yue, et al.
Published: (2023)
by: Wu, Yue, et al.
Published: (2023)
Similar Items
-
MUBox: A Critical Evaluation Framework of Deep Machine Unlearning
by: Li, Xiang, et al.
Published: (2025) -
Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
by: Berrayana, Lina, et al.
Published: (2025) -
Robust and Explainable Divide-and-Conquer Learning for Intrusion Detection
by: Zhou, Yan, et al.
Published: (2026) -
DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
by: Qiu, Han, et al.
Published: (2020) -
Memory Injection Attacks on LLM Agents via Query-Only Interaction
by: Dong, Shen, et al.
Published: (2025)