A Unified Framework for the Evaluation of LLM Agentic Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Pengyu, Li, Lijun, Lyu, Yaxing, Luo, Qianxin, Yang, Jingyi, Liu, Yi, Hui, Tingfeng, Yuan, Xinyu, Sun, Li, Su, Sen, Shao, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems
by: Zhu, Pengyu, et al.
Published: (2025)
by: Zhu, Pengyu, et al.
Published: (2025)
The Necessity of a Unified Framework for LLM-Based Agent Evaluation
by: Zhu, Pengyu, et al.
Published: (2026)
by: Zhu, Pengyu, et al.
Published: (2026)
DecIF: Improving Instruction-Following through Meta-Decomposition
by: Hui, Tingfeng, et al.
Published: (2025)
by: Hui, Tingfeng, et al.
Published: (2025)
Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics
by: Hui, Tingfeng, et al.
Published: (2026)
by: Hui, Tingfeng, et al.
Published: (2026)
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
by: Sun, Jingyi, et al.
Published: (2024)
by: Sun, Jingyi, et al.
Published: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
by: Dong, Peijie, et al.
Published: (2025)
by: Dong, Peijie, et al.
Published: (2025)
STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models
by: Guo, Shaoxiong, et al.
Published: (2025)
by: Guo, Shaoxiong, et al.
Published: (2025)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
by: Li, Miao, et al.
Published: (2024)
by: Li, Miao, et al.
Published: (2024)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
by: Liu, Yuexiao, et al.
Published: (2025)
by: Liu, Yuexiao, et al.
Published: (2025)
Balancing Interpretability and Performance in Reinforcement Learning: An Adaptive Spectral Based Linear Approach
by: Yi, Qianxin, et al.
Published: (2025)
by: Yi, Qianxin, et al.
Published: (2025)
RadFabric: Agentic AI System with Reasoning Capability for Radiology
by: Chen, Wenting, et al.
Published: (2025)
by: Chen, Wenting, et al.
Published: (2025)
A Unified Agentic Framework for Evaluating Conditional Image Generation
by: Wang, Jifang, et al.
Published: (2025)
by: Wang, Jifang, et al.
Published: (2025)
Smaller Language Models Are Better Instruction Evolvers
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment
by: Zhao, Hongjue, et al.
Published: (2026)
by: Zhao, Hongjue, et al.
Published: (2026)
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration
by: Li, Chunxiao, et al.
Published: (2026)
by: Li, Chunxiao, et al.
Published: (2026)
ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning
by: Wang, Xiaoxuan, et al.
Published: (2026)
by: Wang, Xiaoxuan, et al.
Published: (2026)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent
by: Zhu, Pengyu, et al.
Published: (2025)
by: Zhu, Pengyu, et al.
Published: (2025)
Microscopic Insight of the High‐Entropy Effect on the Lithium Storage Performance and Rate Capability of Spinel Oxide
by: Man Zhao, et al.
Published: (2025)
by: Man Zhao, et al.
Published: (2025)
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
by: Wei, Qianshan, et al.
Published: (2026)
by: Wei, Qianshan, et al.
Published: (2026)
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
by: Zhang, Zaibin, et al.
Published: (2024)
by: Zhang, Zaibin, et al.
Published: (2024)
MoFa: A Unified Performance Modeling Framework for LLM Pretraining
by: Zhao, Lu, et al.
Published: (2025)
by: Zhao, Lu, et al.
Published: (2025)
Noise-BERT: A Unified Perturbation-Robust Framework with Noise Alignment Pre-training for Noisy Slot Filling Task
by: Zhao, Jinxu, et al.
Published: (2024)
by: Zhao, Jinxu, et al.
Published: (2024)
Task-Oriented Communication for Human Action Understanding via Edge-Cloud Co-Inference
by: Liu, Jingyi, et al.
Published: (2026)
by: Liu, Jingyi, et al.
Published: (2026)
Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework
by: Li, Jialin, et al.
Published: (2025)
by: Li, Jialin, et al.
Published: (2025)
Realization of Friedrich-Wintgen QBIC with high Q-factors based on acoustic-solid coupling and sensing applications
by: Wu, Bowei, et al.
Published: (2025)
by: Wu, Bowei, et al.
Published: (2025)
OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
by: Wei, Yongxian, et al.
Published: (2025)
by: Wei, Yongxian, et al.
Published: (2025)
Cooperative Visual-LiDAR Extrinsic Calibration Technology for Intersection Vehicle-Infrastructure: A review
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
Counterfactual LLM-based Framework for Measuring Rhetorical Style
by: Qiu, Jingyi, et al.
Published: (2025)
by: Qiu, Jingyi, et al.
Published: (2025)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
Similar Items
-
Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems
by: Zhu, Pengyu, et al.
Published: (2025) -
The Necessity of a Unified Framework for LLM-Based Agent Evaluation
by: Zhu, Pengyu, et al.
Published: (2026) -
DecIF: Improving Instruction-Following through Meta-Decomposition
by: Hui, Tingfeng, et al.
Published: (2025) -
Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
by: Miao, Ziqi, et al.
Published: (2025) -
STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics
by: Hui, Tingfeng, et al.
Published: (2026)