SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhou, Yifan, Zhang, Zhentao, Cheng, Ziming, Zhang, Shuo, Lan, Qizhen, Chen, Zhangquan, Yang, Zhi, QianyuXu, Chen, Ronghao, Wang, Huacan, Hu, Sen |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
par: Wu, Tingyu, et autres
Publié: (2026)
par: Wu, Tingyu, et autres
Publié: (2026)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
par: Li, Kuan, et autres
Publié: (2026)
par: Li, Kuan, et autres
Publié: (2026)
UIPress: Bringing Optical Token Compression to UI-to-Code Generation
par: Dai, Dasen, et autres
Publié: (2026)
par: Dai, Dasen, et autres
Publié: (2026)
CloneMem: Benchmarking Long-Term Memory for AI Clones
par: Hu, Sen, et autres
Publié: (2026)
par: Hu, Sen, et autres
Publié: (2026)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
par: Ni, Ziyi, et autres
Publié: (2025)
par: Ni, Ziyi, et autres
Publié: (2025)
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
par: Li, Xiangyi, et autres
Publié: (2026)
par: Li, Xiangyi, et autres
Publié: (2026)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
par: Han, Tingxu, et autres
Publié: (2026)
par: Han, Tingxu, et autres
Publié: (2026)
SkillMaster: Toward Autonomous Skill Mastery in LLM Agents
par: Yang, Min, et autres
Publié: (2026)
par: Yang, Min, et autres
Publié: (2026)
EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines
par: Zhang, Shuo, et autres
Publié: (2026)
par: Zhang, Shuo, et autres
Publié: (2026)
SkillGen: Verified Inference-Time Agent Skill Synthesis
par: Ma, Yuchen, et autres
Publié: (2026)
par: Ma, Yuchen, et autres
Publié: (2026)
MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences
par: Wang, Qihao, et autres
Publié: (2026)
par: Wang, Qihao, et autres
Publié: (2026)
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
par: Lei, Yingtie, et autres
Publié: (2026)
par: Lei, Yingtie, et autres
Publié: (2026)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
par: Jiang, Yukun, et autres
Publié: (2026)
par: Jiang, Yukun, et autres
Publié: (2026)
SkillGen: Learning Domain Skills for In-Context Sequential Decision Making
par: Ding, Ruomeng, et autres
Publié: (2025)
par: Ding, Ruomeng, et autres
Publié: (2025)
Uni-NTFM: A Unified Foundation Model for EEG Signal Representation Learning
par: Chen, Zhisheng, et autres
Publié: (2025)
par: Chen, Zhisheng, et autres
Publié: (2025)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
par: Jin, Chang, et autres
Publié: (2026)
par: Jin, Chang, et autres
Publié: (2026)
Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
par: Li, Hao, et autres
Publié: (2026)
par: Li, Hao, et autres
Publié: (2026)
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
par: Zhang, Ziao, et autres
Publié: (2026)
par: Zhang, Ziao, et autres
Publié: (2026)
EpochX: Building the Infrastructure for an Emergent Agent Civilization
par: Wang, Huacan, et autres
Publié: (2026)
par: Wang, Huacan, et autres
Publié: (2026)
SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
par: Cho, Hongcheol, et autres
Publié: (2026)
par: Cho, Hongcheol, et autres
Publié: (2026)
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
par: Zhong, Shanshan, et autres
Publié: (2026)
par: Zhong, Shanshan, et autres
Publié: (2026)
SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents
par: Hu, Wentao, et autres
Publié: (2026)
par: Hu, Wentao, et autres
Publié: (2026)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
par: Chen, Shiqi, et autres
Publié: (2026)
par: Chen, Shiqi, et autres
Publié: (2026)
SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
par: Lin, Jiaye, et autres
Publié: (2025)
par: Lin, Jiaye, et autres
Publié: (2025)
Remote Sensing-Oriented World Model
par: Lu, Yuxi, et autres
Publié: (2025)
par: Lu, Yuxi, et autres
Publié: (2025)
SkillRouter: Skill Routing for LLM Agents at Scale
par: Zheng, YanZhao, et autres
Publié: (2026)
par: Zheng, YanZhao, et autres
Publié: (2026)
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
par: Bian, Haonan, et autres
Publié: (2026)
par: Bian, Haonan, et autres
Publié: (2026)
Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives
par: Li, Daiqiang, et autres
Publié: (2026)
par: Li, Daiqiang, et autres
Publié: (2026)
Harnessing LLM Agents with Skill Programs
par: Liu, Hongjun, et autres
Publié: (2026)
par: Liu, Hongjun, et autres
Publié: (2026)
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
par: Zhou, Chenyu, et autres
Publié: (2026)
par: Zhou, Chenyu, et autres
Publié: (2026)
FORTIS: Benchmarking Over-Privilege in Agent Skills
par: Li, Shawn, et autres
Publié: (2026)
par: Li, Shawn, et autres
Publié: (2026)
SkillTester: Benchmarking Utility and Security of Agent Skills
par: Wang, Leye, et autres
Publié: (2026)
par: Wang, Leye, et autres
Publié: (2026)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
par: Wang, Jingxing, et autres
Publié: (2026)
par: Wang, Jingxing, et autres
Publié: (2026)
HomeFlow: A Data Flywheel for Smart Home Agent Training with Verifiable Simulation
par: Gu, Yi, et autres
Publié: (2026)
par: Gu, Yi, et autres
Publié: (2026)
Formal Skill: Programmable Runtime Skills for Efficient and Accurate LLM Agents
par: Zhang, Xi, et autres
Publié: (2026)
par: Zhang, Xi, et autres
Publié: (2026)
GraSP: Graph-Structured Skill Compositions for LLM Agents
par: Xia, Tianle, et autres
Publié: (2026)
par: Xia, Tianle, et autres
Publié: (2026)
SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision
par: Liu, Yuxuan, et autres
Publié: (2026)
par: Liu, Yuxuan, et autres
Publié: (2026)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
par: Yang, Zhi, et autres
Publié: (2026)
par: Yang, Zhi, et autres
Publié: (2026)
QuantaAlpha: An Evolutionary Framework for LLM-Driven Alpha Mining
par: Han, Jun, et autres
Publié: (2026)
par: Han, Jun, et autres
Publié: (2026)
CUA-Skill: Develop Skills for Computer Using Agent
par: Chen, Tianyi, et autres
Publié: (2026)
par: Chen, Tianyi, et autres
Publié: (2026)
Documents similaires
-
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
par: Wu, Tingyu, et autres
Publié: (2026) -
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
par: Li, Kuan, et autres
Publié: (2026) -
UIPress: Bringing Optical Token Compression to UI-to-Code Generation
par: Dai, Dasen, et autres
Publié: (2026) -
CloneMem: Benchmarking Long-Term Memory for AI Clones
par: Hu, Sen, et autres
Publié: (2026) -
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
par: Ni, Ziyi, et autres
Publié: (2025)